Knowledge base Developers
Chat, Responses and Messages
Three ways to ask a model something, in the three formats clients already speak, all running on the same models and billed the same way.
Chat Completions
The OpenAI chat format, and the one almost every tool supports. Send the conversation so far and get the next message back.
POST https://nymbot.ai/api/v1/chat/completions — needs an API key.
Only model and messages are required. A sampling parameter the chosen
model does not take is dropped without an error, so one request body works across models. The
model list gives each model's supported_parameters. What is refused rather than
dropped is anything the model cannot do at all: pictures for a model that cannot see, tools for
a model that cannot call them.
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | nymbot/auto (or auto) for Nymbot's routing on the standard balance, or a catalog model id such as anthropic/claude-sonnet-5 on the Pro balance. The app's short names and aliases are accepted too. May end in a suffix. |
messages | array | Yes | The conversation. Roles system, developer (treated as system), user, assistant and tool. Content is a string or a list of text and image_url parts; pictures only in user messages. input_audio and file parts are refused. |
stream | boolean | No | Send the answer as it is written. See streaming. |
stream_options | object | No | {"include_usage": true} adds a final chunk with token counts and the cost. |
max_tokensmax_completion_tokens | integer | No | The most tokens to write. Lowered to the model's maximum if higher. Also sets how much is held from your balance, so a smaller number needs less credit to start. |
temperaturetop_p | number | No | Sampling controls: temperature from 0 to 2, top_p from 0 to 1. |
stop | string or array | No | Text that ends the answer: a string or at most 4 strings, each at most 256 characters. Not used by nymbot/auto. |
seed | integer | No | For repeatable sampling, where the model supports it. |
presence_penaltyfrequency_penalty | number | No | Repetition controls, each from -2 to 2. |
response_format | object | No | {"type": "json_object"} or {"type": "json_schema", "json_schema": {…}}, where the model supports it. Not used by nymbot/auto. |
toolstool_choiceparallel_tool_calls | array, string or object, boolean | No | Function calling. See tool calls. A web_search tool turns on web search. At most 128 tools and 512 KB of definitions, nested at most 64 levels deep; more is a 400. |
reasoning_effortreasoning | string, object | No | "minimal", "low", "medium" or "high", or {"effort": "high"}. "none" or {"enabled": false} turns it off. See reasoning. |
plugins | array | No | [{"id": "web", "max_results": 5}] always searches the web first. Up to 10 results. |
n | integer | No | Only 1. Anything else returns 400. |
logit_biasusermetadata | object, string, object | No | Accepted and not sent on. logit_bias maps at most 300 token ids to numbers from -100 to 100; user is at most 256 characters; metadata holds at most 16 string values, keys up to 64 characters and values up to 512. |
The reply is an ordinary chat.completion, with the cost in usage.cost
(in dollars) and in the nymbot object. model is the resolved model id,
so a short name comes back as the full one. If the model reasoned before answering, its
reasoning is in message.reasoning_content, separate from the answer.
Response
{
"id": "chatcmpl-5f1c0a9e27d84b3c",
"object": "chat.completion",
"created": 1790726400,
"model": "anthropic/claude-sonnet-5",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "A Lightning invoice is a one-time payment request..."
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1240,
"completion_tokens": 380,
"total_tokens": 1620,
"prompt_tokens_details": { "cached_tokens": 0 },
"completion_tokens_details": { "reasoning_tokens": 0 },
"cost": 0.01895
},
"nymbot": {
"balance": "pro",
"charged_credits": 0.162,
"charged_sats": 16.2,
"balance_credits": 412.425,
"balance_sats": 41242.5
}
}
finish_reason is worked out from what came back, since providers report it
differently: tool_calls when the model asked for tools, length when
the answer used every token allowed (or the model spent them all reasoning and wrote no
answer), content_filter when the provider refused, and stop
otherwise. usage.completion_tokens_details.reasoning_tokens is always 0: hidden
reasoning tokens are counted, and charged, in completion_tokens.
| Status | When |
|---|---|
400 | No model or messages (missing_required_parameter); an audio or file part, or pictures for a model that cannot see them (unsupported_content); more than 20 pictures (too_many_images); a sampling parameter of the wrong type or out of range (invalid_value); a picture link that is not public (invalid_image_url); tools on a model that cannot call them (unsupported_tool); n other than 1; a :thinking suffix on a model that cannot reason (model_not_found); or the provider refused the request (upstream_rejected). |
402 | The balance the model spends cannot cover the hold. |
403 | The worst case does not fit the key's cap (key_limit_reached). Lower max_tokens or raise the cap. |
404 | No model by that name (model_not_found). |
429 | The key's rate limit, or the provider's (upstream_rate_limited). |
502, 503 | The provider failed (upstream_error) or is overloaded (upstream_overloaded, with Retry-After). Nothing is charged unless the provider billed for the attempt. |
The status codes every endpoint shares are listed under errors.
cURL
curl https://nymbot.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NYMBOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"messages": [
{"role": "system", "content": "Answer in one short paragraph."},
{"role": "user", "content": "What is a Lightning invoice?"}
],
"max_tokens": 400
}'
Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])
reply = client.chat.completions.create(
model="anthropic/claude-sonnet-5",
messages=[
{"role": "system", "content": "Answer in one short paragraph."},
{"role": "user", "content": "What is a Lightning invoice?"},
],
max_tokens=400,
)
print(reply.choices[0].message.content)
print(reply.usage.cost)
JavaScript
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });
const reply = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5",
messages: [
{ role: "system", content: "Answer in one short paragraph." },
{ role: "user", content: "What is a Lightning invoice?" },
],
max_tokens: 400,
});
console.log(reply.choices[0].message.content);
console.log(reply.usage.cost);
Streaming
With "stream": true the answer arrives as server-sent events while the model writes
it. Each event is a chat.completion.chunk on a data: line, and the
stream ends with data: [DONE]. Reasoning arrives in
delta.reasoning_content, the answer in delta.content.
Stream anything that may take more than about 100 seconds, such as a long answer, a large
max_tokens or a reasoning model. A request that is not streamed sends nothing until
the answer is complete, and the network between you and Nymbot may close an idle connection
after about 100 seconds; the model still finishes and what it wrote is charged, but the answer
is lost. A stream sends keep-alives, so it stays open for as long as the model writes.
Stream
data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{"content":"A Lightning invoice"},"finish_reason":null}]}
: keep-alive
data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[],"usage":{"prompt_tokens":1240,"completion_tokens":380,"total_tokens":1620,"cost":0.01895},"nymbot":{"balance":"pro","charged_credits":0.162,"charged_sats":16.2,"balance_credits":412.425,"balance_sats":41242.5}}
data: [DONE]
- Lines starting with
:are keep-alive comments, sent every 15 seconds while the model is thinking. SSE clients skip them. - Ask for
"stream_options": {"include_usage": true}to get the last chunk above, with an emptychoices, the token counts, the cost and thenymbotobject. - The charge settles after the stream ends. If you close the connection early, the model is not stopped: Nymbot reads the rest of the provider's stream, for up to 25 seconds, to get its token count, and you pay what the provider reports. Without that count the charge is estimated from your input and what was written, plus the whole output allowance for a model that reasons.
- An error before the first chunk, such as
401or402, comes back as ordinary JSON with its status code, not as a stream. An error after the stream has started arrives as a lastdata: {"error": {…}}event, and the stream ends without[DONE]. - Requests with tools, and models on OpenAI's Responses transport, do not stream from the provider. They still answer with a valid stream, sent once the answer is complete.
cURL
curl -N https://nymbot.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NYMBOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nymbot/auto",
"stream": true,
"stream_options": {"include_usage": true},
"messages": [{"role": "user", "content": "Write a haiku about sats."}]
}'
Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])
stream = client.chat.completions.create(
model="nymbot/auto",
stream=True,
stream_options={"include_usage": True},
messages=[{"role": "user", "content": "Write a haiku about sats."}],
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)
elif chunk.usage:
print("\ncost:", chunk.usage.cost)
JavaScript
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });
const stream = await client.chat.completions.create({
model: "nymbot/auto",
stream: true,
stream_options: { include_usage: true },
messages: [{ role: "user", content: "Write a haiku about sats." }],
});
for await (const chunk of stream) {
if (chunk.choices.length) process.stdout.write(chunk.choices[0].delta.content || "");
else if (chunk.usage) console.log("\ncost:", chunk.usage.cost);
}
Tool calls
Describe functions in tools and the model can ask for one to be called instead of
answering. You run the function, add its result as a tool message with the same
tool_call_id, and send the conversation again. Nymbot never runs your functions;
it passes the model's request back to you.
tool_choice takes "auto", "none",
"required" or {"type": "function", "function": {"name": "…"}}.
The model list marks which models can call tools (capabilities.tools). Tools are
refused with 400 unsupported_tool on nymbot/auto and on
the few catalog models that run on OpenAI's Responses transport, which Nymbot cannot pass tools
to.
With "stream": true, a request with tools runs in one piece and is then sent as a
normal chunk sequence: the role, one chunk carrying every tool call with its index, id, name and
arguments, and the finish chunk. Clients that read streamed tool calls handle it as usual.
Response, in part
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_7d2e",
"type": "function",
"function": { "name": "get_invoice_status", "arguments": "{\"invoice_id\":\"a41f\"}" }
}
]
},
"finish_reason": "tool_calls"
}
]
cURL
curl https://nymbot.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NYMBOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"messages": [{"role": "user", "content": "Has invoice a41f been paid?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_invoice_status",
"description": "Look up whether an invoice is paid.",
"parameters": {
"type": "object",
"properties": {"invoice_id": {"type": "string"}},
"required": ["invoice_id"]
}
}
}]
}'
Python
import json, os
from openai import OpenAI
client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])
tools = [{
"type": "function",
"function": {
"name": "get_invoice_status",
"description": "Look up whether an invoice is paid.",
"parameters": {
"type": "object",
"properties": {"invoice_id": {"type": "string"}},
"required": ["invoice_id"],
},
},
}]
messages = [{"role": "user", "content": "Has invoice a41f been paid?"}]
reply = client.chat.completions.create(model="anthropic/claude-sonnet-5", messages=messages, tools=tools)
call = reply.choices[0].message.tool_calls[0]
args = json.loads(call.function.arguments)
messages.append(reply.choices[0].message)
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps({"paid": True})})
final = client.chat.completions.create(model="anthropic/claude-sonnet-5", messages=messages, tools=tools)
print(final.choices[0].message.content)
JavaScript
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });
const tools = [{
type: "function",
function: {
name: "get_invoice_status",
description: "Look up whether an invoice is paid.",
parameters: {
type: "object",
properties: { invoice_id: { type: "string" } },
required: ["invoice_id"],
},
},
}];
const messages = [{ role: "user", content: "Has invoice a41f been paid?" }];
const reply = await client.chat.completions.create({ model: "anthropic/claude-sonnet-5", messages, tools });
const call = reply.choices[0].message.tool_calls[0];
const args = JSON.parse(call.function.arguments);
messages.push(reply.choices[0].message);
messages.push({ role: "tool", tool_call_id: call.id, content: JSON.stringify({ paid: true }) });
const final = await client.chat.completions.create({ model: "anthropic/claude-sonnet-5", messages, tools });
console.log(final.choices[0].message.content);
Pictures in a request
Models with capabilities.vision can read pictures. Add an image_url
part to a user message, with either a public https:// link or a
data:image/…;base64, URL. Up to 20 pictures per request. SVG pictures are
refused, and so is a link over 4,096 characters or one with a user name or password in it
(400 invalid_image_url). An optional detail is
auto, low or high.
With nymbot/auto, a request with a picture in it is routed to a standard model that
can see. A catalog model without capabilities.vision refuses pictures with
400 unsupported_content. Pictures can only be in user messages. A
link must point to a public host; Nymbot passes the picture to the model's provider and does
not keep it.
A picture is charged as the input tokens the provider counts for it, like the rest of the request.
cURL
curl https://nymbot.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NYMBOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": "https://example.com/receipt.jpg"}}
]
}]
}'
Python
import base64, os
from openai import OpenAI
client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])
with open("receipt.jpg", "rb") as f:
data_url = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode()
reply = client.chat.completions.create(
model="anthropic/claude-sonnet-5",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": data_url}},
],
}],
)
print(reply.choices[0].message.content)
JavaScript
import { readFile } from "node:fs/promises";
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });
const dataUrl = "data:image/jpeg;base64," + (await readFile("receipt.jpg")).toString("base64");
const reply = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5",
messages: [{
role: "user",
content: [
{ type: "text", text: "What is in this picture?" },
{ type: "image_url", image_url: { url: dataUrl } },
],
}],
});
console.log(reply.choices[0].message.content);
Reasoning
Models with capabilities.reasoning can think before they answer. Ask for more or less
of it with reasoning_effort ("minimal", "low",
"medium" or "high") or "reasoning": {"effort": "high"}, or
add :thinking to the model name, which means high effort. On a model without
reasoning the setting is ignored. With nymbot/auto, :thinking sends the
request to the standard reasoning route.
On Anthropic models the effort becomes a thinking budget of about 1,000, 2,000, 8,000 or 16,000
tokens, never more than max_tokens allows. Thinking is left off when
tool_choice forces a tool, and when the request continues a tool loop (its last
message is a tool result), because Anthropic needs the earlier signed thinking to resume it.
The same applies on the Responses and Messages endpoints.
The reasoning comes back in message.reasoning_content, or
delta.reasoning_content when streaming, never mixed into the answer. Reasoning is
output and is charged as such, inside completion_tokens. Some providers do not
return the reasoning text at all, and it is still charged.
cURL
curl https://nymbot.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NYMBOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"reasoning_effort": "high",
"messages": [{"role": "user", "content": "Is 2^61 - 1 prime? Show why."}]
}'
Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])
reply = client.chat.completions.create(
model="anthropic/claude-sonnet-5",
reasoning_effort="high",
messages=[{"role": "user", "content": "Is 2^61 - 1 prime? Show why."}],
)
message = reply.choices[0].message
print(getattr(message, "reasoning_content", None))
print(message.content)
JavaScript
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });
const reply = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5",
reasoning_effort: "high",
messages: [{ role: "user", content: "Is 2^61 - 1 prime? Show why." }],
});
console.log(reply.choices[0].message.reasoning_content);
console.log(reply.choices[0].message.content);
Model suffixes
A suffix on the model name changes how the request is handled, without another field. They
work on full ids and short names alike, as in anthropic/claude-sonnet-5:online.
| Suffix | Effect |
|---|---|
:online | Searches the web first, like plugins: [{"id": "web"}]. See web search. |
:thinking | High reasoning effort; on nymbot/auto, the reasoning route. On a model that cannot reason, 400 model_not_found with “no endpoints found”. |
:nitro, :floor, :exacto, :extended | Accepted and ignored. Each catalog model has one route, so there is no faster, cheaper or longer one to pick; the suffixes are allowed so model names copied from other services still work. |
Any other suffix is ignored. The whole name is tried first, so a model whose id really contains a
colon still works; failing that, suffixes are taken off the end one at a time until a model
matches. A model name may be at most 200 characters long with at most 4 suffixes; a longer
one is refused with 400 invalid_value.
cURL
curl https://nymbot.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NYMBOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5:thinking",
"messages": [{"role": "user", "content": "Plan a three-day trip to Lisbon."}]
}'
Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])
reply = client.chat.completions.create(
model="anthropic/claude-sonnet-5:thinking",
messages=[{"role": "user", "content": "Plan a three-day trip to Lisbon."}],
)
print(reply.choices[0].message.content)
JavaScript
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });
const reply = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5:thinking",
messages: [{ role: "user", content: "Plan a three-day trip to Lisbon." }],
});
console.log(reply.choices[0].message.content);
Web search
Any chat model can answer from the live web. Nymbot searches for the last user message (its first 2,000 characters), reads the best pages and gives them to the model with the question, marked as outside content the model should not take instructions from. There are two modes:
- Always search:
"plugins": [{"id": "web", "max_results": 5}], or a:onlinesuffix on the model. - Search when it helps:
"tools": [{"type": "web_search", "parameters": {"max_results": 5}}], also accepted asweb_search_previeworopenrouter:web_search. Nymbot searches only when the question looks like it needs current information, the same test the app uses.
max_results is 5 by default and at most 10. The sources come back in
nymbot.web_search.sources, each with a title, snippet and
url, and as url_citation annotations on the message.
Each search that runs costs $0.008, converted to sats, on top of the tokens, and it is part of the hold. The pages it read are input tokens too, so an answer from the web costs more than the same question asked cold, sometimes several times more.
Response, in part
"nymbot": {
"balance": "pro",
"charged_credits": 0.431,
"charged_sats": 43.1,
"balance_credits": 411.994,
"balance_sats": 41199.4,
"web_search": {
"sources": [
{ "title": "Lightning Network - Wikipedia", "snippet": "The Lightning Network is a payment protocol...", "url": "https://en.wikipedia.org/wiki/Lightning_Network" }
]
}
}
cURL
curl https://nymbot.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NYMBOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"plugins": [{"id": "web", "max_results": 5}],
"messages": [{"role": "user", "content": "What changed in the latest Bitcoin Core release?"}]
}'
Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])
reply = client.chat.completions.create(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "What changed in the latest Bitcoin Core release?"}],
extra_body={"plugins": [{"id": "web", "max_results": 5}]},
)
print(reply.choices[0].message.content)
for source in reply.model_extra["nymbot"]["web_search"]["sources"]:
print(source["url"])
JavaScript
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });
const reply = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5",
plugins: [{ id: "web", max_results: 5 }],
messages: [{ role: "user", content: "What changed in the latest Bitcoin Core release?" }],
});
console.log(reply.choices[0].message.content);
for (const source of reply.nymbot.web_search.sources) console.log(source.url);
Responses API
OpenAI's newer format, used by the OpenAI Agents SDK and by Codex. It runs on the same models, billing and features as Chat Completions.
POST https://nymbot.ai/api/v1/responses — needs an API key.
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | As for Chat Completions, suffixes included. |
input | string or array | Yes | A string, or a list of items: messages (roles user, assistant, system, developer) with input_text, input_image and output_text parts, and function_call and function_call_output items for tools. input_image takes an image_url link or data URL, not a file id. reasoning and web_search_call items are skipped; item_reference is refused. |
instructions | string | No | System instructions. |
max_output_tokens | integer | No | The most tokens to write. |
temperaturetop_p | number | No | Sampling controls, where the model takes them. |
toolstool_choiceparallel_tool_calls | array, string or object, boolean | No | Function tools, in the Responses shape ({"type": "function", "name": …, "parameters": …}). A web_search or web_search_preview tool turns on web search when it helps. Other built-in tools are refused. |
reasoning | object | No | {"effort": "minimal" | "low" | "medium" | "high"}. xhigh and max mean high; none turns it off. |
text.formatresponse_format | object | No | Structured output, as a JSON schema or json_object. |
metadata | object | No | Returned unchanged in the response. At most 16 string values, keys up to 64 characters and values up to 512. |
stream | boolean | No | Stream events as described below. |
store | boolean | No | Ignored. Nothing is stored, and the response always says "store": false. |
previous_response_idconversationbackground | string, object, boolean | No | Not supported: 400 unsupported_parameter. Responses are not stored, so send the whole conversation in input every time. |
Response
{
"id": "resp_8c1e4b0f9a2d4e61",
"object": "response",
"created_at": 1790726400,
"status": "completed",
"model": "anthropic/claude-sonnet-5",
"output": [
{
"type": "message",
"id": "msg_2b7f",
"role": "assistant",
"status": "completed",
"content": [{ "type": "output_text", "text": "A Lightning invoice is...", "annotations": [] }]
}
],
"output_text": "A Lightning invoice is...",
"usage": {
"input_tokens": 1240,
"input_tokens_details": { "cached_tokens": 0 },
"output_tokens": 380,
"output_tokens_details": { "reasoning_tokens": 0 },
"total_tokens": 1620
},
"incomplete_details": null,
"error": null,
"instructions": null,
"store": false,
"previous_response_id": null,
"metadata": {},
"nymbot": { "balance": "pro", "charged_credits": 0.162, "charged_sats": 16.2, "balance_credits": 412.425, "balance_sats": 41242.5 }
}
A tool request shows up in output as a function_call item with
call_id, name and arguments; send the result back as a
function_call_output item with the same call_id. Reasoning, when the
model returns it, is a reasoning item with reasoning_text content,
listed first. status is incomplete when the answer hit
max_output_tokens or the provider refused, with
incomplete_details.reason set to max_output_tokens or
content_filter. The response also echoes the request's settings (temperature,
tools, tool choice and so on) as OpenAI's does, and carries the nymbot cost
object.
Streamed, each event is an event: line and a data: line with a
sequence_number, in this order: response.created,
response.in_progress, response.output_item.added,
response.content_part.added, any number of response.output_text.delta,
response.output_text.done, response.content_part.done,
response.output_item.done, and finally response.completed with the
usage and cost. An answer cut short ends with response.incomplete instead, and a
failure after the stream started with response.failed. Reasoning streams as its
own item with response.reasoning_text.delta and .done. Tool calls
come after the message, each as an item with response.function_call_arguments.delta
and .done.
| Status | When |
|---|---|
400 | No model or input; previous_response_id, conversation, background or an item_reference (unsupported_parameter); an unsupported tool or content type. |
402, 403, 404, 429, 502, 503 | As for Chat Completions. |
cURL
curl https://nymbot.ai/api/v1/responses \
-H "Authorization: Bearer $NYMBOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"instructions": "Answer in one short paragraph.",
"input": "What is a Lightning invoice?"
}'
Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])
response = client.responses.create(
model="anthropic/claude-sonnet-5",
instructions="Answer in one short paragraph.",
input="What is a Lightning invoice?",
)
print(response.output_text)
JavaScript
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });
const response = await client.responses.create({
model: "anthropic/claude-sonnet-5",
instructions: "Answer in one short paragraph.",
input: "What is a Lightning invoice?",
});
console.log(response.output_text);
Anthropic Messages
Anthropic's format, for the Anthropic SDKs and Claude Code. It works with every model in the catalog, not only Claude: the request is translated, run through the same pipeline, and translated back.
POST https://nymbot.ai/api/v1/messages — needs an API key, as x-api-key or Authorization: Bearer.
The anthropic-version and anthropic-beta headers are accepted and
ignored. Anthropic model names are matched to the catalog, so Claude Code and the SDKs work
with the names they already use:
- A name the catalog knows, such as
claude-sonnet-5oranthropic/claude-opus-5, is used as it is. - Otherwise a date (
-20260514),-latest, a version tag such as-v1, a bracketed tag such as[1m]and ananthropic/oranthropic.prefix are taken off, and dots and dashes in the version are tried both ways (claude-haiku-4-5findsclaude-haiku-4.5). - If that still matches nothing, the family (Opus, Sonnet or Haiku) is used, as long as the catalog's version is the same or newer than the one asked for.
- A name that matches nothing returns
404not_found_error.
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | A catalog model id, or an Anthropic model name. |
max_tokens | integer | Yes | The most tokens to write. |
messages | array | Yes | user and assistant turns, with text, image (base64 or URL source), tool_use and tool_result blocks. thinking blocks from earlier turns are accepted and dropped. |
system | string or array | No | System prompt, as a string or text blocks. |
temperaturetop_p | number | No | Sampling controls, where the model takes them. top_k is accepted and dropped. |
stop_sequences | array of strings | No | Text that ends the answer. At most 4 strings, each at most 256 characters. |
toolstool_choice | array, object | No | Tools with name, description and input_schema. A web_search server tool turns on web search; Anthropic's other built-in tools (bash, text editor, computer use) are refused with unsupported_tool. tool_choice takes auto, any, tool or none, and disable_parallel_tool_use. |
thinking | object | No | {"type": "enabled", "budget_tokens": 8192}, {"type": "adaptive"} or {"type": "disabled"}. The budget picks an effort level: under 2,048 minimal, from 2,048 low, from 8,192 medium, from 16,384 high. Adaptive uses output_config.effort, or high. |
stream | boolean | No | Stream in Anthropic's event format. |
metadata | object | No | Accepted and ignored. |
Response
{
"id": "msg_01c7a2f93e5b4d08",
"type": "message",
"role": "assistant",
"model": "anthropic/claude-sonnet-5",
"content": [{ "type": "text", "text": "A Lightning invoice is..." }],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 1240,
"output_tokens": 380,
"cache_read_input_tokens": 0,
"cache_creation_input_tokens": 0
},
"nymbot": { "balance": "pro", "charged_credits": 0.162, "charged_sats": 16.2, "balance_credits": 412.425, "balance_sats": 41242.5 }
}
content can also hold tool_use blocks and a thinking
block, whose signature is empty. stop_reason is
end_turn, max_tokens, tool_use or refusal,
worked out the same way as finish_reason on
Chat Completions; stop_sequence is always
null, even when a stop sequence ended the answer. input_tokens counts
fresh input only; cached input is in the two cache fields. The cost is in the
nymbot object and the X-Nymbot-Cost-Sats header.
Streamed, the events are Anthropic's: message_start,
content_block_start, ping, content_block_delta
(text_delta, input_json_delta or thinking_delta),
content_block_stop, message_delta with the stop reason, the usage and
the nymbot cost object, and message_stop. A ping is also
sent every 15 seconds while the model is working. Tool calls arrive after the text, each as a
tool_use block with its whole input in one input_json_delta.
Errors on this endpoint use Anthropic's format:
{"type": "error", "error": {"type": "not_found_error", "message": "…"}}. A
short balance is 402 billing_error, an overloaded provider
503 overloaded_error. A failure after the stream has started is sent
as an error event.
cURL
curl https://nymbot.ai/api/v1/messages \
-H "x-api-key: $NYMBOT_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "What is a Lightning invoice?"}]
}'
Python
import os
import anthropic
client = anthropic.Anthropic(base_url="https://nymbot.ai/api", api_key=os.environ["NYMBOT_API_KEY"])
message = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": "What is a Lightning invoice?"}],
)
print(message.content[0].text)
JavaScript
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ baseURL: "https://nymbot.ai/api", apiKey: process.env.NYMBOT_API_KEY });
const message = await client.messages.create({
model: "claude-sonnet-5",
max_tokens: 1024,
messages: [{ role: "user", content: "What is a Lightning invoice?" }],
});
console.log(message.content[0].text);
Counting tokens
An estimate of how many input tokens a Messages request would use, so a client can check before it sends. It is free, but still needs a key.
POST https://nymbot.ai/api/v1/messages/count_tokens — needs an API key. Free.
The body is the same as for Messages, without max_tokens;
the model name has to resolve. The count is an estimate: the characters of the system prompt,
messages, tool calls and tool definitions divided by four, plus 1,600 for each picture. It is
not the provider's own tokenizer, so the real count can differ. A key that has reached its cap
can still use it.
Response
{ "input_tokens": 318 }
cURL
curl https://nymbot.ai/api/v1/messages/count_tokens \
-H "x-api-key: $NYMBOT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5",
"messages": [{"role": "user", "content": "What is a Lightning invoice?"}]
}'
Python
import os
import anthropic
client = anthropic.Anthropic(base_url="https://nymbot.ai/api", api_key=os.environ["NYMBOT_API_KEY"])
count = client.messages.count_tokens(
model="claude-sonnet-5",
messages=[{"role": "user", "content": "What is a Lightning invoice?"}],
)
print(count.input_tokens)
JavaScript
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ baseURL: "https://nymbot.ai/api", apiKey: process.env.NYMBOT_API_KEY });
const count = await client.messages.countTokens({
model: "claude-sonnet-5",
messages: [{ role: "user", content: "What is a Lightning invoice?" }],
});
console.log(count.input_tokens);
Listing models
Every model and generator the API accepts, with what it costs. The list is read from the same live catalog as the app's picker, so it is always what the server will run.
GET https://nymbot.ai/api/v1/models — no key needed. Cached for five minutes.
GET /api/v1/models/{id} returns one entry.
| Field | Type | Required | Description |
|---|---|---|---|
type | string (query) | No | chat (the default), image, video, audio, embedding or all. Several can be given, as image,video. |
Prices are what you pay, with the fee and margin already in, in dollars and in sats at the
current Bitcoin price. balance says which balance the model spends.
nymbot_key is the model's short name in the app. created is always
0, since the catalog does not record when a model was added. A model without published token
rates is priced per_request instead.
nymbot/auto is always first, priced as variable, with a
routes list giving each standard route's rates. It lists vision and reasoning, but
not tools.
GET /api/v1/models/{id} accepts the same names and aliases as a request, and returns
the entry for the model they resolve to.
Response
{
"object": "list",
"data": [
{
"id": "anthropic/claude-sonnet-5",
"object": "model",
"type": "chat",
"owned_by": "anthropic",
"name": "Claude Sonnet 5",
"created": 0,
"context_length": 1000000,
"max_output_tokens": 64000,
"architecture": { "input_modalities": ["text", "image"], "output_modalities": ["text"] },
"supported_parameters": ["max_tokens", "temperature", "tools", "tool_choice", "reasoning", "response_format", "stop"],
"capabilities": { "vision": true, "video": false, "tools": true, "reasoning": true, "web_search": true },
"balance": "pro",
"pricing": {
"type": "per_token",
"currency": "USD",
"input_per_1M_tokens": 4.725,
"output_per_1M_tokens": 23.625,
"cache_read_per_1M_tokens": 0.4725,
"sats_input_per_1M_tokens": 4038,
"sats_output_per_1M_tokens": 20192
},
"description": "...",
"nymbot_key": "claude-sonnet"
}
]
}
The other types:
- Image entries have
capabilities(accepts_image_url,requires_image_url,edit) and a priceper_generation. - Video entries have
max_duration_secondsandresolutions, and a priceper_secondfor each resolution. - Audio entries have
audio_typespeechortranscription, pricedper_1k_charsorper_minute. - Embedding entries have
dimensions,context_length,max_inputsand a price per million input tokens.
Prices marked "estimated": true are the app's estimate for a generator whose price
is not published. An unknown type returns 400.
cURL
curl "https://nymbot.ai/api/v1/models?type=chat"
Python
import requests
models = requests.get("https://nymbot.ai/api/v1/models", params={"type": "chat"}).json()["data"]
for m in models:
print(m["id"], m["balance"], m["pricing"].get("input_per_1M_tokens"))
JavaScript
const res = await fetch("https://nymbot.ai/api/v1/models?type=chat");
const { data } = await res.json();
for (const m of data) console.log(m.id, m.balance, m.pricing.input_per_1M_tokens);