Skip to the content
Back to Nymbot

Knowledge base Developers

Chat, Responses and Messages

Three ways to ask a model something, in the three formats clients already speak, all running on the same models and billed the same way.

Chat Completions

The OpenAI chat format, and the one almost every tool supports. Send the conversation so far and get the next message back.

POST https://nymbot.ai/api/v1/chat/completions — needs an API key.

Only model and messages are required. A sampling parameter the chosen model does not take is dropped without an error, so one request body works across models. The model list gives each model's supported_parameters. What is refused rather than dropped is anything the model cannot do at all: pictures for a model that cannot see, tools for a model that cannot call them.

FieldTypeRequiredDescription
modelstringYesnymbot/auto (or auto) for Nymbot's routing on the standard balance, or a catalog model id such as anthropic/claude-sonnet-5 on the Pro balance. The app's short names and aliases are accepted too. May end in a suffix.
messagesarrayYesThe conversation. Roles system, developer (treated as system), user, assistant and tool. Content is a string or a list of text and image_url parts; pictures only in user messages. input_audio and file parts are refused.
streambooleanNoSend the answer as it is written. See streaming.
stream_optionsobjectNo{"include_usage": true} adds a final chunk with token counts and the cost.
max_tokens
max_completion_tokens
integerNoThe most tokens to write. Lowered to the model's maximum if higher. Also sets how much is held from your balance, so a smaller number needs less credit to start.
temperature
top_p
numberNoSampling controls: temperature from 0 to 2, top_p from 0 to 1.
stopstring or arrayNoText that ends the answer: a string or at most 4 strings, each at most 256 characters. Not used by nymbot/auto.
seedintegerNoFor repeatable sampling, where the model supports it.
presence_penalty
frequency_penalty
numberNoRepetition controls, each from -2 to 2.
response_formatobjectNo{"type": "json_object"} or {"type": "json_schema", "json_schema": {…}}, where the model supports it. Not used by nymbot/auto.
tools
tool_choice
parallel_tool_calls
array, string or object, booleanNoFunction calling. See tool calls. A web_search tool turns on web search. At most 128 tools and 512 KB of definitions, nested at most 64 levels deep; more is a 400.
reasoning_effort
reasoning
string, objectNo"minimal", "low", "medium" or "high", or {"effort": "high"}. "none" or {"enabled": false} turns it off. See reasoning.
pluginsarrayNo[{"id": "web", "max_results": 5}] always searches the web first. Up to 10 results.
nintegerNoOnly 1. Anything else returns 400.
logit_bias
user
metadata
object, string, objectNoAccepted and not sent on. logit_bias maps at most 300 token ids to numbers from -100 to 100; user is at most 256 characters; metadata holds at most 16 string values, keys up to 64 characters and values up to 512.

The reply is an ordinary chat.completion, with the cost in usage.cost (in dollars) and in the nymbot object. model is the resolved model id, so a short name comes back as the full one. If the model reasoned before answering, its reasoning is in message.reasoning_content, separate from the answer.

Response

{
  "id": "chatcmpl-5f1c0a9e27d84b3c",
  "object": "chat.completion",
  "created": 1790726400,
  "model": "anthropic/claude-sonnet-5",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "A Lightning invoice is a one-time payment request..."
      },
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 1240,
    "completion_tokens": 380,
    "total_tokens": 1620,
    "prompt_tokens_details": { "cached_tokens": 0 },
    "completion_tokens_details": { "reasoning_tokens": 0 },
    "cost": 0.01895
  },
  "nymbot": {
    "balance": "pro",
    "charged_credits": 0.162,
    "charged_sats": 16.2,
    "balance_credits": 412.425,
    "balance_sats": 41242.5
  }
}

finish_reason is worked out from what came back, since providers report it differently: tool_calls when the model asked for tools, length when the answer used every token allowed (or the model spent them all reasoning and wrote no answer), content_filter when the provider refused, and stop otherwise. usage.completion_tokens_details.reasoning_tokens is always 0: hidden reasoning tokens are counted, and charged, in completion_tokens.

StatusWhen
400No model or messages (missing_required_parameter); an audio or file part, or pictures for a model that cannot see them (unsupported_content); more than 20 pictures (too_many_images); a sampling parameter of the wrong type or out of range (invalid_value); a picture link that is not public (invalid_image_url); tools on a model that cannot call them (unsupported_tool); n other than 1; a :thinking suffix on a model that cannot reason (model_not_found); or the provider refused the request (upstream_rejected).
402The balance the model spends cannot cover the hold.
403The worst case does not fit the key's cap (key_limit_reached). Lower max_tokens or raise the cap.
404No model by that name (model_not_found).
429The key's rate limit, or the provider's (upstream_rate_limited).
502, 503The provider failed (upstream_error) or is overloaded (upstream_overloaded, with Retry-After). Nothing is charged unless the provider billed for the attempt.

The status codes every endpoint shares are listed under errors.

cURL

curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "messages": [
      {"role": "system", "content": "Answer in one short paragraph."},
      {"role": "user", "content": "What is a Lightning invoice?"}
    ],
    "max_tokens": 400
  }'

Python

import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[
        {"role": "system", "content": "Answer in one short paragraph."},
        {"role": "user", "content": "What is a Lightning invoice?"},
    ],
    max_tokens=400,
)
print(reply.choices[0].message.content)
print(reply.usage.cost)

JavaScript

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5",
  messages: [
    { role: "system", content: "Answer in one short paragraph." },
    { role: "user", content: "What is a Lightning invoice?" },
  ],
  max_tokens: 400,
});
console.log(reply.choices[0].message.content);
console.log(reply.usage.cost);

Streaming

With "stream": true the answer arrives as server-sent events while the model writes it. Each event is a chat.completion.chunk on a data: line, and the stream ends with data: [DONE]. Reasoning arrives in delta.reasoning_content, the answer in delta.content.

Stream anything that may take more than about 100 seconds, such as a long answer, a large max_tokens or a reasoning model. A request that is not streamed sends nothing until the answer is complete, and the network between you and Nymbot may close an idle connection after about 100 seconds; the model still finishes and what it wrote is charged, but the answer is lost. A stream sends keep-alives, so it stays open for as long as the model writes.

Stream

data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{"content":"A Lightning invoice"},"finish_reason":null}]}

: keep-alive

data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[],"usage":{"prompt_tokens":1240,"completion_tokens":380,"total_tokens":1620,"cost":0.01895},"nymbot":{"balance":"pro","charged_credits":0.162,"charged_sats":16.2,"balance_credits":412.425,"balance_sats":41242.5}}

data: [DONE]
  • Lines starting with : are keep-alive comments, sent every 15 seconds while the model is thinking. SSE clients skip them.
  • Ask for "stream_options": {"include_usage": true} to get the last chunk above, with an empty choices, the token counts, the cost and the nymbot object.
  • The charge settles after the stream ends. If you close the connection early, the model is not stopped: Nymbot reads the rest of the provider's stream, for up to 25 seconds, to get its token count, and you pay what the provider reports. Without that count the charge is estimated from your input and what was written, plus the whole output allowance for a model that reasons.
  • An error before the first chunk, such as 401 or 402, comes back as ordinary JSON with its status code, not as a stream. An error after the stream has started arrives as a last data: {"error": {…}} event, and the stream ends without [DONE].
  • Requests with tools, and models on OpenAI's Responses transport, do not stream from the provider. They still answer with a valid stream, sent once the answer is complete.

cURL

curl -N https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nymbot/auto",
    "stream": true,
    "stream_options": {"include_usage": true},
    "messages": [{"role": "user", "content": "Write a haiku about sats."}]
  }'

Python

import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

stream = client.chat.completions.create(
    model="nymbot/auto",
    stream=True,
    stream_options={"include_usage": True},
    messages=[{"role": "user", "content": "Write a haiku about sats."}],
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="", flush=True)
    elif chunk.usage:
        print("\ncost:", chunk.usage.cost)

JavaScript

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const stream = await client.chat.completions.create({
  model: "nymbot/auto",
  stream: true,
  stream_options: { include_usage: true },
  messages: [{ role: "user", content: "Write a haiku about sats." }],
});
for await (const chunk of stream) {
  if (chunk.choices.length) process.stdout.write(chunk.choices[0].delta.content || "");
  else if (chunk.usage) console.log("\ncost:", chunk.usage.cost);
}

Tool calls

Describe functions in tools and the model can ask for one to be called instead of answering. You run the function, add its result as a tool message with the same tool_call_id, and send the conversation again. Nymbot never runs your functions; it passes the model's request back to you.

tool_choice takes "auto", "none", "required" or {"type": "function", "function": {"name": "…"}}. The model list marks which models can call tools (capabilities.tools). Tools are refused with 400 unsupported_tool on nymbot/auto and on the few catalog models that run on OpenAI's Responses transport, which Nymbot cannot pass tools to.

With "stream": true, a request with tools runs in one piece and is then sent as a normal chunk sequence: the role, one chunk carrying every tool call with its index, id, name and arguments, and the finish chunk. Clients that read streamed tool calls handle it as usual.

Response, in part

"choices": [
  {
    "index": 0,
    "message": {
      "role": "assistant",
      "content": null,
      "tool_calls": [
        {
          "id": "call_7d2e",
          "type": "function",
          "function": { "name": "get_invoice_status", "arguments": "{\"invoice_id\":\"a41f\"}" }
        }
      ]
    },
    "finish_reason": "tool_calls"
  }
]

cURL

curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "messages": [{"role": "user", "content": "Has invoice a41f been paid?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_invoice_status",
        "description": "Look up whether an invoice is paid.",
        "parameters": {
          "type": "object",
          "properties": {"invoice_id": {"type": "string"}},
          "required": ["invoice_id"]
        }
      }
    }]
  }'

Python

import json, os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

tools = [{
    "type": "function",
    "function": {
        "name": "get_invoice_status",
        "description": "Look up whether an invoice is paid.",
        "parameters": {
            "type": "object",
            "properties": {"invoice_id": {"type": "string"}},
            "required": ["invoice_id"],
        },
    },
}]
messages = [{"role": "user", "content": "Has invoice a41f been paid?"}]

reply = client.chat.completions.create(model="anthropic/claude-sonnet-5", messages=messages, tools=tools)
call = reply.choices[0].message.tool_calls[0]
args = json.loads(call.function.arguments)

messages.append(reply.choices[0].message)
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps({"paid": True})})

final = client.chat.completions.create(model="anthropic/claude-sonnet-5", messages=messages, tools=tools)
print(final.choices[0].message.content)

JavaScript

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const tools = [{
  type: "function",
  function: {
    name: "get_invoice_status",
    description: "Look up whether an invoice is paid.",
    parameters: {
      type: "object",
      properties: { invoice_id: { type: "string" } },
      required: ["invoice_id"],
    },
  },
}];
const messages = [{ role: "user", content: "Has invoice a41f been paid?" }];

const reply = await client.chat.completions.create({ model: "anthropic/claude-sonnet-5", messages, tools });
const call = reply.choices[0].message.tool_calls[0];
const args = JSON.parse(call.function.arguments);

messages.push(reply.choices[0].message);
messages.push({ role: "tool", tool_call_id: call.id, content: JSON.stringify({ paid: true }) });

const final = await client.chat.completions.create({ model: "anthropic/claude-sonnet-5", messages, tools });
console.log(final.choices[0].message.content);

Pictures in a request

Models with capabilities.vision can read pictures. Add an image_url part to a user message, with either a public https:// link or a data:image/…;base64, URL. Up to 20 pictures per request. SVG pictures are refused, and so is a link over 4,096 characters or one with a user name or password in it (400 invalid_image_url). An optional detail is auto, low or high.

With nymbot/auto, a request with a picture in it is routed to a standard model that can see. A catalog model without capabilities.vision refuses pictures with 400 unsupported_content. Pictures can only be in user messages. A link must point to a public host; Nymbot passes the picture to the model's provider and does not keep it.

A picture is charged as the input tokens the provider counts for it, like the rest of the request.

cURL

curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "What is in this picture?"},
        {"type": "image_url", "image_url": {"url": "https://example.com/receipt.jpg"}}
      ]
    }]
  }'

Python

import base64, os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

with open("receipt.jpg", "rb") as f:
    data_url = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode()

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this picture?"},
            {"type": "image_url", "image_url": {"url": data_url}},
        ],
    }],
)
print(reply.choices[0].message.content)

JavaScript

import { readFile } from "node:fs/promises";
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const dataUrl = "data:image/jpeg;base64," + (await readFile("receipt.jpg")).toString("base64");

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5",
  messages: [{
    role: "user",
    content: [
      { type: "text", text: "What is in this picture?" },
      { type: "image_url", image_url: { url: dataUrl } },
    ],
  }],
});
console.log(reply.choices[0].message.content);

Reasoning

Models with capabilities.reasoning can think before they answer. Ask for more or less of it with reasoning_effort ("minimal", "low", "medium" or "high") or "reasoning": {"effort": "high"}, or add :thinking to the model name, which means high effort. On a model without reasoning the setting is ignored. With nymbot/auto, :thinking sends the request to the standard reasoning route.

On Anthropic models the effort becomes a thinking budget of about 1,000, 2,000, 8,000 or 16,000 tokens, never more than max_tokens allows. Thinking is left off when tool_choice forces a tool, and when the request continues a tool loop (its last message is a tool result), because Anthropic needs the earlier signed thinking to resume it. The same applies on the Responses and Messages endpoints.

The reasoning comes back in message.reasoning_content, or delta.reasoning_content when streaming, never mixed into the answer. Reasoning is output and is charged as such, inside completion_tokens. Some providers do not return the reasoning text at all, and it is still charged.

cURL

curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "reasoning_effort": "high",
    "messages": [{"role": "user", "content": "Is 2^61 - 1 prime? Show why."}]
  }'

Python

import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    reasoning_effort="high",
    messages=[{"role": "user", "content": "Is 2^61 - 1 prime? Show why."}],
)
message = reply.choices[0].message
print(getattr(message, "reasoning_content", None))
print(message.content)

JavaScript

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5",
  reasoning_effort: "high",
  messages: [{ role: "user", content: "Is 2^61 - 1 prime? Show why." }],
});
console.log(reply.choices[0].message.reasoning_content);
console.log(reply.choices[0].message.content);

Model suffixes

A suffix on the model name changes how the request is handled, without another field. They work on full ids and short names alike, as in anthropic/claude-sonnet-5:online.

SuffixEffect
:onlineSearches the web first, like plugins: [{"id": "web"}]. See web search.
:thinkingHigh reasoning effort; on nymbot/auto, the reasoning route. On a model that cannot reason, 400 model_not_found with “no endpoints found”.
:nitro, :floor, :exacto, :extendedAccepted and ignored. Each catalog model has one route, so there is no faster, cheaper or longer one to pick; the suffixes are allowed so model names copied from other services still work.

Any other suffix is ignored. The whole name is tried first, so a model whose id really contains a colon still works; failing that, suffixes are taken off the end one at a time until a model matches. A model name may be at most 200 characters long with at most 4 suffixes; a longer one is refused with 400 invalid_value.

cURL

curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5:thinking",
    "messages": [{"role": "user", "content": "Plan a three-day trip to Lisbon."}]
  }'

Python

import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5:thinking",
    messages=[{"role": "user", "content": "Plan a three-day trip to Lisbon."}],
)
print(reply.choices[0].message.content)

JavaScript

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5:thinking",
  messages: [{ role: "user", content: "Plan a three-day trip to Lisbon." }],
});
console.log(reply.choices[0].message.content);

Any chat model can answer from the live web. Nymbot searches for the last user message (its first 2,000 characters), reads the best pages and gives them to the model with the question, marked as outside content the model should not take instructions from. There are two modes:

  • Always search: "plugins": [{"id": "web", "max_results": 5}], or a :online suffix on the model.
  • Search when it helps: "tools": [{"type": "web_search", "parameters": {"max_results": 5}}], also accepted as web_search_preview or openrouter:web_search. Nymbot searches only when the question looks like it needs current information, the same test the app uses.

max_results is 5 by default and at most 10. The sources come back in nymbot.web_search.sources, each with a title, snippet and url, and as url_citation annotations on the message.

Each search that runs costs $0.008, converted to sats, on top of the tokens, and it is part of the hold. The pages it read are input tokens too, so an answer from the web costs more than the same question asked cold, sometimes several times more.

Response, in part

"nymbot": {
  "balance": "pro",
  "charged_credits": 0.431,
  "charged_sats": 43.1,
  "balance_credits": 411.994,
  "balance_sats": 41199.4,
  "web_search": {
    "sources": [
      { "title": "Lightning Network - Wikipedia", "snippet": "The Lightning Network is a payment protocol...", "url": "https://en.wikipedia.org/wiki/Lightning_Network" }
    ]
  }
}

cURL

curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "plugins": [{"id": "web", "max_results": 5}],
    "messages": [{"role": "user", "content": "What changed in the latest Bitcoin Core release?"}]
  }'

Python

import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "What changed in the latest Bitcoin Core release?"}],
    extra_body={"plugins": [{"id": "web", "max_results": 5}]},
)
print(reply.choices[0].message.content)
for source in reply.model_extra["nymbot"]["web_search"]["sources"]:
    print(source["url"])

JavaScript

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5",
  plugins: [{ id: "web", max_results: 5 }],
  messages: [{ role: "user", content: "What changed in the latest Bitcoin Core release?" }],
});
console.log(reply.choices[0].message.content);
for (const source of reply.nymbot.web_search.sources) console.log(source.url);

Responses API

OpenAI's newer format, used by the OpenAI Agents SDK and by Codex. It runs on the same models, billing and features as Chat Completions.

POST https://nymbot.ai/api/v1/responses — needs an API key.

FieldTypeRequiredDescription
modelstringYesAs for Chat Completions, suffixes included.
inputstring or arrayYesA string, or a list of items: messages (roles user, assistant, system, developer) with input_text, input_image and output_text parts, and function_call and function_call_output items for tools. input_image takes an image_url link or data URL, not a file id. reasoning and web_search_call items are skipped; item_reference is refused.
instructionsstringNoSystem instructions.
max_output_tokensintegerNoThe most tokens to write.
temperature
top_p
numberNoSampling controls, where the model takes them.
tools
tool_choice
parallel_tool_calls
array, string or object, booleanNoFunction tools, in the Responses shape ({"type": "function", "name": …, "parameters": …}). A web_search or web_search_preview tool turns on web search when it helps. Other built-in tools are refused.
reasoningobjectNo{"effort": "minimal" | "low" | "medium" | "high"}. xhigh and max mean high; none turns it off.
text.format
response_format
objectNoStructured output, as a JSON schema or json_object.
metadataobjectNoReturned unchanged in the response. At most 16 string values, keys up to 64 characters and values up to 512.
streambooleanNoStream events as described below.
storebooleanNoIgnored. Nothing is stored, and the response always says "store": false.
previous_response_id
conversation
background
string, object, booleanNoNot supported: 400 unsupported_parameter. Responses are not stored, so send the whole conversation in input every time.

Response

{
  "id": "resp_8c1e4b0f9a2d4e61",
  "object": "response",
  "created_at": 1790726400,
  "status": "completed",
  "model": "anthropic/claude-sonnet-5",
  "output": [
    {
      "type": "message",
      "id": "msg_2b7f",
      "role": "assistant",
      "status": "completed",
      "content": [{ "type": "output_text", "text": "A Lightning invoice is...", "annotations": [] }]
    }
  ],
  "output_text": "A Lightning invoice is...",
  "usage": {
    "input_tokens": 1240,
    "input_tokens_details": { "cached_tokens": 0 },
    "output_tokens": 380,
    "output_tokens_details": { "reasoning_tokens": 0 },
    "total_tokens": 1620
  },
  "incomplete_details": null,
  "error": null,
  "instructions": null,
  "store": false,
  "previous_response_id": null,
  "metadata": {},
  "nymbot": { "balance": "pro", "charged_credits": 0.162, "charged_sats": 16.2, "balance_credits": 412.425, "balance_sats": 41242.5 }
}

A tool request shows up in output as a function_call item with call_id, name and arguments; send the result back as a function_call_output item with the same call_id. Reasoning, when the model returns it, is a reasoning item with reasoning_text content, listed first. status is incomplete when the answer hit max_output_tokens or the provider refused, with incomplete_details.reason set to max_output_tokens or content_filter. The response also echoes the request's settings (temperature, tools, tool choice and so on) as OpenAI's does, and carries the nymbot cost object.

Streamed, each event is an event: line and a data: line with a sequence_number, in this order: response.created, response.in_progress, response.output_item.added, response.content_part.added, any number of response.output_text.delta, response.output_text.done, response.content_part.done, response.output_item.done, and finally response.completed with the usage and cost. An answer cut short ends with response.incomplete instead, and a failure after the stream started with response.failed. Reasoning streams as its own item with response.reasoning_text.delta and .done. Tool calls come after the message, each as an item with response.function_call_arguments.delta and .done.

StatusWhen
400No model or input; previous_response_id, conversation, background or an item_reference (unsupported_parameter); an unsupported tool or content type.
402, 403, 404, 429, 502, 503As for Chat Completions.

cURL

curl https://nymbot.ai/api/v1/responses \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "instructions": "Answer in one short paragraph.",
    "input": "What is a Lightning invoice?"
  }'

Python

import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

response = client.responses.create(
    model="anthropic/claude-sonnet-5",
    instructions="Answer in one short paragraph.",
    input="What is a Lightning invoice?",
)
print(response.output_text)

JavaScript

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const response = await client.responses.create({
  model: "anthropic/claude-sonnet-5",
  instructions: "Answer in one short paragraph.",
  input: "What is a Lightning invoice?",
});
console.log(response.output_text);

Anthropic Messages

Anthropic's format, for the Anthropic SDKs and Claude Code. It works with every model in the catalog, not only Claude: the request is translated, run through the same pipeline, and translated back.

POST https://nymbot.ai/api/v1/messages — needs an API key, as x-api-key or Authorization: Bearer.

The anthropic-version and anthropic-beta headers are accepted and ignored. Anthropic model names are matched to the catalog, so Claude Code and the SDKs work with the names they already use:

  • A name the catalog knows, such as claude-sonnet-5 or anthropic/claude-opus-5, is used as it is.
  • Otherwise a date (-20260514), -latest, a version tag such as -v1, a bracketed tag such as [1m] and an anthropic/ or anthropic. prefix are taken off, and dots and dashes in the version are tried both ways (claude-haiku-4-5 finds claude-haiku-4.5).
  • If that still matches nothing, the family (Opus, Sonnet or Haiku) is used, as long as the catalog's version is the same or newer than the one asked for.
  • A name that matches nothing returns 404 not_found_error.
FieldTypeRequiredDescription
modelstringYesA catalog model id, or an Anthropic model name.
max_tokensintegerYesThe most tokens to write.
messagesarrayYesuser and assistant turns, with text, image (base64 or URL source), tool_use and tool_result blocks. thinking blocks from earlier turns are accepted and dropped.
systemstring or arrayNoSystem prompt, as a string or text blocks.
temperature
top_p
numberNoSampling controls, where the model takes them. top_k is accepted and dropped.
stop_sequencesarray of stringsNoText that ends the answer. At most 4 strings, each at most 256 characters.
tools
tool_choice
array, objectNoTools with name, description and input_schema. A web_search server tool turns on web search; Anthropic's other built-in tools (bash, text editor, computer use) are refused with unsupported_tool. tool_choice takes auto, any, tool or none, and disable_parallel_tool_use.
thinkingobjectNo{"type": "enabled", "budget_tokens": 8192}, {"type": "adaptive"} or {"type": "disabled"}. The budget picks an effort level: under 2,048 minimal, from 2,048 low, from 8,192 medium, from 16,384 high. Adaptive uses output_config.effort, or high.
streambooleanNoStream in Anthropic's event format.
metadataobjectNoAccepted and ignored.

Response

{
  "id": "msg_01c7a2f93e5b4d08",
  "type": "message",
  "role": "assistant",
  "model": "anthropic/claude-sonnet-5",
  "content": [{ "type": "text", "text": "A Lightning invoice is..." }],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 1240,
    "output_tokens": 380,
    "cache_read_input_tokens": 0,
    "cache_creation_input_tokens": 0
  },
  "nymbot": { "balance": "pro", "charged_credits": 0.162, "charged_sats": 16.2, "balance_credits": 412.425, "balance_sats": 41242.5 }
}

content can also hold tool_use blocks and a thinking block, whose signature is empty. stop_reason is end_turn, max_tokens, tool_use or refusal, worked out the same way as finish_reason on Chat Completions; stop_sequence is always null, even when a stop sequence ended the answer. input_tokens counts fresh input only; cached input is in the two cache fields. The cost is in the nymbot object and the X-Nymbot-Cost-Sats header.

Streamed, the events are Anthropic's: message_start, content_block_start, ping, content_block_delta (text_delta, input_json_delta or thinking_delta), content_block_stop, message_delta with the stop reason, the usage and the nymbot cost object, and message_stop. A ping is also sent every 15 seconds while the model is working. Tool calls arrive after the text, each as a tool_use block with its whole input in one input_json_delta.

Errors on this endpoint use Anthropic's format: {"type": "error", "error": {"type": "not_found_error", "message": "…"}}. A short balance is 402 billing_error, an overloaded provider 503 overloaded_error. A failure after the stream has started is sent as an error event.

cURL

curl https://nymbot.ai/api/v1/messages \
  -H "x-api-key: $NYMBOT_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "What is a Lightning invoice?"}]
  }'

Python

import os
import anthropic

client = anthropic.Anthropic(base_url="https://nymbot.ai/api", api_key=os.environ["NYMBOT_API_KEY"])

message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "What is a Lightning invoice?"}],
)
print(message.content[0].text)

JavaScript

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({ baseURL: "https://nymbot.ai/api", apiKey: process.env.NYMBOT_API_KEY });

const message = await client.messages.create({
  model: "claude-sonnet-5",
  max_tokens: 1024,
  messages: [{ role: "user", content: "What is a Lightning invoice?" }],
});
console.log(message.content[0].text);

Counting tokens

An estimate of how many input tokens a Messages request would use, so a client can check before it sends. It is free, but still needs a key.

POST https://nymbot.ai/api/v1/messages/count_tokens — needs an API key. Free.

The body is the same as for Messages, without max_tokens; the model name has to resolve. The count is an estimate: the characters of the system prompt, messages, tool calls and tool definitions divided by four, plus 1,600 for each picture. It is not the provider's own tokenizer, so the real count can differ. A key that has reached its cap can still use it.

Response

{ "input_tokens": 318 }

cURL

curl https://nymbot.ai/api/v1/messages/count_tokens \
  -H "x-api-key: $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "messages": [{"role": "user", "content": "What is a Lightning invoice?"}]
  }'

Python

import os
import anthropic

client = anthropic.Anthropic(base_url="https://nymbot.ai/api", api_key=os.environ["NYMBOT_API_KEY"])

count = client.messages.count_tokens(
    model="claude-sonnet-5",
    messages=[{"role": "user", "content": "What is a Lightning invoice?"}],
)
print(count.input_tokens)

JavaScript

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({ baseURL: "https://nymbot.ai/api", apiKey: process.env.NYMBOT_API_KEY });

const count = await client.messages.countTokens({
  model: "claude-sonnet-5",
  messages: [{ role: "user", content: "What is a Lightning invoice?" }],
});
console.log(count.input_tokens);

Listing models

Every model and generator the API accepts, with what it costs. The list is read from the same live catalog as the app's picker, so it is always what the server will run.

GET https://nymbot.ai/api/v1/models — no key needed. Cached for five minutes.

GET /api/v1/models/{id} returns one entry.

FieldTypeRequiredDescription
typestring (query)Nochat (the default), image, video, audio, embedding or all. Several can be given, as image,video.

Prices are what you pay, with the fee and margin already in, in dollars and in sats at the current Bitcoin price. balance says which balance the model spends. nymbot_key is the model's short name in the app. created is always 0, since the catalog does not record when a model was added. A model without published token rates is priced per_request instead.

nymbot/auto is always first, priced as variable, with a routes list giving each standard route's rates. It lists vision and reasoning, but not tools.

GET /api/v1/models/{id} accepts the same names and aliases as a request, and returns the entry for the model they resolve to.

Response

{
  "object": "list",
  "data": [
    {
      "id": "anthropic/claude-sonnet-5",
      "object": "model",
      "type": "chat",
      "owned_by": "anthropic",
      "name": "Claude Sonnet 5",
      "created": 0,
      "context_length": 1000000,
      "max_output_tokens": 64000,
      "architecture": { "input_modalities": ["text", "image"], "output_modalities": ["text"] },
      "supported_parameters": ["max_tokens", "temperature", "tools", "tool_choice", "reasoning", "response_format", "stop"],
      "capabilities": { "vision": true, "video": false, "tools": true, "reasoning": true, "web_search": true },
      "balance": "pro",
      "pricing": {
        "type": "per_token",
        "currency": "USD",
        "input_per_1M_tokens": 4.725,
        "output_per_1M_tokens": 23.625,
        "cache_read_per_1M_tokens": 0.4725,
        "sats_input_per_1M_tokens": 4038,
        "sats_output_per_1M_tokens": 20192
      },
      "description": "...",
      "nymbot_key": "claude-sonnet"
    }
  ]
}

The other types:

  • Image entries have capabilities (accepts_image_url, requires_image_url, edit) and a price per_generation.
  • Video entries have max_duration_seconds and resolutions, and a price per_second for each resolution.
  • Audio entries have audio_type speech or transcription, priced per_1k_chars or per_minute.
  • Embedding entries have dimensions, context_length, max_inputs and a price per million input tokens.

Prices marked "estimated": true are the app's estimate for a generator whose price is not published. An unknown type returns 400.

cURL

curl "https://nymbot.ai/api/v1/models?type=chat"

Python

import requests

models = requests.get("https://nymbot.ai/api/v1/models", params={"type": "chat"}).json()["data"]
for m in models:
    print(m["id"], m["balance"], m["pricing"].get("input_per_1M_tokens"))

JavaScript

const res = await fetch("https://nymbot.ai/api/v1/models?type=chat");
const { data } = await res.json();
for (const m of data) console.log(m.id, m.balance, m.pricing.input_per_1M_tokens);