# Chat, Responses and Messages

Three ways to ask a model something, in the three formats clients already speak, all running on the same models and billed the same way.

## Chat Completions

The OpenAI chat format, and the one almost every tool supports. Send the conversation so far and get the next message back.

`POST https://nymbot.ai/api/v1/chat/completions` — needs an API key.

Only `model` and `messages` are required. A sampling parameter the chosen model does not take is dropped without an error, so one request body works across models. The model list gives each model's `supported_parameters`. What is refused rather than dropped is anything the model cannot do at all: pictures for a model that cannot see, tools for a model that cannot call them.

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `model` | string | Yes | `nymbot/auto` (or `auto`) for Nymbot's routing on the standard balance, or a catalog model id such as `anthropic/claude-sonnet-5` on the Pro balance. The app's short names and aliases are accepted too. May end in a [suffix](#suffixes). |
| `messages` | array | Yes | The conversation. Roles `system`, `developer` (treated as system), `user`, `assistant` and `tool`. Content is a string or a list of `text` and `image_url` parts; pictures only in user messages. `input_audio` and file parts are refused. |
| `stream` | boolean | No | Send the answer as it is written. See [streaming](#streaming). |
| `stream_options` | object | No | `{"include_usage": true}` adds a final chunk with token counts and the cost. |
| `max_tokens` `max_completion_tokens` | integer | No | The most tokens to write. Lowered to the model's maximum if higher. Also sets how much is held from your balance, so a smaller number needs less credit to start. |
| `temperature` `top_p` | number | No | Sampling controls: `temperature` from 0 to 2, `top_p` from 0 to 1. |
| `stop` | string or array | No | Text that ends the answer: a string or at most 4 strings, each at most 256 characters. Not used by `nymbot/auto`. |
| `seed` | integer | No | For repeatable sampling, where the model supports it. |
| `presence_penalty` `frequency_penalty` | number | No | Repetition controls, each from -2 to 2. |
| `response_format` | object | No | `{"type": "json_object"}` or `{"type": "json_schema", "json_schema": {…}}`, where the model supports it. Not used by `nymbot/auto`. |
| `tools` `tool_choice` `parallel_tool_calls` | array, string or object, boolean | No | Function calling. See [tool calls](#tools). A `web_search` tool turns on [web search](#web-search). At most 128 tools and 512 KB of definitions, nested at most 64 levels deep; more is a `400`. |
| `reasoning_effort` `reasoning` | string, object | No | `"minimal"`, `"low"`, `"medium"` or `"high"`, or `{"effort": "high"}`. `"none"` or `{"enabled": false}` turns it off. See [reasoning](#reasoning). |
| `plugins` | array | No | `[{"id": "web", "max_results": 5}]` always searches the web first. Up to 10 results. |
| `n` | integer | No | Only 1. Anything else returns `400`. |
| `logit_bias` `user` `metadata` | object, string, object | No | Accepted and not sent on. `logit_bias` maps at most 300 token ids to numbers from -100 to 100; `user` is at most 256 characters; `metadata` holds at most 16 string values, keys up to 64 characters and values up to 512. |

The reply is an ordinary `chat.completion`, with the cost in `usage.cost` (in dollars) and in the `nymbot` object. `model` is the resolved model id, so a short name comes back as the full one. If the model reasoned before answering, its reasoning is in `message.reasoning_content`, separate from the answer.

Response

```
{
  "id": "chatcmpl-5f1c0a9e27d84b3c",
  "object": "chat.completion",
  "created": 1790726400,
  "model": "anthropic/claude-sonnet-5",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "A Lightning invoice is a one-time payment request..."
      },
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 1240,
    "completion_tokens": 380,
    "total_tokens": 1620,
    "prompt_tokens_details": { "cached_tokens": 0 },
    "completion_tokens_details": { "reasoning_tokens": 0 },
    "cost": 0.01895
  },
  "nymbot": {
    "balance": "pro",
    "charged_credits": 0.162,
    "charged_sats": 16.2,
    "balance_credits": 412.425,
    "balance_sats": 41242.5
  }
}
```

`finish_reason` is worked out from what came back, since providers report it differently: `tool_calls` when the model asked for tools, `length` when the answer used every token allowed (or the model spent them all reasoning and wrote no answer), `content_filter` when the provider refused, and `stop` otherwise. `usage.completion_tokens_details.reasoning_tokens` is always 0: hidden reasoning tokens are counted, and charged, in `completion_tokens`.

| Status | When |
| --- | --- |
| `400` | No `model` or `messages` (`missing_required_parameter`); an audio or file part, or pictures for a model that cannot see them (`unsupported_content`); more than 20 pictures (`too_many_images`); a sampling parameter of the wrong type or out of range (`invalid_value`); a picture link that is not public (`invalid_image_url`); tools on a model that cannot call them (`unsupported_tool`); `n` other than 1; a `:thinking` suffix on a model that cannot reason (`model_not_found`); or the provider refused the request (`upstream_rejected`). |
| `402` | The balance the model spends cannot cover the hold. |
| `403` | The worst case does not fit the key's cap (`key_limit_reached`). Lower `max_tokens` or raise the cap. |
| `404` | No model by that name (`model_not_found`). |
| `429` | The key's rate limit, or the provider's (`upstream_rate_limited`). |
| `502`, `503` | The provider failed (`upstream_error`) or is overloaded (`upstream_overloaded`, with `Retry-After`). Nothing is charged unless the provider billed for the attempt. |

The status codes every endpoint shares are listed under [errors](https://nymbot.ai/docs/api/#errors).

cURL

```
curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "messages": [
      {"role": "system", "content": "Answer in one short paragraph."},
      {"role": "user", "content": "What is a Lightning invoice?"}
    ],
    "max_tokens": 400
  }'
```

Python

```
import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[
        {"role": "system", "content": "Answer in one short paragraph."},
        {"role": "user", "content": "What is a Lightning invoice?"},
    ],
    max_tokens=400,
)
print(reply.choices[0].message.content)
print(reply.usage.cost)
```

JavaScript

```
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5",
  messages: [
    { role: "system", content: "Answer in one short paragraph." },
    { role: "user", content: "What is a Lightning invoice?" },
  ],
  max_tokens: 400,
});
console.log(reply.choices[0].message.content);
console.log(reply.usage.cost);
```

## Streaming

With `"stream": true` the answer arrives as server-sent events while the model writes it. Each event is a `chat.completion.chunk` on a `data:` line, and the stream ends with `data: [DONE]`. Reasoning arrives in `delta.reasoning_content`, the answer in `delta.content`.

Stream anything that may take more than about 100 seconds, such as a long answer, a large `max_tokens` or a reasoning model. A request that is not streamed sends nothing until the answer is complete, and the network between you and Nymbot may close an idle connection after about 100 seconds; the model still finishes and what it wrote is charged, but the answer is lost. A stream sends keep-alives, so it stays open for as long as the model writes.

Stream

```
data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{"content":"A Lightning invoice"},"finish_reason":null}]}

: keep-alive

data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"id":"chatcmpl-5f1c0a9e27d84b3c","object":"chat.completion.chunk","created":1790726400,"model":"anthropic/claude-sonnet-5","choices":[],"usage":{"prompt_tokens":1240,"completion_tokens":380,"total_tokens":1620,"cost":0.01895},"nymbot":{"balance":"pro","charged_credits":0.162,"charged_sats":16.2,"balance_credits":412.425,"balance_sats":41242.5}}

data: [DONE]
```

- Lines starting with `:` are keep-alive comments, sent every 15 seconds while the model is thinking. SSE clients skip them.
- Ask for `"stream_options": {"include_usage": true}` to get the last chunk above, with an empty `choices`, the token counts, the cost and the `nymbot` object.
- The charge settles after the stream ends. If you close the connection early, the model is not stopped: Nymbot reads the rest of the provider's stream, for up to 25 seconds, to get its token count, and you pay what the provider reports. Without that count the charge is estimated from your input and what was written, plus the whole output allowance for a model that reasons.
- An error before the first chunk, such as `401` or `402`, comes back as ordinary JSON with its status code, not as a stream. An error after the stream has started arrives as a last `data: {"error": {…}}` event, and the stream ends without `[DONE]`.
- Requests with tools, and models on OpenAI's Responses transport, do not stream from the provider. They still answer with a valid stream, sent once the answer is complete.

cURL

```
curl -N https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nymbot/auto",
    "stream": true,
    "stream_options": {"include_usage": true},
    "messages": [{"role": "user", "content": "Write a haiku about sats."}]
  }'
```

Python

```
import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

stream = client.chat.completions.create(
    model="nymbot/auto",
    stream=True,
    stream_options={"include_usage": True},
    messages=[{"role": "user", "content": "Write a haiku about sats."}],
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="", flush=True)
    elif chunk.usage:
        print("\ncost:", chunk.usage.cost)
```

JavaScript

```
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const stream = await client.chat.completions.create({
  model: "nymbot/auto",
  stream: true,
  stream_options: { include_usage: true },
  messages: [{ role: "user", content: "Write a haiku about sats." }],
});
for await (const chunk of stream) {
  if (chunk.choices.length) process.stdout.write(chunk.choices[0].delta.content || "");
  else if (chunk.usage) console.log("\ncost:", chunk.usage.cost);
}
```

## Tool calls

Describe functions in `tools` and the model can ask for one to be called instead of answering. You run the function, add its result as a `tool` message with the same `tool_call_id`, and send the conversation again. Nymbot never runs your functions; it passes the model's request back to you.

`tool_choice` takes `"auto"`, `"none"`, `"required"` or `{"type": "function", "function": {"name": "…"}}`. The model list marks which models can call tools (`capabilities.tools`). Tools are refused with `400` `unsupported_tool` on `nymbot/auto` and on the few catalog models that run on OpenAI's Responses transport, which Nymbot cannot pass tools to.

With `"stream": true`, a request with tools runs in one piece and is then sent as a normal chunk sequence: the role, one chunk carrying every tool call with its index, id, name and arguments, and the finish chunk. Clients that read streamed tool calls handle it as usual.

Response, in part

```
"choices": [
  {
    "index": 0,
    "message": {
      "role": "assistant",
      "content": null,
      "tool_calls": [
        {
          "id": "call_7d2e",
          "type": "function",
          "function": { "name": "get_invoice_status", "arguments": "{\"invoice_id\":\"a41f\"}" }
        }
      ]
    },
    "finish_reason": "tool_calls"
  }
]
```

cURL

```
curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "messages": [{"role": "user", "content": "Has invoice a41f been paid?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_invoice_status",
        "description": "Look up whether an invoice is paid.",
        "parameters": {
          "type": "object",
          "properties": {"invoice_id": {"type": "string"}},
          "required": ["invoice_id"]
        }
      }
    }]
  }'
```

Python

```
import json, os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

tools = [{
    "type": "function",
    "function": {
        "name": "get_invoice_status",
        "description": "Look up whether an invoice is paid.",
        "parameters": {
            "type": "object",
            "properties": {"invoice_id": {"type": "string"}},
            "required": ["invoice_id"],
        },
    },
}]
messages = [{"role": "user", "content": "Has invoice a41f been paid?"}]

reply = client.chat.completions.create(model="anthropic/claude-sonnet-5", messages=messages, tools=tools)
call = reply.choices[0].message.tool_calls[0]
args = json.loads(call.function.arguments)

messages.append(reply.choices[0].message)
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps({"paid": True})})

final = client.chat.completions.create(model="anthropic/claude-sonnet-5", messages=messages, tools=tools)
print(final.choices[0].message.content)
```

JavaScript

```
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const tools = [{
  type: "function",
  function: {
    name: "get_invoice_status",
    description: "Look up whether an invoice is paid.",
    parameters: {
      type: "object",
      properties: { invoice_id: { type: "string" } },
      required: ["invoice_id"],
    },
  },
}];
const messages = [{ role: "user", content: "Has invoice a41f been paid?" }];

const reply = await client.chat.completions.create({ model: "anthropic/claude-sonnet-5", messages, tools });
const call = reply.choices[0].message.tool_calls[0];
const args = JSON.parse(call.function.arguments);

messages.push(reply.choices[0].message);
messages.push({ role: "tool", tool_call_id: call.id, content: JSON.stringify({ paid: true }) });

const final = await client.chat.completions.create({ model: "anthropic/claude-sonnet-5", messages, tools });
console.log(final.choices[0].message.content);
```

## Pictures in a request

Models with `capabilities.vision` can read pictures. Add an `image_url` part to a user message, with either a public `https://` link or a `data:image/…;base64,` URL. Up to 20 pictures per request. SVG pictures are refused, and so is a link over 4,096 characters or one with a user name or password in it (`400` `invalid_image_url`). An optional `detail` is `auto`, `low` or `high`.

With `nymbot/auto`, a request with a picture in it is routed to a standard model that can see. A catalog model without `capabilities.vision` refuses pictures with `400` `unsupported_content`. Pictures can only be in user messages. A link must point to a public host; Nymbot passes the picture to the model's provider and does not keep it.

A picture is charged as the input tokens the provider counts for it, like the rest of the request.

cURL

```
curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "What is in this picture?"},
        {"type": "image_url", "image_url": {"url": "https://example.com/receipt.jpg"}}
      ]
    }]
  }'
```

Python

```
import base64, os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

with open("receipt.jpg", "rb") as f:
    data_url = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode()

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this picture?"},
            {"type": "image_url", "image_url": {"url": data_url}},
        ],
    }],
)
print(reply.choices[0].message.content)
```

JavaScript

```
import { readFile } from "node:fs/promises";
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const dataUrl = "data:image/jpeg;base64," + (await readFile("receipt.jpg")).toString("base64");

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5",
  messages: [{
    role: "user",
    content: [
      { type: "text", text: "What is in this picture?" },
      { type: "image_url", image_url: { url: dataUrl } },
    ],
  }],
});
console.log(reply.choices[0].message.content);
```

## Reasoning

Models with `capabilities.reasoning` can think before they answer. Ask for more or less of it with `reasoning_effort` (`"minimal"`, `"low"`, `"medium"` or `"high"`) or `"reasoning": {"effort": "high"}`, or add `:thinking` to the model name, which means high effort. On a model without reasoning the setting is ignored. With `nymbot/auto`, `:thinking` sends the request to the standard reasoning route.

On Anthropic models the effort becomes a thinking budget of about 1,000, 2,000, 8,000 or 16,000 tokens, never more than `max_tokens` allows. Thinking is left off when `tool_choice` forces a tool, and when the request continues a tool loop (its last message is a tool result), because Anthropic needs the earlier signed thinking to resume it. The same applies on the Responses and Messages endpoints.

The reasoning comes back in `message.reasoning_content`, or `delta.reasoning_content` when streaming, never mixed into the answer. Reasoning is output and is charged as such, inside `completion_tokens`. Some providers do not return the reasoning text at all, and it is still charged.

cURL

```
curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "reasoning_effort": "high",
    "messages": [{"role": "user", "content": "Is 2^61 - 1 prime? Show why."}]
  }'
```

Python

```
import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    reasoning_effort="high",
    messages=[{"role": "user", "content": "Is 2^61 - 1 prime? Show why."}],
)
message = reply.choices[0].message
print(getattr(message, "reasoning_content", None))
print(message.content)
```

JavaScript

```
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5",
  reasoning_effort: "high",
  messages: [{ role: "user", content: "Is 2^61 - 1 prime? Show why." }],
});
console.log(reply.choices[0].message.reasoning_content);
console.log(reply.choices[0].message.content);
```

## Model suffixes

A suffix on the model name changes how the request is handled, without another field. They work on full ids and short names alike, as in `anthropic/claude-sonnet-5:online`.

| Suffix | Effect |
| --- | --- |
| `:online` | Searches the web first, like `plugins: [{"id": "web"}]`. See [web search](#web-search). |
| `:thinking` | High reasoning effort; on `nymbot/auto`, the reasoning route. On a model that cannot reason, `400` `model_not_found` with “no endpoints found”. |
| `:nitro`, `:floor`, `:exacto`, `:extended` | Accepted and ignored. Each catalog model has one route, so there is no faster, cheaper or longer one to pick; the suffixes are allowed so model names copied from other services still work. |

Any other suffix is ignored. The whole name is tried first, so a model whose id really contains a colon still works; failing that, suffixes are taken off the end one at a time until a model matches. A model name may be at most 200 characters long with at most 4 suffixes; a longer one is refused with `400` `invalid_value`.

cURL

```
curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5:thinking",
    "messages": [{"role": "user", "content": "Plan a three-day trip to Lisbon."}]
  }'
```

Python

```
import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5:thinking",
    messages=[{"role": "user", "content": "Plan a three-day trip to Lisbon."}],
)
print(reply.choices[0].message.content)
```

JavaScript

```
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5:thinking",
  messages: [{ role: "user", content: "Plan a three-day trip to Lisbon." }],
});
console.log(reply.choices[0].message.content);
```

## Web search

Any chat model can answer from the live web. Nymbot searches for the last user message (its first 2,000 characters), reads the best pages and gives them to the model with the question, marked as outside content the model should not take instructions from. There are two modes:

- **Always search**: `"plugins": [{"id": "web", "max_results": 5}]`, or a `:online` suffix on the model.
- **Search when it helps**: `"tools": [{"type": "web_search", "parameters": {"max_results": 5}}]`, also accepted as `web_search_preview` or `openrouter:web_search`. Nymbot searches only when the question looks like it needs current information, the same test the app uses.

`max_results` is 5 by default and at most 10. The sources come back in `nymbot.web_search.sources`, each with a `title`, `snippet` and `url`, and as `url_citation` annotations on the message.

Each search that runs costs $0.008, converted to sats, on top of the tokens, and it is part of the hold. The pages it read are input tokens too, so an answer from the web costs more than the same question asked cold, sometimes several times more.

Response, in part

```
"nymbot": {
  "balance": "pro",
  "charged_credits": 0.431,
  "charged_sats": 43.1,
  "balance_credits": 411.994,
  "balance_sats": 41199.4,
  "web_search": {
    "sources": [
      { "title": "Lightning Network - Wikipedia", "snippet": "The Lightning Network is a payment protocol...", "url": "https://en.wikipedia.org/wiki/Lightning_Network" }
    ]
  }
}
```

cURL

```
curl https://nymbot.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "plugins": [{"id": "web", "max_results": 5}],
    "messages": [{"role": "user", "content": "What changed in the latest Bitcoin Core release?"}]
  }'
```

Python

```
import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

reply = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "What changed in the latest Bitcoin Core release?"}],
    extra_body={"plugins": [{"id": "web", "max_results": 5}]},
)
print(reply.choices[0].message.content)
for source in reply.model_extra["nymbot"]["web_search"]["sources"]:
    print(source["url"])
```

JavaScript

```
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const reply = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5",
  plugins: [{ id: "web", max_results: 5 }],
  messages: [{ role: "user", content: "What changed in the latest Bitcoin Core release?" }],
});
console.log(reply.choices[0].message.content);
for (const source of reply.nymbot.web_search.sources) console.log(source.url);
```

## Responses API

OpenAI's newer format, used by the OpenAI Agents SDK and by Codex. It runs on the same models, billing and features as Chat Completions.

`POST https://nymbot.ai/api/v1/responses` — needs an API key.

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `model` | string | Yes | As for [Chat Completions](#chat-completions), suffixes included. |
| `input` | string or array | Yes | A string, or a list of items: messages (roles `user`, `assistant`, `system`, `developer`) with `input_text`, `input_image` and `output_text` parts, and `function_call` and `function_call_output` items for tools. `input_image` takes an `image_url` link or data URL, not a file id. `reasoning` and `web_search_call` items are skipped; `item_reference` is refused. |
| `instructions` | string | No | System instructions. |
| `max_output_tokens` | integer | No | The most tokens to write. |
| `temperature` `top_p` | number | No | Sampling controls, where the model takes them. |
| `tools` `tool_choice` `parallel_tool_calls` | array, string or object, boolean | No | Function tools, in the Responses shape (`{"type": "function", "name": …, "parameters": …}`). A `web_search` or `web_search_preview` tool turns on [web search](#web-search) when it helps. Other built-in tools are refused. |
| `reasoning` | object | No | `{"effort": "minimal" \| "low" \| "medium" \| "high"}`. `xhigh` and `max` mean `high`; `none` turns it off. |
| `text.format` `response_format` | object | No | Structured output, as a JSON schema or `json_object`. |
| `metadata` | object | No | Returned unchanged in the response. At most 16 string values, keys up to 64 characters and values up to 512. |
| `stream` | boolean | No | Stream events as described below. |
| `store` | boolean | No | Ignored. Nothing is stored, and the response always says `"store": false`. |
| `previous_response_id` `conversation` `background` | string, object, boolean | No | Not supported: `400` `unsupported_parameter`. Responses are not stored, so send the whole conversation in `input` every time. |

Response

```
{
  "id": "resp_8c1e4b0f9a2d4e61",
  "object": "response",
  "created_at": 1790726400,
  "status": "completed",
  "model": "anthropic/claude-sonnet-5",
  "output": [
    {
      "type": "message",
      "id": "msg_2b7f",
      "role": "assistant",
      "status": "completed",
      "content": [{ "type": "output_text", "text": "A Lightning invoice is...", "annotations": [] }]
    }
  ],
  "output_text": "A Lightning invoice is...",
  "usage": {
    "input_tokens": 1240,
    "input_tokens_details": { "cached_tokens": 0 },
    "output_tokens": 380,
    "output_tokens_details": { "reasoning_tokens": 0 },
    "total_tokens": 1620
  },
  "incomplete_details": null,
  "error": null,
  "instructions": null,
  "store": false,
  "previous_response_id": null,
  "metadata": {},
  "nymbot": { "balance": "pro", "charged_credits": 0.162, "charged_sats": 16.2, "balance_credits": 412.425, "balance_sats": 41242.5 }
}
```

A tool request shows up in `output` as a `function_call` item with `call_id`, `name` and `arguments`; send the result back as a `function_call_output` item with the same `call_id`. Reasoning, when the model returns it, is a `reasoning` item with `reasoning_text` content, listed first. `status` is `incomplete` when the answer hit `max_output_tokens` or the provider refused, with `incomplete_details.reason` set to `max_output_tokens` or `content_filter`. The response also echoes the request's settings (temperature, tools, tool choice and so on) as OpenAI's does, and carries the `nymbot` cost object.

Streamed, each event is an `event:` line and a `data:` line with a `sequence_number`, in this order: `response.created`, `response.in_progress`, `response.output_item.added`, `response.content_part.added`, any number of `response.output_text.delta`, `response.output_text.done`, `response.content_part.done`, `response.output_item.done`, and finally `response.completed` with the usage and cost. An answer cut short ends with `response.incomplete` instead, and a failure after the stream started with `response.failed`. Reasoning streams as its own item with `response.reasoning_text.delta` and `.done`. Tool calls come after the message, each as an item with `response.function_call_arguments.delta` and `.done`.

| Status | When |
| --- | --- |
| `400` | No `model` or `input`; `previous_response_id`, `conversation`, `background` or an `item_reference` (`unsupported_parameter`); an unsupported tool or content type. |
| `402`, `403`, `404`, `429`, `502`, `503` | As for Chat Completions. |

cURL

```
curl https://nymbot.ai/api/v1/responses \
  -H "Authorization: Bearer $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "instructions": "Answer in one short paragraph.",
    "input": "What is a Lightning invoice?"
  }'
```

Python

```
import os
from openai import OpenAI

client = OpenAI(base_url="https://nymbot.ai/api/v1", api_key=os.environ["NYMBOT_API_KEY"])

response = client.responses.create(
    model="anthropic/claude-sonnet-5",
    instructions="Answer in one short paragraph.",
    input="What is a Lightning invoice?",
)
print(response.output_text)
```

JavaScript

```
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://nymbot.ai/api/v1", apiKey: process.env.NYMBOT_API_KEY });

const response = await client.responses.create({
  model: "anthropic/claude-sonnet-5",
  instructions: "Answer in one short paragraph.",
  input: "What is a Lightning invoice?",
});
console.log(response.output_text);
```

## Anthropic Messages

Anthropic's format, for the Anthropic SDKs and Claude Code. It works with every model in the catalog, not only Claude: the request is translated, run through the same pipeline, and translated back.

`POST https://nymbot.ai/api/v1/messages` — needs an API key, as `x-api-key` or `Authorization: Bearer`.

The `anthropic-version` and `anthropic-beta` headers are accepted and ignored. Anthropic model names are matched to the catalog, so Claude Code and the SDKs work with the names they already use:

- A name the catalog knows, such as `claude-sonnet-5` or `anthropic/claude-opus-5`, is used as it is.
- Otherwise a date (`-20260514`), `-latest`, a version tag such as `-v1`, a bracketed tag such as `[1m]` and an `anthropic/` or `anthropic.` prefix are taken off, and dots and dashes in the version are tried both ways (`claude-haiku-4-5` finds `claude-haiku-4.5`).
- If that still matches nothing, the family (Opus, Sonnet or Haiku) is used, as long as the catalog's version is the same or newer than the one asked for.
- A name that matches nothing returns `404` `not_found_error`.

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `model` | string | Yes | A catalog model id, or an Anthropic model name. |
| `max_tokens` | integer | Yes | The most tokens to write. |
| `messages` | array | Yes | `user` and `assistant` turns, with `text`, `image` (base64 or URL source), `tool_use` and `tool_result` blocks. `thinking` blocks from earlier turns are accepted and dropped. |
| `system` | string or array | No | System prompt, as a string or text blocks. |
| `temperature` `top_p` | number | No | Sampling controls, where the model takes them. `top_k` is accepted and dropped. |
| `stop_sequences` | array of strings | No | Text that ends the answer. At most 4 strings, each at most 256 characters. |
| `tools` `tool_choice` | array, object | No | Tools with `name`, `description` and `input_schema`. A `web_search` server tool turns on [web search](#web-search); Anthropic's other built-in tools (bash, text editor, computer use) are refused with `unsupported_tool`. `tool_choice` takes `auto`, `any`, `tool` or `none`, and `disable_parallel_tool_use`. |
| `thinking` | object | No | `{"type": "enabled", "budget_tokens": 8192}`, `{"type": "adaptive"}` or `{"type": "disabled"}`. The budget picks an effort level: under 2,048 minimal, from 2,048 low, from 8,192 medium, from 16,384 high. Adaptive uses `output_config.effort`, or high. |
| `stream` | boolean | No | Stream in Anthropic's event format. |
| `metadata` | object | No | Accepted and ignored. |

Response

```
{
  "id": "msg_01c7a2f93e5b4d08",
  "type": "message",
  "role": "assistant",
  "model": "anthropic/claude-sonnet-5",
  "content": [{ "type": "text", "text": "A Lightning invoice is..." }],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 1240,
    "output_tokens": 380,
    "cache_read_input_tokens": 0,
    "cache_creation_input_tokens": 0
  },
  "nymbot": { "balance": "pro", "charged_credits": 0.162, "charged_sats": 16.2, "balance_credits": 412.425, "balance_sats": 41242.5 }
}
```

`content` can also hold `tool_use` blocks and a `thinking` block, whose `signature` is empty. `stop_reason` is `end_turn`, `max_tokens`, `tool_use` or `refusal`, worked out the same way as `finish_reason` on [Chat Completions](#chat-completions); `stop_sequence` is always `null`, even when a stop sequence ended the answer. `input_tokens` counts fresh input only; cached input is in the two cache fields. The cost is in the `nymbot` object and the `X-Nymbot-Cost-Sats` header.

Streamed, the events are Anthropic's: `message_start`, `content_block_start`, `ping`, `content_block_delta` (`text_delta`, `input_json_delta` or `thinking_delta`), `content_block_stop`, `message_delta` with the stop reason, the usage and the `nymbot` cost object, and `message_stop`. A `ping` is also sent every 15 seconds while the model is working. Tool calls arrive after the text, each as a `tool_use` block with its whole input in one `input_json_delta`.

Errors on this endpoint use Anthropic's format: `{"type": "error", "error": {"type": "not_found_error", "message": "…"}}`. A short balance is `402` `billing_error`, an overloaded provider `503` `overloaded_error`. A failure after the stream has started is sent as an `error` event.

cURL

```
curl https://nymbot.ai/api/v1/messages \
  -H "x-api-key: $NYMBOT_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "What is a Lightning invoice?"}]
  }'
```

Python

```
import os
import anthropic

client = anthropic.Anthropic(base_url="https://nymbot.ai/api", api_key=os.environ["NYMBOT_API_KEY"])

message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "What is a Lightning invoice?"}],
)
print(message.content[0].text)
```

JavaScript

```
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({ baseURL: "https://nymbot.ai/api", apiKey: process.env.NYMBOT_API_KEY });

const message = await client.messages.create({
  model: "claude-sonnet-5",
  max_tokens: 1024,
  messages: [{ role: "user", content: "What is a Lightning invoice?" }],
});
console.log(message.content[0].text);
```

## Counting tokens

An estimate of how many input tokens a Messages request would use, so a client can check before it sends. It is free, but still needs a key.

`POST https://nymbot.ai/api/v1/messages/count_tokens` — needs an API key. Free.

The body is the same as for [Messages](#messages), without `max_tokens`; the model name has to resolve. The count is an estimate: the characters of the system prompt, messages, tool calls and tool definitions divided by four, plus 1,600 for each picture. It is not the provider's own tokenizer, so the real count can differ. A key that has reached its cap can still use it.

Response

```
{ "input_tokens": 318 }
```

cURL

```
curl https://nymbot.ai/api/v1/messages/count_tokens \
  -H "x-api-key: $NYMBOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "messages": [{"role": "user", "content": "What is a Lightning invoice?"}]
  }'
```

Python

```
import os
import anthropic

client = anthropic.Anthropic(base_url="https://nymbot.ai/api", api_key=os.environ["NYMBOT_API_KEY"])

count = client.messages.count_tokens(
    model="claude-sonnet-5",
    messages=[{"role": "user", "content": "What is a Lightning invoice?"}],
)
print(count.input_tokens)
```

JavaScript

```
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({ baseURL: "https://nymbot.ai/api", apiKey: process.env.NYMBOT_API_KEY });

const count = await client.messages.countTokens({
  model: "claude-sonnet-5",
  messages: [{ role: "user", content: "What is a Lightning invoice?" }],
});
console.log(count.input_tokens);
```

## Listing models

Every model and generator the API accepts, with what it costs. The list is read from the same live catalog as the app's picker, so it is always what the server will run.

`GET https://nymbot.ai/api/v1/models` — no key needed. Cached for five minutes.

`GET /api/v1/models/{id}` returns one entry.

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `type` | string (query) | No | `chat` (the default), `image`, `video`, `audio`, `embedding` or `all`. Several can be given, as `image,video`. |

Prices are what you pay, with the fee and margin already in, in dollars and in sats at the current Bitcoin price. `balance` says which balance the model spends. `nymbot_key` is the model's short name in the app. `created` is always 0, since the catalog does not record when a model was added. A model without published token rates is priced `per_request` instead.

`nymbot/auto` is always first, priced as `variable`, with a `routes` list giving each standard route's rates. It lists vision and reasoning, but not tools.

`GET /api/v1/models/{id}` accepts the same names and aliases as a request, and returns the entry for the model they resolve to.

Response

```
{
  "object": "list",
  "data": [
    {
      "id": "anthropic/claude-sonnet-5",
      "object": "model",
      "type": "chat",
      "owned_by": "anthropic",
      "name": "Claude Sonnet 5",
      "created": 0,
      "context_length": 1000000,
      "max_output_tokens": 64000,
      "architecture": { "input_modalities": ["text", "image"], "output_modalities": ["text"] },
      "supported_parameters": ["max_tokens", "temperature", "tools", "tool_choice", "reasoning", "response_format", "stop"],
      "capabilities": { "vision": true, "video": false, "tools": true, "reasoning": true, "web_search": true },
      "balance": "pro",
      "pricing": {
        "type": "per_token",
        "currency": "USD",
        "input_per_1M_tokens": 4.725,
        "output_per_1M_tokens": 23.625,
        "cache_read_per_1M_tokens": 0.4725,
        "sats_input_per_1M_tokens": 4038,
        "sats_output_per_1M_tokens": 20192
      },
      "description": "...",
      "nymbot_key": "claude-sonnet"
    }
  ]
}
```

The other types:

- **Image** entries have `capabilities` (`accepts_image_url`, `requires_image_url`, `edit`) and a price `per_generation`.
- **Video** entries have `max_duration_seconds` and `resolutions`, and a price `per_second` for each resolution.
- **Audio** entries have `audio_type` `speech` or `transcription`, priced `per_1k_chars` or `per_minute`.
- **Embedding** entries have `dimensions`, `context_length`, `max_inputs` and a price per million input tokens.

Prices marked `"estimated": true` are the app's estimate for a generator whose price is not published. An unknown `type` returns `400`.

cURL

```
curl "https://nymbot.ai/api/v1/models?type=chat"
```

Python

```
import requests

models = requests.get("https://nymbot.ai/api/v1/models", params={"type": "chat"}).json()["data"]
for m in models:
    print(m["id"], m["balance"], m["pricing"].get("input_per_1M_tokens"))
```

JavaScript

```
const res = await fetch("https://nymbot.ai/api/v1/models?type=chat");
const { data } = await res.json();
for (const m of data) console.log(m.id, m.balance, m.pricing.input_per_1M_tokens);
```
