OpenAI-compatible chat
POST /v1/chat/completions takes a list of messages and returns the next assistant message. It is
stateless: for a follow-up, send the earlier user and assistant messages again yourself.
from openai import OpenAI
client = OpenAI(base_url="https://dummydomain/v1", api_key="YOUR_API_KEY")
history = [{"role": "user", "content": "Explain why leaves look green."}]first = client.chat.completions.create(model="Fast", messages=history, max_tokens=512)history.append({"role": "assistant", "content": first.choices[0].message.content})
history.append({"role": "user", "content": "Now explain it to a ten-year-old."})second = client.chat.completions.create(model="Fast", messages=history)print(second.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://dummydomain/v1", apiKey: process.env.API_KEY });
const history = [{ role: "user", content: "Explain why leaves look green." }];const first = await client.chat.completions.create({ model: "Fast", messages: history, max_tokens: 512 });history.push({ role: "assistant", content: first.choices[0].message.content });curl --fail-with-body "https://dummydomain/v1/chat/completions" \ -H "Authorization: Bearer $API_KEY" -H 'Content-Type: application/json' \ -d '{"model":"Fast","max_tokens":512, "messages":[{"role":"user","content":"Explain why leaves look green."}]}'Streaming
Section titled “Streaming”Set stream: true. Replies arrive as server-sent events with choices[].delta, ending with
data: [DONE]. Add stream_options: {"include_usage": true} to receive token usage at the end.
stream = client.chat.completions.create( model="Fast", stream=True, messages=[{"role": "user", "content": "Write a two-line welcome note."}],)for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True)curl --no-buffer "https://dummydomain/v1/chat/completions" \ -H "Authorization: Bearer $API_KEY" -H 'Content-Type: application/json' \ -d '{"model":"Fast","stream":true,"messages":[{"role":"user","content":"Write a two-line welcome note."}]}'Parameters
Section titled “Parameters”| Field | Notes |
|---|---|
model |
Instant, Fast, Smart or Genius (when your server has it) |
messages |
system, developer, user, assistant and tool roles |
max_tokens / max_completion_tokens |
Positive integer; if you send both, they must match |
temperature |
0 to 2 |
top_p |
0 to 1 |
presence_penalty, frequency_penalty |
−2 to 2 |
stop |
A string or up to 4 strings |
seed |
Signed 64-bit integer |
response_format |
See Structured JSON output |
tools, tool_choice, parallel_tool_calls |
On tiers with capabilities.tools; your code runs the functions |
n other than 1, functions, function_call and logprobs are refused.
Errors
Section titled “Errors”Errors return an error object with message, type, code, param where relevant, and a
request_id. Keep the X-Request-ID response header when you report a problem.
| Status | What to do |
|---|---|
| 400 | Fix the request; the code and param say what is wrong |
| 401 | Check the key and that you are sending Authorization: Bearer |
| 429 | Wait for Retry-After if present, then retry with back-off |
| 5xx | Retry with back-off, only if repeating the request is safe |