Skip to content

Build your own agents on ZenithAI

An agent is a program that reads something, thinks about it, and takes a step, again and again until the work is done. With ZenithAI, the thinking and reading happen on your own server, and the steps are taken by your code. Think of ZenithAI as a very capable munshi (clerk) who reads, checks and drafts, while your program decides what gets filed, sent or paid.

You need to Use
Understand text, decide, draft POST /v1/chat/completions
Get fields you can trust in code response_format with a JSON schema
Read photos and screenshots image_url parts on /v1/chat/completions
Read PDFs exactly, or scans POST /pdf, POST /ocr
Listen to audio, watch video POST /chat with attachments
Look things up on the web POST /websearch
Make Word, Excel or PowerPoint files POST /chat or POST /generate, then download
Act in your own systems Function calling (tools), Smart and Genius

The full list, with limits, is in Built-in tools.

With function calling, the model can ask your code to run a function. The loop is always the same:

  1. Send the task, plus a description of your functions.
  2. If the model replies with tool_calls, run those functions in your code.
  3. Send the results back as tool messages and ask again.
  4. Stop when the model gives a plain answer, or when you reach your step limit.
import json
from openai import OpenAI
client = OpenAI(base_url="https://dummydomain/v1", api_key="YOUR_API_KEY")
# Your own function. ZenithAI never runs it; your code does.
def get_order_status(order_id: str) -> dict:
orders = {"SO-1042": {"status": "dispatched", "expected": "09-10-2026", "courier": "Blue Line"}}
return orders.get(order_id, {"status": "not found"})
FUNCTIONS = {"get_order_status": get_order_status}
TOOLS = [{
"type": "function",
"function": {
"name": "get_order_status",
"description": "Look up a sales order by its number, such as SO-1042.",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
"additionalProperties": False,
},
},
}]
messages = [
{"role": "system", "content": "You help the sales desk. Use the tools for facts. Never guess an order status."},
{"role": "user", "content": "A customer asks where order SO-1042 is. Reply in two lines."},
]
for step in range(5): # a hard limit, so a confused agent can't loop for ever
choice = client.chat.completions.create(model="Smart", messages=messages, tools=TOOLS).choices[0]
if choice.finish_reason == "length":
raise RuntimeError("The reply was cut off. Raise max_tokens or ask for less.")
msg = choice.message
if not msg.tool_calls:
print(msg.content)
break
messages.append({"role": "assistant", "content": msg.content or "",
"tool_calls": [c.model_dump() for c in msg.tool_calls]})
for call in msg.tool_calls:
fn = FUNCTIONS.get(call.function.name)
result = fn(**json.loads(call.function.arguments)) if fn else {"error": "unknown function"}
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
else:
print("Stopped after 5 steps. Send this one to a person.")

A real run of the Python loop above. Press Replay to watch the model ask for the function, then answer.

Recorded on 07-10-2026 against a ZenithAI test server, running the code above. The times are real.

What went in

The question in the code: "A customer asks where order SO-1042 is."

What the agent did

  1. POST /v1/chat/completionsSmart0.3 s

    Smart asked your code to run get_order_status.

    Request and response
    Request
    {
      "messages": [
        {
          "role": "system",
          "content": "You help the sales desk. Use the tools for facts. Never guess an order status."
        },
        {
          "role": "user",
          "content": "A customer asks where order SO-1042 is. Reply in two lines."
        }
      ],
      "tools": [
        {
          "type": "function",
          "function": {
            "name": "get_order_status",
            "description": "Look up a sales order by its number, such as SO-1042.",
            "parameters": {
              "type": "object",
              "properties": {
                "order_id": {
                  "type": "string"
                }
              },
              "required": [
                "order_id"
              ],
              "additionalProperties": false
            }
          }
        }
      ]
    }
    Response 200
    {
      "choices": [
        {
          "finish_reason": "tool_calls",
          "index": 0,
          "message": {
            "content": null,
            "role": "assistant",
            "tool_calls": [
              {
                "function": {
                  "arguments": "{\"order_id\":\"SO-1042\"}",
                  "name": "get_order_status"
                },
                "index": 0,
                "type": "function"
              }
            ]
          }
        }
      ],
      "object": "chat.completion",
      "usage": {
        "completion_tokens": 24,
        "prompt_tokens": 115,
        "total_tokens": 139
      }
    }
  2. POST /v1/chat/completionsSmart0.3 s

    Smart read the function result and wrote the answer.

    Request and response
    Request
    {
      "messages": [
        {
          "role": "system",
          "content": "You help the sales desk. Use the tools for facts. Never guess an order status."
        },
        {
          "role": "user",
          "content": "A customer asks where order SO-1042 is. Reply in two lines."
        },
        {
          "role": "assistant",
          "content": "",
          "tool_calls": [
            {
              "function": {
                "arguments": "{\"order_id\":\"SO-1042\"}",
                "name": "get_order_status"
              },
              "type": "function",
              "index": 0
            }
          ]
        },
        {
          "role": "tool",
          "tool_call_id": "call_v2j6ecy5",
          "content": "{\"status\": \"dispatched\", \"expected\": \"09-10-2026\", \"courier\": \"Blue Line\"}"
        }
      ],
      "tools": [
        {
          "type": "function",
          "function": {
            "name": "get_order_status",
            "description": "Look up a sales order by its number, such as SO-1042.",
            "parameters": {
              "type": "object",
              "properties": {
                "order_id": {
                  "type": "string"
                }
              },
              "required": [
                "order_id"
              ],
              "additionalProperties": false
            }
          }
        }
      ]
    }
    Response 200
    {
      "choices": [
        {
          "finish_reason": "stop",
          "index": 0,
          "message": {
            "content": "Your order SO-1042 has been dispatched via Blue Line.\nIt is expected to arrive by 09-10-2026.",
            "role": "assistant"
          }
        }
      ],
      "object": "chat.completion",
      "usage": {
        "completion_tokens": 39,
        "prompt_tokens": 179,
        "total_tokens": 218
      }
    }

What the program printed

Your order SO-1042 has been dispatched via Blue Line.
It is expected to arrive by 09-10-2026.
  • The model suggests, your code decides. Payments, emails to customers and changes in your ERP should go through your own checks, and a person when the stakes are high.
  • Check the shape. Ask for JSON with a schema and validate it in code before you use it. See Structured JSON output.
  • Check facts in code. Formats such as GSTIN, PAN, IFSC and phone numbers are rules, so test them with code, not with the model.
  • Treat input as data. A web form, an email or a PDF can contain text like “ignore your instructions”. Put such content inside the user message, say clearly that it is data, and never let it choose which function runs without your checks.
  • Cap the steps and the time. Use a step limit, a timeout on every call, and a deadline when you poll a job.
  • Handle cut-off answers. finish_reason: "length", truncated: true or a stream without [DONE] means the answer is not complete.
  • Respect limits. By default each person can make 10 requests a minute to /chat, /generate and /pdf together, and your administrator can change this. /v1/chat/completions, /websearch and OCR have their own limits. On 429, wait for Retry-After, then try again. Run a few jobs at a time, not hundreds.
  • Log the reason. Store the model’s answer next to the action your code took, so anyone can see later why something happened.

Common questions

Does ZenithAI host or run my agent?

No. Your agent is your own program, such as a Python script, a scheduled job or a service. It calls ZenithAI endpoints for the thinking, reading and file making, and your code takes the actions. Your data stays on your server either way.

Which ZenithAI tier should an agent use?

Use Instant or Fast for quick sorting, checks and pulling out fields. Use Smart when the agent must call your functions or reason over long material, and Genius for the hardest cases. Check capabilities in GET /v1/models first.

How do I stop an agent from doing something wrong?

Let the model suggest and let your code decide. Check every answer against a schema, cap the number of steps, keep a person in the loop for money, legal or customer-facing actions, and log what the agent did and why.

Can an agent call our ERP or CRM?

Yes, through function calling on Smart or Genius. You describe your functions, the model asks for one with arguments, your code calls the ERP or CRM and sends back the result. ZenithAI never reaches your systems directly.