Skip to content

Agent example: check site videos

The job: a supervisor records a short walk-through of a site, a warehouse aisle or a shop floor. The agent watches it, notes what it sees, and returns a checklist: helmets, blocked exits, spills, loose cables. Your code then raises tickets for anything marked “fail”.

Tools used: video reading on POST /chat, then structured JSON on /v1/chat/completions.

  • It takes one frame every second, up to 60 frames, so about the first minute.
  • It turns speech in the soundtrack into text, when there is a soundtrack.
  • It reads MP4, MOV, M4V, MKV, AVI and WebM files. Inside them, H.264 and H.265 (HEVC) work, so phone and most CCTV videos go in as they are. VP8, VP9, MPEG-4 and Motion JPEG work too.
  • A video can be up to 100 MB and 5 minutes long.
  • Send one video per message.

The agent works in two steps. First it asks ZenithAI to describe the video in plain words. Then it turns that description into checklist JSON your code can trust.

import base64, os
from pathlib import Path
from typing import Literal
import requests
from openai import OpenAI
from pydantic import BaseModel
BASE = "https://dummydomain"
KEY = os.environ["API_KEY"]
client = OpenAI(base_url=BASE + "/v1", api_key=KEY)
def watch(clip: Path) -> str:
r = requests.post(f"{BASE}/chat", headers={"Authorization": "Bearer " + KEY}, timeout=900, json={
"message": "Describe this site walk-through second by second. Note people and whether they wear "
"helmets and safety vests, exits and whether they are clear, spills, loose cables, "
"and anything said about safety. Say clearly when something is not visible.",
"attachments": [{"name": clip.name, "data": base64.b64encode(clip.read_bytes()).decode()}],
"stream": False,
})
r.raise_for_status()
return r.json()["reply"]
class Item(BaseModel):
check: Literal["helmets", "safety_vests", "exits_clear", "no_spills", "cables_secured"]
result: Literal["pass", "fail", "not_visible"]
evidence: str
class Checklist(BaseModel):
items: list[Item]
summary: str
def to_checklist(description: str) -> Checklist:
reply = client.chat.completions.create(
model="Fast",
messages=[{"role": "system", "content": "Fill the safety checklist only from the description. "
"Use not_visible when the description does not show it."},
{"role": "user", "content": description}],
response_format={"type": "json_schema", "json_schema": {
"name": "Checklist", "schema": Checklist.model_json_schema(), "strict": True}},
)
return Checklist.model_validate_json(reply.choices[0].message.content)
checklist = to_checklist(watch(Path("site-walk.mp4")))
for item in checklist.items:
if item.result == "fail":
print("Raise a ticket:", item.check, "|", item.evidence)
print(checklist.summary)

A real run of the code above, on made-up sample files. Press Replay to watch the steps in order.

Recorded on 07-10-2026 against a ZenithAI test server, running the code above. The times are real.

What went in

site-walk.mp4, a drawn warehouse aisle with a spoken note, saved the way an Android phone saves video (H.264)

What the agent did

  1. POST /chat11.7 s

    Sent site-walk.mp4 to ZenithAI, which watched and listened to it and answered.

    Request and response
    Request
    {
      "message": "Describe this site walk-through second by second. Note people and whether they wear helmets and safety vests, exits and whether they are clear, spills, loose cables, and anything said about safety. Say clearly when something is not visible.",
      "attachments": [
        {
          "name": "site-walk.mp4",
          "data": "<file, 186,234 bytes>"
        }
      ],
      "stream": false
    }
    Response 200
    {
      "attempted_route": "general",
      "current_route": "general",
      "finish_reason": "stop",
      "max_tokens": 4096,
      "num_predict": 4096,
      "reply": "Based on the site walk-through video provided, here is a second-by-second description of the scene:\n\n*   **00:00 - 00:13:** A warehouse aisle is shown. A single worker is visible standing near a storage rack on the left side of the aisle. The worker is wearing an **orange safety vest** but is **not wearing a helmet**. At the end of the aisle, there is a green **FIRE EXIT** sign, but the exit path is **blocked by several cardboard boxes**. The floor appears dry and clean, with no visible spills or loose cables.\n\n**Safety Summary:**\n*   **PPE:** The worker is wearing a safety vest but is missing a helmet.\n*   **Exits:** The fire exit is obstructed by boxes.\n*   **Housekeeping:** The floor is d ...",
      "route_disagreement": false,
      "session_id": "773240cf-c238-47b6-b11d-2d4c55018a72",
      "shadow_confidence": 0.7,
      "shadow_evidence_status": "n/a",
      "shadow_reason": "no tool route necessary; general knowledge",
      "shadow_route": "general"
    }
  2. POST /v1/chat/completionsFast4.9 s

    Fast returned JSON that matches the Checklist schema.

    Request and response
    Request
    {
      "messages": [
        {
          "role": "system",
          "content": "Fill the safety checklist only from the description. Use not_visible when the description does not show it."
        },
        {
          "role": "user",
          "content": "Based on the site walk-through video provided, here is a second-by-second description of the scene:\n\n*   **00:00 - 00:13:** A warehouse aisle is shown. A single worker is visible standing near a storage rack on the left side of the aisle. The worker is wearing an **orange safety vest** but is **not wearing a helmet**. At the end of the aisle, there is a green **FIRE EXIT** sign, but the exit path is **blocked by several cardboard boxes**. The floor appears dry and clean, with no visible spills or loose cables.\n\n**Safety Summary:**\n*   **PPE:** The worker is wearing a safety vest but is missing a helmet.\n*   **Exits:** The fire exit is obstructed by boxes.\n*   **Housekeeping:** The floor is d ..."
        }
      ],
      "response_format": {
        "type": "json_schema",
        "json_schema": {
          "name": "Checklist",
          "schema": {
            "$defs": {
              "Item": {
                "properties": {
                  "check": {
                    "enum": [
                      "helmets",
                      "safety_vests",
                      "exits_clear",
                      "no_spills",
                      "cables_secured"
                    ],
                    "title": "Check",
                    "type": "string"
                  },
                  "result": {
                    "enum": [
                      "pass",
                      "fail",
                      "not_visible"
                    ],
                    "title": "Result",
                    "type": "string"
                  },
                  "evidence": {
                    "title": "Evidence",
                    "type": "string"
                  }
                },
                "required": [
                  "check",
                  "result",
                  "evidence"
                ],
                "title": "Item",
                "type": "object"
              }
            },
            "properties": {
              "items": {
                "items": {
                  "$ref": "#/$defs/Item"
                },
                "title": "Items",
                "type": "array"
              },
              "summary": {
                "title": "Summary",
                "type": "string"
              }
            },
            "required": [
              "items",
              "summary"
            ],
            "title": "Checklist",
            "type": "object"
          },
          "strict": true
        }
      }
    }
    Response 200
    {
      "choices": [
        {
          "finish_reason": "stop",
          "index": 0,
          "message": {
            "content": "{\"items\": [{\"check\": \"helmets\", \"result\": \"fail\", \"evidence\": \"The worker is not wearing a helmet.\"}, {\"check\": \"safety_vests\", \"result\": \"pass\", \"evidence\": \"The worker is wearing an orange safety vest.\"}, {\"check\": \"exits_clear\", \"result\": \"fail\", \"evidence\": \"The fire exit path is blocked by several cardboard boxes.\"}, {\"check\": \"no_spills\", \"result\": \"pass\", \"evidence\": \"The floor appears dry and clean, with no visible spills.\"}, {\"check\": \"cables_secured\", \"result\": \"pass\", \"evidence\": \"No loose cables are visible.\"}], \"summary\": \"Safety violations identified include a missing helmet for the worker and an obstructed fire exit; housekeeping standards for spills and cables are met.\"}",
            "role": "assistant"
          }
        }
      ],
      "object": "chat.completion",
      "usage": {
        "completion_tokens": 174,
        "prompt_tokens": 444,
        "total_tokens": 618
      }
    }

What the program printed

Raise a ticket: helmets | The worker is not wearing a helmet.
Raise a ticket: exits_clear | The fire exit path is blocked by several cardboard boxes.
Safety violations identified include a missing helmet for the worker and an obstructed fire exit; housekeeping standards for spills and cables are met.
  1. See, then sort. /chat watches and listens. A second, cheaper call on Fast turns the description into fixed checklist values.
  2. “Not visible” is an allowed answer. The agent doesn’t have to guess about things the camera missed.
  3. Your code raises the tickets. The model reports; your program decides who is told.

POST /chat answers 400 with a fixed error_code, so your code can tell the person what to do:

error_code What it means What to do
video_too_large The file is over 100 MB. Cut it into shorter clips, or record at a lower quality.
video_too_long The video is over 5 minutes. Cut it into clips of a minute or two.
video_unreadable The file is damaged or not really a video. Check that it plays on your computer, then send it again.
video_unavailable This server can’t read videos at the moment. Ask your ZenithAI administrator to check video support.

Checking that a product demo video shows the right steps, reviewing a short training clip, describing CCTV clips for an incident report, or checking that a delivery video shows the seal intact.

Common questions

How long a video can ZenithAI watch?

A video can be up to 5 minutes and 100 MB. ZenithAI looks at one frame per second, up to 60 frames, so it sees about the first minute of the picture. For longer recordings, cut the video into one-minute clips in your own code and send them one by one.

Can ZenithAI read a video recorded on a phone?

Yes. Send the file just as the phone saved it. Android phones usually save H.264 MP4 and iPhones save H.265 MOV, and ZenithAI reads both. There is no need to convert the video first.

Does ZenithAI hear what people say in a video?

Yes, when the video has a soundtrack. The speech is turned into text and used with the frames, so an answer can cover both what was shown and what was said.