Boxes are fractions
{x,y,w,h} of image width and height, top-left origin. Model observations only — no depth estimate and no navigation guarantee.
A photograph becomes communication choices about what is actually in it — not a generic word list. Images produce proposed communication, never delivered speech.
/see picks object, menu or scene automatically; you can also name the mode.
| Input | What comes back |
|---|---|
| An object — shoes, gloves, food | Object names, body-use associations where relevant, and choices for requests, help, actions and refusal. |
| A menu photograph | The item names and prices it reads off the page, with categories and ordering sentences. Several pages stay in one catalog, and a misread is surfaced rather than hidden. |
| A scene | Visible objects, optional image rectangles and spatial relations, plus choices relevant to that scene. |
/object "/path/to/shoes.jpg"
/menu "/path/to/menu-1.jpg" "/path/to/menu-2.jpg"
/scene "/path/to/room.jpg"
/see "/path/to/photo.jpg"After a scan, ordinary partner input keeps using the vision model together with the saved observations: shoes → “I need my shoes” → “Do you need help?” → relevant follow-up choices. No new image is sent, and the runtime does not claim to observe a changed scene. /unsee returns to the text model and keeps the conversation.
{x,y,w,h} of image width and height, top-left origin. Model observations only — no depth estimate and no navigation guarantee.
Relations expose indices into the object list, so two cups are two cups. An ambiguous group name yields a null index rather than a guess.
Separate from the proposed sentence. Selecting an ordering tile speaks the user's words; it never submits a purchase or payment.
Recognition and OCR can be wrong. Warnings and the original item text are exposed for your review interface. An incomplete result fails explicitly rather than substituting a fabricated menu.
Up to eight images per scan, 5 MiB each and 12 MiB combined. Sentences are never truncated; replies that break the word limit are reported as excluded.
Image bytes exist only during the request. Session snapshots keep hashes, type, size and observations — never the image, a path or base64.
The text-only child model does not gain image recognition by configuration. A local Ollama model can process images on your host, but it uses HTTP, so --offline disables that connection too. A loopback endpoint alone does not prove offline inference.