Perslis Accessibility
03 / SEE TO SPEAK

Point the camera. Get words about that thing.

A photograph becomes communication choices about what is actually in it — not a generic word list. Images produce proposed communication, never delivered speech.

THREE KINDS OF LOOKING

The runtime tells them apart.

/see picks object, menu or scene automatically; you can also name the mode.

InputWhat comes back
An object — shoes, gloves, foodObject names, body-use associations where relevant, and choices for requests, help, actions and refusal.
A menu photographThe item names and prices it reads off the page, with categories and ordering sentences. Several pages stay in one catalog, and a misread is surfaced rather than hidden.
A sceneVisible objects, optional image rectangles and spatial relations, plus choices relevant to that scene.
/object "/path/to/shoes.jpg"
/menu "/path/to/menu-1.jpg" "/path/to/menu-2.jpg"
/scene "/path/to/room.jpg"
/see "/path/to/photo.jpg"
THE CONVERSATION CONTINUES

Seeing is not a dead end.

After a scan, ordinary partner input keeps using the vision model together with the saved observations: shoes → “I need my shoes” → “Do you need help?” → relevant follow-up choices. No new image is sent, and the runtime does not claim to observe a changed scene. /unsee returns to the text model and keeps the conversation.

WHAT THE DATA PROMISES

Observations, with their limits stated.

01

Boxes are fractions

{x,y,w,h} of image width and height, top-left origin. Model observations only — no depth estimate and no navigation guarantee.

02

Instances stay distinct

Relations expose indices into the object list, so two cups are two cups. An ambiguous group name yields a null index rather than a guess.

03

Prices are display data

Separate from the proposed sentence. Selecting an ordering tile speaks the user's words; it never submits a purchase or payment.

04

Wrong is reported, not hidden

Recognition and OCR can be wrong. Warnings and the original item text are exposed for your review interface. An incomplete result fails explicitly rather than substituting a fabricated menu.

05

Limits are explicit

Up to eight images per scan, 5 MiB each and 12 MiB combined. Sentences are never truncated; replies that break the word limit are reported as excluded.

06

Photos are not kept

Image bytes exist only during the request. Session snapshots keep hashes, type, size and observations — never the image, a path or base64.

Offline and images are different questions

The text-only child model does not gain image recognition by configuration. A local Ollama model can process images on your host, but it uses HTTP, so --offline disables that connection too. A loopback endpoint alone does not prove offline inference.

ContinueModels & privacy