You do not install our runtime. You call it.
Your app talks to the TinkySpeak API with a key we issue. You never hold our code, you never ship a model, and your symbols, your voice and your interface stay entirely yours.
A pilot, issued by hand.
There is no self-serve key console today. You tell us what you are building, we issue a key and we work through your deployment with you. If you would rather run the engine inside your own infrastructure instead of calling ours, say so when you write — that is a conversation about terms, not a download link.
Six steps, start to finish.
| Step | What you do |
|---|---|
| 1 Ask for a key | Tell us what you are building and who it is for. We issue a key for the pilot and talk to you about your deployment — this is not a self-serve console yet. |
| 2 Keep the key on your server | Put it in TINKYSPEAK_API_TOKEN and send it as Authorization: Bearer …. Never ship a key inside an app binary: anyone can read it out. Your backend owns accounts and session ownership when you serve more than one person. |
| 3 Open a session per conversation | A session carries the languages, the communication profile and the memory. Keep its id for as long as the conversation lasts; delete it when you are done. |
| 4 Run the five-step loop | hear → render → select → speak → confirm. That is the whole conversation surface, and it is the same whether a person is choosing by touch, switch, or eye gaze. |
| 5 Point your own symbols at it | The runtime never fetches or draws artwork. You map a choice to your library and return a descriptor; unmatched choices keep their emoji. |
| 6 Add the camera when you want it | Send frames from your own camera and the same tiles come back about what is in front of the person — a menu, a shelf, an object. |
Where it lives, and where it must not.
# Your server, not the device in someone's hands.
export TINKYSPEAK_API_TOKEN="the key we issue you"A key identifies your application, not a person. Requests from a browser must come from an origin you listed; a non-loopback address requires a token. No conversation request bodies are logged.
One per conversation.
POST /v1/sessions
Authorization: Bearer $TINKYSPEAK_API_TOKEN
Content-Type: application/json
{ "language": "en", "partnerLanguage": "en",
"profile": { "age": 7, "level": "sentence", "choices": 6 } }
→ { "id": "SESSION_ID", "choices": [], ... }Five calls, and your app owns the middle of it.
// 1 — what the other person said
const state = await api.hear({ text: 'Would you like tea or coffee?' });
renderYourTiles(state.choices); // six complete sentences
// 2 — the person picks a tile in YOUR interface
const selected = await api.select({ choiceId: tappedTileId });
// 3 — turn the draft into a delivery request
const pending = await api.speak({ draftId: selected.draft.id });
// 4 — YOUR speech engine says it, in your voice, on your device
await yourSpeechEngine(pending.speech.text, pending.speech.language);
// 5 — tell the runtime what actually happened
await api.confirmSpeech({ speechId: pending.speech.id, outcome: 'spoken' });Each id is single-use, so a stale tap cannot deliver yesterday's sentence. Concurrent actions return 409 session_busy. If delivery fails, confirm with failed or cancelled — never blindly retry.
The words are ours to propose. The pictures are yours.
A choice is words plus an identity. What it looks like is a decision only your app can make, because only your app knows the person, their vocabulary and the artwork they already recognise.
// Every choice arrives with an id, a label, a sentence and an emoji.
// You decide what it looks like.
function yourArtworkFor(choice) {
const hit = myLibrary.lookup(choice.label, choice.sentence);
if (!hit) return null; // null keeps the emoji fallback
return { kind: 'symbol', library: 'my-aac-library', id: hit.id, alt: hit.alt };
// or: { kind: 'image', uri: '/my-art/water.png', alt: 'My water cup' }
}Return null and the emoji stays. Matching a whole generated sentence is not a universal symbol dictionary — provide a language-aware lookup and let it fail honestly rather than showing the wrong picture. Changing artwork never changes the words that get spoken.
Give it eyes when you want them.
Vision is configured separately from the conversation model, so you can run a camera feature without changing how the words are proposed.
// A photo from YOUR camera or photo picker, as base64.
const state = await api.scan({
images: [{ mimeType: 'image/jpeg', data: base64Frame }],
mode: 'menu', // auto | object | menu | scene
});
renderYourTiles(state.choices); // tiles about what is in the picture
state.vision.categories; // the whole catalog, paged
state.vision.objects; // { name, box, bodyPart }Up to eight images per scan, 5 MiB each. Boxes are fractions of the image, top-left origin. Prices are display data, never a purchase. Recognition can be wrong, so warnings and the original text come back for your review interface, and an incomplete read fails rather than inventing a menu.