Speech to text, live
Real-time captions or sentence-by-sentence capture, straight from your microphone with per-result confidence.
Open RecogniseRecognition, voice synthesis, and a lip-synced avatar, tuned for Australian customers, sovereign workloads, and the vocabulary generic engines stumble on: Centrelink, Medicare, Nguyen, Surry Hills, BPay, NBN, rego, and every reference number in between.

Real-time captions or sentence-by-sentence capture, straight from your microphone with per-result confidence.
Open RecogniseHear Camlin Voice, an Australian studio voice, on contact-centre lines — ATO, rego, Medicare, callbacks — with lip-sync for avatars.
Open SpeakWatch a 3D avatar speak your line in sync — every clip ships with the lip-sync data that drives it.
Open AvatarREST reference for recognition, synthesis and the avatar track, with curl and JavaScript you can paste straight in.
Open DocsAvatar
Every clip Camlin Speech synthesises can carry a second track: frame-by-frame mouth positions in the standard ARKit blendshape format, timed to the audio — through pauses, spelled-out reference numbers, all of it. Point that track at an avatar and it speaks your words on screen: web self-service with a person to look at, kiosks, and agent-assist views where a caller sees who’s talking.
Use it with your own rig — anything that takes ARKit blendshapes works — or with the hosted 3D avatar from Camlin Connect, the platform Camlin Speech is part of. The demo drives that avatar live, in your browser, from the same API response you’d get.
What’s in the track
For developers
cs_live_ for production, cs_test_ for sandboxes and pilots. Each key carries only the scopes you give it — stt, stt:stream, tts, viseme, lexicon — with its own rate limit.
Authorization: Bearer cs_live_…POST a WAV and get a JSON transcript, or open one Socket.IO session and receive words as they commit, tentative words marked, with per-result confidence.
POST /api/v1/asr/recognizePut your names, agencies, suburbs and product terms in a tenant lexicon over the API. It applies to every keyed recognition call from then on — no client change, no retraining wait.
PUT /api/v1/asr/lexicon/{tenant}Camlin Voice returns WAV audio and, with the viseme scope, the ARKit-52 blendshape track aligned to it — pauses and spelled-out codes included. One request, both tracks.
POST /api/v1/tts/synthesizeFine-tuned for Australian English, with lexicon biasing that carries the everyday vocabulary — Centrelink, Medicare, NBN, rego — and the surnames generic engines mishear.
Audio is processed in our Sydney region (AWS ap-southeast-2). Built for Australian government, healthcare, telco and finance workloads that have to stay onshore.
Audio you send and transcripts we return are not stored unless your tenant opts in. Request metadata — route, timing, audio seconds, key id — is kept 90 days, then purged.
The same stack can run in our Sydney cloud, in a dedicated AWS account, or on infrastructure you control. Ask us about the fit.
Details, including what is logged per request and for how long, are in the data retention section of the docs and in our privacy policy.
Who builds it
Camlin Speech is built by Nuamedia, which has built IVR and service-automation platforms since 2011 and has developed Camlin Connect since 2012; the redesigned Camlin Connect v1.0 shipped in February 2026. The recogniser and the voice were shaped on the calls those platforms take — the agencies, surnames, suburbs and reference numbers that turn up in Australian contact centres.
How to judge it
We do not publish an accuracy table. Read-speech benchmarks say little about noisy phone lines, and a number we picked would tell you even less. Run the demos on your own voice, then send us the audio your current engine gets wrong and we will return a side-by-side on it.
Ask for access or a trial, send us your hardest audio to test against, or just start a conversation.
Microphone is only active while you’re recording. Audio is processed in our Sydney region.