Voice V1 is frozen: the next milestone is conversational presence

Updated 26 August 2026prototypeBy Catalin Adelin Iovan Raducu

The Orin-generation experimental Voice V1 was frozen after representative clean physical turns reached about 2.15 seconds from user end-of-speech to first meaningful audio, with a fastest retained turn near 1.51 seconds. This remains historical Orin evidence, not the current Thor runtime.

PROGRESS UPDATEThe Orin Nano implementation is now preserved as historical evidence. Active development has moved to a clean-room Jetson AGX Thor baseline with new verification and identity gates.
Completed
At that checkpoint, the bilingual voice baseline combined persistent six-thread Parakeet ASR, Qwen3.5-4B, selective-INT8 Piper speech, XVF3800 hardware echo cancellation, hardware-AEC-assisted barge-in, qualified endpointing, 233 passing tests plus two subtests, and a tagged recovery anchor.
Known limit
About two seconds is a representative result, not a guaranteed ceiling. Conversational Presence R0 is still research: no vocal backchannel, visual nod, reliable spontaneous end-of-turn policy, robust multi-speaker ownership, governed persistent memory, or physical actuation is claimed yet.
Record type
Engineering milestone
See the current development state

LAB RECORD / J-009 / 17-20 AUGUST 2026

Evopien now has a voice baseline worth freezing.

This record covers the engineering work after the latest website commit on 17 August 2026 at 16:22 CEST. The Orin voice path at that checkpoint reached a representative clean physical response of about 2.15 seconds from user end-of-speech to first meaningful audio. The fastest observed turn in the retained evidence was about 1.51 seconds.

Those numbers must not be collapsed into a guaranteed 1.5-second claim. Healthy physical turns still vary, and the frozen master checkpoint describes an ordinary clean range of roughly 2.0 to 2.7 seconds. The honest public milestone is therefore about two seconds, with 1.51 seconds recorded as a best observed turn rather than the normal ceiling.

Current result

What is frozen, active and still open.

Voice owns acoustic input and output. Conversational Presence is a separate active subsystem that will decide social timing without changing the frozen voice baseline.

ItemStateEvidence
Voice V1Legacy validated

Local bilingual simultaneous capture/playback, endpointing, meaningful first audio, streaming speech, and hardware-AEC-assisted barge-in formed the frozen Orin experimental baseline.

First meaningful audioAbout 2.15 s

Representative clean physical result after user end-of-speech. The fastest observed retained turn was about 1.51 s; this is not a guarantee.

Speech recognitionPersistent

Parakeet TDT 0.6B v3 Q4_K runs as a persistent six-thread CPU worker; same-WAV median was about 0.297 s.

Speech synthesisFrozen

Selective-INT8 Piper uses Amy Medium for English and Claude High for Spanish, about 166 MB RSS and roughly 0.20 s synthesis for typical short opening pieces.

Conversation modelWorking

Qwen3.5-4B UD-Q3_K_XL runs through llama.cpp with reasoning off, 4096 voice context, Q8 K/V and mandatory prompt-cache warm-up.

Conversational Presence R0Active research

Turn state, thinking pauses, floor handoff and acknowledgement are being evaluated separately from Voice V1.

Vocal acknowledgementNot active

No mhm, mm-hm or uh-huh backchannel was enabled in the Orin runtime.

Multi-speaker ownershipNot solved

Direction of arrival remains observe-only and no robust family-room floor owner is claimed.

Governed memory / bodyNot implemented

Selective governed memory, the complete multi-user privacy proof, face or head actuation and physical movement remain future milestones.

Engineering sequence

What changed after J-008.

The work was not one model swap. It was a measured narrowing of the complete response path until the remaining latency no longer justified destabilising the subsystem.

  1. The response boundary was defined honestly

    Timing is measured from physical user end-of-speech to the first acoustic answer that carries real meaning. Filler such as wait, let me think or a synthetic acknowledgement added only to stop the clock does not count.

  2. Parakeet became a persistent six-thread service

    Q4_K CPU recognition remained inside the 8 GB memory envelope while same-WAV median time fell from about 0.764 s at two threads to about 0.297 s at six. GPU ASR remained rejected because its unified-memory risk was not justified.

  3. The endpoint stopped letting random noise own the clock

    The normal pause remained about 0.60 s, but the capture path updated its final speech time only for qualified speech after the turn opened. Environmental noise no longer extended a turn merely because it crossed a raw energy threshold.

  4. English and Spanish TTS candidates were tested, then narrowed

    Scylla, Pocket and other alternatives exposed useful speed and quality trade-offs, but no mixed stack justified continuing model churn. Selective-INT8 Piper with Amy Medium and Claude High became the final bilingual Voice V1 baseline.

  5. Qwen was profiled beyond headline tokens per second

    Prompt cache removed roughly 2.6 to 2.8 seconds from cold prompt evaluation. The resident multimodal projector added about 55 MB PSS and essentially no measured text latency. Batch size, K/V type, pinned clocks and context depth produced only small gains.

  6. Latency shortcuts were rejected when they damaged the product

    Speculative rolling paths reached roughly 1.26 to 1.68 s but were fragile. Forced ultra-short first sentences could reach roughly 1.11 to 1.64 s but degraded answer quality and gamed the metric. Neither entered the frozen Voice V1 baseline.

  7. The final physical voice result became good enough

    Representative recent live turns clustered around 1.95, 2.12 and 2.18 seconds, averaging about 2.08 seconds. The frozen master checkpoint records a representative median near 2.15 seconds and an ordinary clean range around 2.0 to 2.7 seconds.

  8. Voice V1 was tagged and frozen

    The suite reached 233 passing tests plus two subtests. Bash syntax, Python compilation and whitespace validation passed. Further voice tuning now requires a concrete physical-prototype failure or a changed product requirement.

  9. Conversational Presence became the active branch

    The next problem is not simply speaking earlier. It is deciding whether silence means finished, thinking or continuing, and whether a subtle acknowledgement can show attention without taking the floor.

Frozen Orin baseline

Voice V1 on the 8 GB Jetson.

These were the selected choices at that checkpoint. Earlier J-008 details remain historically correct for that checkpoint but no longer describe the active TTS or ASR lifecycle.

Compute
Jetson Orin Nano Super 8 GB · MAXN_SUPER · Ubuntu 24.04.4 · JetPack 7.2.1-b49 · CUDA 13.2
LLM / VLM
Qwen3.5-4B UD-Q3_K_XL · llama.cpp · GPU offload · 4096 voice context · Q8 K/V · reasoning off · multimodal projector resident
ASR
Parakeet TDT 0.6B v3 Q4_K · persistent native CPU worker · six threads · RIGHT XVF channel
TTS
Piper selective INT8 · Amy Medium English · Claude High Spanish · first piece near 48 characters plus slack
Audio
XVF3800 hardware AEC · LEFT speech gate · RIGHT ASR · Pebble V3 playback · hardware-AEC-assisted barge-in and stream cancellation
Endpoint
Approximately 0.60 s normal pause · qualified-speech timing · no random-noise extension
Default vision
Fresh C920 frame on demand · resident multimodal projector · continuous RF-DETR/NvDCF off
Preserved vision
DeepStream 9.1 · RF-DETR Medium 576 FP16 · threshold 0.40 · NvDCF tracking · separate validated mode

Measured timing

The two-second milestone, without hiding its variance.

Component timings explain the architecture, but only physical end-to-end timing supports the user-facing response claim.

ItemStateEvidence
Fastest retained physical turn1.510 s

Best observed first meaningful audio in the retained 18-turn evidence; not the representative result.

Recent representative examples2.08 s average

About 1.95 s, 2.12 s and 2.18 s after the stronger direct-answer prompt and final voice path.

Frozen representative medianAbout 2.15 s

Clean serial physical end-of-speech to first meaningful audio; ordinary healthy turns were roughly 2.0 to 2.7 s.

Pre-Qwen path0.76-0.90 s

Endpoint confirmation, WAV finalisation and ASR consume most of the time before Qwen can begin.

Warm Qwen first sentenceAbout 0.78 s

Representative warm first-sentence median from the retained Qwen profile, including first-token time.

Piper opening synthesisAbout 0.20 s

Typical short opening piece in the frozen selective-INT8 Orin path.

Engineering decision

A lower benchmark number is no longer the highest-value task.

Q4_K_M decoded about 11.6% faster than Q3 in isolation, but the expected opening-audio saving was only tens of milliseconds. Larger batches saved about nine milliseconds. Pinned Jetson clocks added heat for about a quarter-percent decode gain. None changes the interaction as much as correct turn timing.

Sub-second meaningful speech would require safe overlap between live ASR, end-of-turn confidence, speculative Qwen and TTS. That remains a legitimate future architecture, but it is not justified while Voice V1 already meets the current prototype requirement and major robot systems remain unfinished.

Active subsystem

Presence owns social timing, not the answer.

The objective is for Evopien to feel genuinely there: listening, waiting when appropriate, acknowledging carefully and responding at the socially correct moment.

ItemStateEvidence
CONTINUEHold floor

The user is still speaking or has clearly left the sentence open. Evopien remains silent.

THINKINGWait attentively

Silence is not automatically a handoff. The user may still own the conversational floor.

BACKCHANNEL WINDOWAcknowledge carefully

A future cached mhm or subtle visual nod may signal attention without agreeing with the content or ending the turn.

ENDTake floor

The user has handed over the turn and frozen Voice V1 may generate the actual response.

Inbound backchannelLater phase

A human mhm, yeah or okay while Evopien speaks should not always trigger barge-in.

Turn-detector checkpoint

Two small local detectors fit. Neither has earned floor ownership.

Smart Turn v3.2 and LiveKit v1-mini both deploy on the Jetson with acceptable inference cost. Smart Turn used about 176 MB peak RSS and later ran near 88 ms median. LiveKit used about 359 MB peak RSS and ran near 85 ms median. Compute feasibility passed for both.

On a deliberately difficult guided stress corpus at the 600 ms decision point, Smart Turn produced 21 false interruptions across 23 HOLD events; LiveKit produced 23 across 23. A threshold strict enough to remove false interruptions caused 39 unnecessary waits across 41 END events for either detector.

Next controlled sequence

Prove real turn timing once, then choose and freeze.

  1. 01Record roughly 25 to 40 spontaneous English, Spanish and naturally code-switched answers without scripted sentences, instructed pauses or requested prosody.
  2. 02Use the then-current Silero segmentation and audit only obviously suspicious long gaps; do not treat every VAD boundary as conversational truth.
  3. 03Run Smart Turn and LiveKit on identical events and perform one threshold-separation analysis covering false interruptions and unnecessary waits.
  4. 04If one detector is good enough, integrate it first in shadow mode, test it physically and freeze Presence R0.
  5. 06Enable visual listening cues before vocal backchannels when head hardware exists; they can show attention without semantic endorsement or speech interference.
  6. 07Only then enable sparse cached noncommittal acknowledgements such as mhm, with cooldown, cancellation and frozen Voice V1 retaining priority.
  7. 08After Presence is comfortable on the physical head, freeze it and move to governed memory, privacy, face and gaze, ESP32 integration and embodiment.

Open the evidence-gated roadmap for the boundary between the current Presence phase, the governed Core milestone and future embodiment.

Evidence boundary

What this milestone does not claim.

ItemStateEvidence
Guaranteed 1.5 s responseNot claimed

1.51 s is the fastest retained observation. Representative clean operation is about 2.15 s and harder turns can take longer.

Human-like PresenceNot established

The architecture and benchmark path were defined, but backchannels and floor-control influence were not enabled in the Orin runtime.

Content understandingBounded

A neutral acknowledgement must never be presented as agreement, factual verification or safety approval.

Memory or motor authorityNone

Presence cannot fabricate user turns, write long-term memory, bypass governance or issue raw motor commands.

Complete robotNot yet

Governance, selective memory, household privacy, face and gaze, movement, manipulation and real-world embodiment remain incomplete.

Return to the complete journal