5–9 September / J-014
The first integrated foundation is now worth preserving.
Since the September 5 checkpoint, 424 engineering commits have brought conversation, vision and internet research into one working demonstration platform.
I have frozen that baseline and brought it into main. It is the starting point for governed memory and stronger permission boundaries. Useful integration has been achieved; formal acceptance and the remaining defects stay visible.

What changed
Four parts of one working foundation.
5–6 September / Voice and vision
From always-attached video to the right evidence for the question.
J-013 ended with a candidate that attached a selected view of the previous 15 seconds to every ordinary Qwen turn. The following CP08–CP11 work separated live evidence from recent-history evidence, raised live-image quality, expanded English/Spain-Spanish controls and tightened cleanup around the process that actually owned the resources. Ordinary Qwen turns still received ambient live video in CP11. CP15 later introduced the shared text/live/history split, keeping fresh visual payloads out of text-default requests.
That distinction matters. A question about the room now needs fresh evidence. A question about an action ten seconds ago needs a timestamped sequence. Ordinary conversation should not pay for unnecessary visual payloads. It also must not quietly treat a previous model description as a fresh camera observation.
CP12 concentrated on the latency path: the existing cache16 configuration, playback ownership, cancellation, replay and measurement. CP13 added semantic retrieval over the open conversation while keeping Qwen’s working context bounded. Larger prefill and speculative decoding candidates did not meet their advancement conditions; the retained cache16 reference was not replaced simply because an isolated throughput number improved. Replay-source preservation and explicit abort completion addressed conversation lifecycle failures. The local reference was accepted and frozen at f42bb577 before internet integration.
The CP13 retrieval fixture used the real multilingual E5-small encoder over 2,018 synthetic fragments and 51 simulated epochs: ten of ten retrieval checks passed, with a maximum of 63.253 ms. Selected recall is bounded to twelve exchanges and 2,000 characters within the 9,000-character context budget. That tests retrieval, not answer accuracy.
The retained local physical run completed twenty turns and cancelled one, with six context rollovers and three language transitions. Live-vision nearest-rank P50/P95 was 3.182/3.946 seconds in English (nine turns) and 3.380/3.969 seconds in Spain Spanish (seven turns), from speech end to first PCM. The four history turns formed smaller separate cohorts. Correction recall and untagged reasoning reaching speech remained recorded failures. These local results use a different fixture from the later cloud demo.
This session recall is a derived index over the spoken conversation, under Core ownership. It is not governed durable memory, an independent autobiography or permission to write permanent memories. Interrupted or merely scheduled speech must not become something Evopien assumes the person heard.
CP14 and CP15 then introduced a separate optional cloud comparison. The retained CP15 detail12 reference is 1bdc32ca. Groq replaces selected dialogue and visual reasoning while the local speech, camera, perception and Core owners remain. These experiments do not establish a controlled cloud-versus-local speedup: the runs and prompts differ.
6–9 September / Internet
A search result is only the beginning of an answer.
The first local internet work added typed retrieval, provider-specific readers and broader fixtures. The difficult part was the complete chain: admitting a legitimate public question, finding usable sources, opening the right pages, checking the evidence and releasing a useful spoken answer within a practical deadline.
The first 100-case retrieval-only assessment returned some evidence for 65 questions and none for 35. That measured availability, not the quality of final answers. Structured and keyless routes could return useful information, while generic search, page reading and provider/key availability remained inconsistent. Wikipedia route and retrieval fixtures improved, but natural follow-up questions still failed, motivating a more semantic Qwen-controlled research loop.
The experiments exposed empty results, stale-only evidence, unavailable providers, public tasks rejected by Core, grounding failures and timeouts. Direct and efficient research paths were compared with a 100-question harness. Compact structured decisions, scoped candidate identifiers and phase diagnostics made the failures easier to locate. The existence of 100 questions is not a 100-question pass.
An isolated Open WebUI reference then provided another way to test the same user experience. It progressed through bounded search/read workflows, conversational routing, structured decisions and the James Jones DuckDuckGo tool. Streamed Linux terminal access made the answer and its progress visible outside the browser, ready for the existing ASR/TTS path.
On September 9, the James retrieval worker was integrated directly with the local conversation runtime. Core admits a public question, an isolated worker returns evidence, and local Qwen answers through the established speech queue. The final integrated profile does not require an Open WebUI server or a second Open WebUI answer-generation pass.
Qwen, speech and Core can run on Thor. James/DDG public research still contacts external websites and search services. The optional cloud profile additionally sends selected requests to Groq or Exa.
The failed experiments stay in the record.
A headline-only shortcut returned off-topic or empty news and was rolled back to preserve page evidence. A later recovery candidate reduced execution checks from 91 to 68 and fresh-evidence cases from 23/30 to 11/30; the preceding tree was restored. These counts describe execution and evidence availability, not factual-accuracy percentages.
Another mandatory news-date screen discarded two real eight-source responses, leaving no matching source date. That news change was rolled back while speech repairs for times, currencies and domains were retained. Current-news relevance, publication dates and source freshness therefore remain open work.
The first integrated local run also exposed a WAIT/CONTINUE stall. Replay ownership and supervised resumed speech were repaired. Public-query admission, fresh device-clock answers, unsupported search claims and spoken formatting received additional work. Regressions were not promoted just because they were newer.
Two different Exa paths
Retrieving evidence and generating the research answer are separate choices.
The optional Exa MCP search path first returned evidence for local Qwen to synthesize. One weather trial took 15.258 seconds from speech end to first PCM: 4.262 seconds in retrieval and 9.180 seconds from the instrumented local-model request to its first token. This exposed the bottleneck in that trial; it is not a universal local-model timing.
The next path used Exa /answer to generate the cited public answer itself. Two standalone text trials completed in 1.729 and 1.689 seconds. Those were text measurements, not voice timings. The integration then streamed the cited answer through Core’s existing normalization, queue, ledger and cancellation path without another local Qwen rewrite.
Exa receives the admitted public query and bounded date/language controls. The implemented request boundary excludes camera media, audio, private recall and identity state. Speech waits for citations and a usable text boundary; automatic billable retry is disabled. Tests exercise this boundary, but the demo does not prove complete privacy or network security.
The frozen architecture
Two profiles. One local conversation lifecycle.
| Function | Local + James | Cloud |
|---|---|---|
| Dialogue | Local Qwen3.8 through the existing SGLang/cache16 runtime | Groq; the retained log identifies qwen/qwen3.8-27b |
| Selected vision | Native selected live video or recent-history video | Fresh images or timestamped detail12 contact sheets sent to Groq |
| Public research | Isolated James/DDG retrieval; local Qwen writes the answer | Isolated Exa /answer streams a cited answer without a local rewrite |
| Speech and turns | Local Cohere Transcribe, Smart Turn/Core policy, Scylla CPU speech | The same local speech and turn owners |
| Camera and perception | One C920 owner; local DeepStream/TensorRT detection and tracking | The same local acquisition and perception |
| Identity and session | Core, permissions, spoken ledger and temporary recall remain on Thor | Canonical state stays local; selected bounded context can reach Groq |
Canonical identity, permissions and the authoritative ledger remain on Thor. The cloud profile sends selected bounded conversation context and visual evidence to Groq; it is not an all-local privacy mode. Exa has the narrower public-query boundary described above.
These provider names describe the reviewed configuration and log labels. The actual speech recognizer is local resident Cohere Transcribe, operating on utterances; synthesis uses local Scylla INT8 on CPU, the ink voice and 24 kHz output. Older Riva, Parakeet and Magpie status banners do not identify the active backend.
The reSpeaker XVF3800 remains in the playback path so its hardware echo-cancellation reference is preserved. The C920 is acquired once at configured 1280×720/30 FPS. The final run uses RT-DETR R50vd COCO/O365 FP16 detection with the local tracking path; earlier RF-DETR references describe a different configuration.
Live detail and recent history have different budgets.
The temporary camera history now holds 15 seconds at eight samples per second, up to 121 compressed JPEGs in RAM. Local live vision selects eight frames from approximately the latest second at 960×540; local history selects sixteen frames at 640×360.
Cloud detail12 selects one fresh original 720p image for static questions, three for live motion, or up to twelve timestamped history samples in three 2×2 contact sheets. History sheets are prepared off the request path. Text-default requests should carry no new visual payload. Temporary processing is separate from permission to retain permanent visual memories.
9 September / One demo interface
The console exposes the runtime that actually runs.
A shared local console and terminal launcher now offer the local and cloud profiles, typed or voice input, English/Spain-Spanish selection, Start/Stop, stop-speaking and software microphone mute. The console shows runtime stages, software latency, the shared camera preview and the actual visual evidence selected for a model request.
The on-screen avatar reacts to playback amplitude and current person cues. It is an illustrative interface, not a physical head, phoneme-level lip sync, recognition or proof of emotional understanding.
Bounded text/JSON diagnostics can be exported after Stop. Provider keys are not given to the browser; exports exclude credentials and raw audio/video, but conversation text remains sensitive. Hiding a camera preview does not stop capture, and software mute is not a physical privacy disconnect. The session stays in RAM until the next Start or console exit.
Retained physical evidence / 9 September
What the final cloud demonstration exercised.
The recorded session lasted about 8 minutes 39 seconds including startup and shutdown. It began in English and ended in Spain Spanish. Eighteen Groq requests returned HTTP 200; eight Exa answer routes returned sources. The run recorded fifteen confirmed interruptions and one accepted WAIT/CONTINUE pause/resume with four replay units.
Local perception processed 14,172 frames at a reported 28.996 FPS, with no invalid tracking IDs, boxes or metadata. Seventy Cohere requests reported no failures. The session ended normally with exit code 0, no forced stop, its precomputation fence passed and retained media at zero.
Those are bounded Thor-verified execution observations. A completed request does not prove that its answer was correct or fully heard. The final export does not repeat the latest local Qwen + James console qualification; that profile is implemented and inherits earlier local evidence.
| Route | Turns | P50 | Observed range |
|---|---|---|---|
| Cloud dialogue | 5 | 1.4203 s | 1.1916–1.6060 s |
| Cloud live vision | 4 | 1.7264 s | 1.7090–2.0094 s |
| Cloud recent history | 2 | 2.0818 s | 2.0818–3.1339 s |
| Exa research answers | 8 | 3.0035 s | 2.7965–4.8142 s |
Timing starts at detected speech end and ends at first available PCM in software, before device write. Cohorts include completed transcript entries with a latency object; P50 uses nearest rank, the lower middle observation for an even sample count. These small descriptive cohorts are not accepted success distributions, acoustic measurements or a matched local/cloud benchmark.
All 27 transcript latency objects remained qualified=false. The inherited sustained time-to-first-token and ≤1.3-second P95 first-PCM screens failed. Those screens mix routes in this report and need careful interpretation, but no threshold was changed to turn the freeze into a pass. The reporter completing successfully is not the system passing its latency gates.
What is still wrong
The freeze preserves the bugs as well as the useful baseline.
- A Spanish recent-history request routed to text without images and repeated an older scene. Another camera question needed rephrasing before reaching live vision.
- Legitimate encyclopedia and currency questions were rejected as private or unresolved public queries. Rephrasing reached research.
- Returned citations still need factual and publication/freshness review. Having sources is not evidence that every claim is supported.
- Some transcript statuses remained pending or in progress in the finalized export; playback labels and completion flags do not prove complete audible delivery.
- Technical-version speech can still merge digits incorrectly. Interrupted Exa answers followed by replay and Groq continuation need clearer provenance so a changed provider cannot silently inherit a source claim.
- The latest local console, bilingual routing coverage, reliable visual truth, interruption quality and latency tails still need controlled qualification.
MS-07, MS-08 and MS-09 remain open. This is a founder-accepted working demo freeze, not Alpha, validated household use, certification or permission for powered motion.
The next development phase
Decide what deserves to become memory.
The integration is now useful enough to hold steady while the next layer is built. The work starts with design and synthetic fixtures: selective retention, provenance, person and household scope, explicit permission, inspection, correction, deletion and audit.
Core owns memory decisions. Qwen, Groq, voice, vision and the console are replaceable interfaces and providers. A generated answer, stale visual description, interrupted sentence or external summary cannot become an authorized permanent fact simply because it appeared in a conversation.
The Thor identity was already created once at MS-06 with disclosed engineering lineage. This phase does not create it again or import the predecessor’s memories. Approved Foundation v1.2 remains controlling; the later Foundation and Alpha drafts do not silently supersede it.
Head hardware follows through separate motion contracts, local control and independent safety evidence. Memory work, real participants and powered embodiment each keep their applicable gates. The immediate task is to preserve the useful demo, fix its known regressions and make the next memory decisions defensible.
Source boundary and engineering references
J-014 compares website v1.7.0 at 694390dad9136142f107b989bc0a75b1dfd88f32 with engineering work after J-013’s cutoff 443db7f53e623983842917568ab25b0dce44caea: 424 reachable commits through 99571bfb8efe05f12b3ba762bc988290f16df321.
Engineering main and ms09/20260909-demo-console-freeze resolve to that same documentation checkpoint. The physically exercised executable is f3bf57db5ab2a8623bf89c09c1433e43c6493106; the freeze changed documentation only. An unchanged executable is not a complete model, OS or container snapshot.
Reviewed sources include the live programme registers; CP08–CP15 operations records; the CP13 local and CP15 detail12 freezes; the September 6 engineering achievements record; MS-07 implementation and rollback records; ADR-0057–0066; the demo runbook; and docs/operations/EVOPIEN_THOR_CHECKPOINT_2026-09-09.md, record EVD-MS09-20260909-DEMO-FREEZE-01. The final demo was recorded on September 9 at 22:43–22:52 Europe/Madrid. Public figures here are sanitized aggregates; the raw conversation export is not published.

