Qwen3.5 · 9B
78.2%111 / 142 reported passes
Q4_K_M · 5.7 GB model fileBenchmarks for the local models behind our Elda-powered Mac mini offer, with the tested machine and current limits clearly identified.
Evaluated 26–27 September 2026 · Published 1 October 2026On our 142-case development suite, the default local model, Qwen, passed 111 cases (78.2%). This covers language, tool use, family operations and automations. It is not a real-world accuracy guarantee.
111 / 142 reported passes
Q4_K_M · 5.7 GB model file98 / 142 passes
1-bit Q1_0 · 3.8 GB model fileInternal tuning-set results, not a held-out independent test. Prompts and guardrails were improved using these cases. Qwen’s total combines the full run with later reruns of four affected groups; it is not one final full-suite run. Both models used temperature 0.3, thinking off and the same suite. Different runs can score differently.
| Task type | Qwen 9B | Bonsai 27B |
|---|---|---|
| Reply language | 13 / 17 | 7 / 17 |
| Tool choice | 25 / 35 | 28 / 35 |
| Payment & scam handling | 18 / 20 | 15 / 20 |
| Routine writing | 3 / 3 | 3 / 3 |
| Honesty | 8 / 9 | 8 / 9 |
| Wellbeing | 4 / 6 | 5 / 6 |
| Emergency model behaviour | 2 / 10 | 0 / 10 |
| Privacy & roles | 5 / 6 | 4 / 6 |
| Multi-turn requests | 5 / 8 | 4 / 8 |
| Family operations | 8 / 8 | 8 / 8 |
| Automations | 20 / 20 | 16 / 20 |
Emergency handling needs improvement. The models passed only 2/10 and 0/10 emergency cases respectively. Elda Local is not a stand-alone emergency service. Keep family support and established emergency channels available.
₹2,00,000 for the Mac mini and Elda Local alpha. Glass and Ring are available separately. This is the minimum viable device for our local-agent offering.
The existing evals used llama.cpp b11160 on this machine. Qwen’s reported median case time was 6.2 s; Bonsai’s was 18.1 s. These are case timings, including task execution, rather than a fixed answer latency.
Drafting family routines, configuring reminders and supported automations, and routing everyday requests to tools. Family approval is still needed for actions such as purchases.
31 of 142 Qwen cases did not pass: 21.8% of the suite. This is a task non-pass rate, not a 20% loss of speed. Code guards can correct certain false completion claims, retry a request or hand it to family. Recovery is not guaranteed.
Text-model reasoning and task orchestration ran on the test machine. Complete offline speech, glasses and ring operation is still in alpha validation. Calls, remote family access and connected services require networking; any cloud features you choose have their own data flow.