← Back to Elda
Elda Local · Alpha / Research note 01

Your agent.
Your machine.

Benchmarks for the local models behind our Elda-powered Mac mini offer, with the tested machine and current limits clearly identified.

Evaluated 26–27 September 2026 · Published 1 October 2026
EldaLocal reasoning, in development.
01 / Measured results

Everyday tasks. Real limits.

On our 142-case development suite, the default local model, Qwen, passed 111 cases (78.2%). This covers language, tool use, family operations and automations. It is not a real-world accuracy guarantee.

Alternative tested

Bonsai · 27B

69.0%

98 / 142 passes

1-bit Q1_0 · 3.8 GB model file
How to read these numbers

Internal tuning-set results, not a held-out independent test. Prompts and guardrails were improved using these cases. Qwen’s total combines the full run with later reruns of four affected groups; it is not one final full-suite run. Both models used temperature 0.3, thinking off and the same suite. Different runs can score differently.

Results by task type +
Reported passes out of cases in each group
Task typeQwen 9BBonsai 27B
Reply language13 / 177 / 17
Tool choice25 / 3528 / 35
Payment & scam handling18 / 2015 / 20
Routine writing3 / 33 / 3
Honesty8 / 98 / 9
Wellbeing4 / 65 / 6
Emergency model behaviour2 / 100 / 10
Privacy & roles5 / 64 / 6
Multi-turn requests5 / 84 / 8
Family operations8 / 88 / 8
Automations20 / 2016 / 20

Emergency handling needs improvement. The models passed only 2/10 and 0/10 emergency cases respectively. Elda Local is not a stand-alone emergency service. Keep family support and established emergency channels available.

02 / Device & response time

M6 Mac mini. 24 GB. Elda at home.

Minimum viable device · Offered configuration

Elda-powered Mac mini
M6 · 24 GB unified memory

₹2,00,000 for the Mac mini and Elda Local alpha. Glass and Ring are available separately. This is the minimum viable device for our local-agent offering.

Apple M6 Mac mini specifications ↗

Recorded benchmark setup

MacBook Pro
M5 Pro · 24 GB

The existing evals used llama.cpp b11160 on this machine. Qwen’s reported median case time was 6.2 s; Bonsai’s was 18.1 s. These are case timings, including task execution, rather than a fixed answer latency.

03 / What to expect

Useful help, with a way to ask for help.

A good starting point

Drafting family routines, configuring reminders and supported automations, and routing everyday requests to tools. Family approval is still needed for actions such as purchases.

When the model misses

31 of 142 Qwen cases did not pass: 21.8% of the suite. This is a task non-pass rate, not a 20% loss of speed. Code guards can correct certain false completion claims, retry a request or hand it to family. Recovery is not guaranteed.

What “local” means here

Text-model reasoning and task orchestration ran on the test machine. Complete offline speech, glasses and ring operation is still in alpha validation. Calls, remote family access and connected services require networking; any cloud features you choose have their own data flow.