Simam is a Windows desktop app that runs open-weight language models on the user’s own computer and gives them a narrow, supervised job. Each teammate has a role, one working folder, a notes memory and one conversation. It can search, read and list the files in that folder. Writing a file or running a command waits for a person to approve it.
It was built for document-heavy teams, starting with construction project managers, who want an assistant to read site diaries and RFIs and draft notes from them, but who cannot send drawings, contracts and site reports to a third-party service.
The interesting part is not the chat window. It is the layer between a small model and the file system: what the model may touch, when it must ask, and how the software checks that what the model says it did actually happened.
Runs on — the user’s PC only, no account and no cloud fallback · Agent test — 24 of 25 tasks (Qwen3 8B, 6 October 2026) · Tools — six; writes and commands ask first · Triggers — chat, weekday schedule, interval, folder change · Command limit — two minutes, whole process tree stopped · Status — Windows beta, unsigned installer
Each teammate is configured once and then used like a colleague in a chat app. The figure shows one at work: it has read an RFI and is paused on the file it wants to write, with the full content shown and Approve and Deny buttons.
A name, a one-line role, written instructions and a single working folder. File tools are confined to that folder.
Search, read and list are On or Off. Writing files and running commands are Off, Ask or Always allow; both default to Ask.
A notes memory of up to 4,000 characters that the teammate adds to and the user can edit, plus an activity log of every run, tool call, approval and failure.
Three statements of the same problem shaped the project.
Firms that handle drawings, contracts and site reports are increasingly offered agent tools that read and write documents, but most are paid cloud services, and every file they touch leaves the business. The brief was a tool that does the same kind of work with the files never leaving the machine.
Models small enough for an office PC, at seven to eight billion parameters, can write well but use tools badly. In our first smoke test on 24 September 2026, llama3.1:8b completed a simple multi-step file task in only one of three tries, and qwen2.5-coder:7b wrote its tool calls as plain text instead of making them. So the reliability layer was built and measured before any teammate features.
A person must be able to see what the software is about to change before it changes, and to see afterwards what it actually did. Neither is optional for a tool that writes into a project folder.
These decisions constrain every feature that sits on top of them.
The agent loop is written in Rust and every tool call passes through one policy check: Off, Ask or Allow, set per teammate and per tool. For Ask, the loop pauses, shows exactly what will be written or run — writes include a preview of the content — and resumes only on Approve. When the window is closed, the request becomes a Windows notification and waits in the tray. Tightening a permission takes effect on the next call, even in the middle of a run.
A reliability layer sits between the model and the tools. It recovers tool calls a model writes as text; repairs near-miss tool and argument names, and refuses calls with missing arguments; allows one call per round; refuses an identical repeat call; catches a final answer that claims a file was saved when no write happened; and stops any command after two minutes, including the processes it started.
The app downloads and manages its own model engine on a private local port and stops it on exit. Models are open-weight files on disk. There is no account, no telemetry and no cloud fallback. The performance and hardware panel shows only real readings from the machine.
The desktop shell is Tauri 2 in Rust with a React and TypeScript interface. Everything below runs inside it.
A single Rust function that drives any model client, tool runner and approval sink. Most of the reliability logic lives here and is covered by unit tests.
Six: search documents, read file, list directory, write file, run command and remember. File tools are confined to the teammate’s folder; commands start there.
Plain JSON and append-only JSONL files in the app’s data folder. Writes are atomic, and a corrupt file is set aside rather than overwriting good data.
A scheduler checks every 30 seconds and watches folders about once a minute. A folder change only counts once it has settled. Missed scheduled runs are not caught up later: a daily run more than ten minutes late is skipped.
A tray icon shows how many approvals are waiting and can pause all routines. Closing the window keeps the app in the tray while any routine is on. It can start hidden at login, and only one copy runs at a time.
A repeatable test of five file tasks against the real engine, with the reliability layer on or off and any number of runs, so a change can be measured rather than argued.
The test gives a model five everyday file tasks in a fresh folder: summarise a site note into a file, find which of three files names the crane operator, open a file from a loosely worded name, count .log files with a shell command, and answer a sum without using tools. Each task runs five times.
25 September, three runs per task: qwen3:8b 14/15, llama3.1:8b 7/15, qwen2.5-coder:7b 3/15.
Same day, same test: qwen3:8b 15/15, qwen2.5-coder:7b 12/15, llama3.1:8b 8/15. The layer’s main effect was on qwen2.5-coder, whose text-written calls are now recovered.
6 October 2026, five runs per task, after the claim check was added: qwen3:8b 24 of 25; qwen2.5-coder:7b 19, 20 and 20 of 25 over three runs.
qwen3:8b clears it comfortably; qwen2.5-coder:7b sits at the bar; llama3.1:8b is below it. The app never suggests a model below the bar for teammates, and if one is chosen it shows a warning that the model may make mistakes with files.
It ran on one machine: an Intel Core i9-11900KF with an NVIDIA RTX 4070 and 32 GB of RAM. Five tasks is a narrow test and five runs is a small sample. Passing it does not mean a model’s written content is correct, only that it used the tools properly and reached the expected answer. That is why approvals show the content in full.
None of these was hypothetical. Each was found in a test or in use, and each changed the code or the process.
The repeat-call guard refused an identical call within a run. That broke a normal pattern: a read fails because a file does not exist yet, the model writes the file, then reads it again to check. The second read was refused as a repeat. What changed, 25 September 2026: the guard now resets after any successful change to the folder, while still refusing an immediate identical repeat of the change itself.
In a live session on 6 October, qwen2.5-coder:7b replied that a daily note “has been written to Reports/daily-note-2026-10-06.md” without calling any tool. Nothing had been written. What changed, the same day: the loop now spots a final answer that names a file as written or saved when no write succeeded in that run. It sends the model one correction, and if the claim is repeated the reply carries a visible note that no file was written. The same session showed this model drifting in longer conversations, copying the shape of its own earlier “done” replies. That is one reason the recommended model is qwen3:8b.
qwen3:8b had been downloaded and evaluated, yet the installed app listed only one model. The model files had been written by a copy of the app started from inside another packaged Windows application, and Windows silently redirects such a process’s writes to AppData into that package’s private storage. The normally installed app reads the real folder, so it never saw them. What changed, 6 October 2026: the files were moved, and the lesson was procedural rather than a code change — test builds are now started the way a user starts them.
One test run scored qwen3:8b at 9 of 25, against 23 of 25 on 25 September. The pattern gave it away: the first two tasks passed, then every attempt failed. The engine had stopped partway through. The run was discarded and recorded as invalid. What changed, 6 October 2026: the harness now prints every failing attempt with its errors, so a stopped engine shows up as errors rather than as a low score. The valid re-run scored 24 of 25.
The system panel said “Integrated” graphics on a machine with an RTX 4070, because graphics detection had never been written and “Integrated” was the fallback text. It had been wrong in every build since the panel was added. The house rule is that telemetry must be real. What changed, 6 October 2026: the app reads the actual graphics card from Windows, ignoring virtual and software adapters, and says “Not detected” when it cannot tell.
These are positions the product takes, not gaps waiting to be filled.
Teammates do not browse the web or click through other programs’ screens. Their tools are the six listed above.
File tools cannot leave the teammate’s folder. Commands start there but can reach other files on the PC, and the app says so in the folder setting. That is why commands default to Ask.
Routines run only while the PC is on and the app is running, in the window or the tray. Missed scheduled runs are skipped, not caught up.
A model can misread a document. Approval shows exactly what will be written so a person can catch it; the activity log shows what happened.
The app does not offer a cloud model as a fallback when a local one struggles.
Windows only, and the installer is not yet code-signed, so Windows shows a warning on install.
The work ahead is about reaching ordinary laptops and one specific industry.
Replace the downloaded engine with a small llama.cpp server bundled in the installer, so only the model downloads on first run.
Re-run the same test on two- to four-billion-parameter models to find which keep this reliability on an 8 GB laptop with no dedicated graphics card.
Turn a folder of site documents into a weekly report grounded in those documents, using the same approval and logging model.
Sign the installer before any public download is offered.
A concise record of what the project set out to prove, the evidence available today, and the next responsible investment step.
Document-heavy teams want AI help with routine reading and reporting, but cannot send project documents to cloud services, and small local models are not reliable enough to trust with files unsupervised.
It tests whether useful, supervised document work is possible entirely on an office PC with open models. That is a prerequisite for offering AI assistance to clients who cannot or will not use cloud tools.
A Windows desktop app with a managed local model engine; a Rust agent loop with per-tool permissions and approvals; a measured reliability layer; teammates with notes memory and an activity log; scheduled and folder-triggered routines with tray notifications; and a repeatable agent test harness.
Writes and commands ask by default. File tools are confined to one folder. Reliability was measured before features were built. No cloud fallback. Missed routine runs are not caught up. Models are recommended only above an 80% test pass rate. What the product does not do is stated publicly.
qwen3:8b completed 24 of 25 attempts at five everyday file tasks on 6 October 2026, five runs per task, on one machine: an i9-11900KF with an RTX 4070 and 32 GB of RAM. The reliability layer took qwen2.5-coder:7b from 3/15 to 12/15 on the same test.
A working Windows beta, exercised on a demo construction project of site diaries and an RFI. Approvals, routines, the activity log and the false-claim check all behave as described in the screenshots. It has not yet been used by external testers.
A short pilot with one construction or consultancy team on their own document folder, measuring which tasks the teammates complete unaided, how often approvals are declined, and whether the slimmer engine and smaller models hold up on their hardware.