Simam AI Lab Our applied AI research division is now open. Visit the lab
Simam AI Lab product · Windows beta

AI teammates that work on the files in your folders, on your own PC, and ask before they change anything.

Simam is a Windows desktop app that runs open-weight language models on the user’s own computer and gives them a narrow, supervised job. Each teammate has a role, one working folder, a notes memory and one conversation. It can search, read and list the files in that folder. Writing a file or running a command waits for a person to approve it.

It was built for document-heavy teams, starting with construction project managers, who want an assistant to read site diaries and RFIs and draft notes from them, but who cannot send drawings, contracts and site reports to a third-party service.

The interesting part is not the chat window. It is the layer between a small model and the file system: what the model may touch, when it must ask, and how the software checks that what the model says it did actually happened.

Runs on — the user’s PC only, no account and no cloud fallback  ·  Agent test — 24 of 25 tasks (Qwen3 8B, 6 October 2026)  ·  Tools — six; writes and commands ask first  ·  Triggers — chat, weekday schedule, interval, folder change  ·  Command limit — two minutes, whole process tree stopped  ·  Status — Windows beta, unsigned installer

Overview

A teammate is a model, a folder, a set of permissions and a record of everything it did.

Each teammate is configured once and then used like a colleague in a chat app. The figure shows one at work: it has read an RFI and is paused on the file it wants to write, with the full content shown and Approve and Deny buttons.

Role and folder

A name, a one-line role, written instructions and a single working folder. File tools are confined to that folder.

Permissions per tool

Search, read and list are On or Off. Writing files and running commands are Off, Ask or Always allow; both default to Ask.

Memory and record

A notes memory of up to 4,000 characters that the teammate adds to and the user can edit, plus an activity log of every run, tool call, approval and failure.

The brief

The problem was trust and reliability, not intelligence.

Three statements of the same problem shaped the project.

Document work should not leave the building

Firms that handle drawings, contracts and site reports are increasingly offered agent tools that read and write documents, but most are paid cloud services, and every file they touch leaves the business. The brief was a tool that does the same kind of work with the files never leaving the machine.

Small local models are unreliable agents out of the box

Models small enough for an office PC, at seven to eight billion parameters, can write well but use tools badly. In our first smoke test on 24 September 2026, llama3.1:8b completed a simple multi-step file task in only one of three tries, and qwen2.5-coder:7b wrote its tool calls as plain text instead of making them. So the reliability layer was built and measured before any teammate features.

An assistant that acts on files has to ask, and has to be checkable

A person must be able to see what the software is about to change before it changes, and to see afterwards what it actually did. Neither is optional for a tool that writes into a project folder.

Architecture

Three decisions do most of the work: ask before acting, check what the model claims, and keep everything on the machine.

These decisions constrain every feature that sits on top of them.

Approval is part of the loop, not a dialog bolted on

The agent loop is written in Rust and every tool call passes through one policy check: Off, Ask or Allow, set per teammate and per tool. For Ask, the loop pauses, shows exactly what will be written or run — writes include a preview of the content — and resumes only on Approve. When the window is closed, the request becomes a Windows notification and waits in the tray. Tightening a permission takes effect on the next call, even in the middle of a run.

The software checks the model, not the other way round

A reliability layer sits between the model and the tools. It recovers tool calls a model writes as text; repairs near-miss tool and argument names, and refuses calls with missing arguments; allows one call per round; refuses an identical repeat call; catches a final answer that claims a file was saved when no write happened; and stops any command after two minutes, including the processes it started.

Local means local, including the engine

The app downloads and manages its own model engine on a private local port and stops it on exit. Models are open-weight files on disk. There is no account, no telemetry and no cloud fallback. The performance and hardware panel shows only real readings from the machine.

Components

Six parts, each small enough to test on its own.

The desktop shell is Tauri 2 in Rust with a React and TypeScript interface. Everything below runs inside it.

Agent loop

A single Rust function that drives any model client, tool runner and approval sink. Most of the reliability logic lives here and is covered by unit tests.

Tools

Six: search documents, read file, list directory, write file, run command and remember. File tools are confined to the teammate’s folder; commands start there.

Teammate store

Plain JSON and append-only JSONL files in the app’s data folder. Writes are atomic, and a corrupt file is set aside rather than overwriting good data.

Routines

A scheduler checks every 30 seconds and watches folders about once a minute. A folder change only counts once it has settled. Missed scheduled runs are not caught up later: a daily run more than ten minutes late is skipped.

Tray and notifications

A tray icon shows how many approvals are waiting and can pause all routines. Closing the window keeps the app in the tray while any routine is on. It can start hidden at login, and only one copy runs at a time.

Agent test harness

A repeatable test of five file tasks against the real engine, with the reliability layer on or off and any number of runs, so a change can be measured rather than argued.

Evidence

Reliability was measured with a fixed test, and the test’s limits are stated.

The test gives a model five everyday file tasks in a fresh folder: summarise a site note into a file, find which of three files names the crane operator, open a file from a loosely worded name, count .log files with a shell command, and answer a sum without using tools. Each task runs five times.

Before the reliability layer

25 September, three runs per task: qwen3:8b 14/15, llama3.1:8b 7/15, qwen2.5-coder:7b 3/15.

With the layer

Same day, same test: qwen3:8b 15/15, qwen2.5-coder:7b 12/15, llama3.1:8b 8/15. The layer’s main effect was on qwen2.5-coder, whose text-written calls are now recovered.

Current beta

6 October 2026, five runs per task, after the claim check was added: qwen3:8b 24 of 25; qwen2.5-coder:7b 19, 20 and 20 of 25 over three runs.

The bar for recommending a model is 80%

qwen3:8b clears it comfortably; qwen2.5-coder:7b sits at the bar; llama3.1:8b is below it. The app never suggests a model below the bar for teammates, and if one is chosen it shows a warning that the model may make mistakes with files.

What the test does not prove

It ran on one machine: an Intel Core i9-11900KF with an NVIDIA RTX 4070 and 32 GB of RAM. Five tasks is a narrow test and five runs is a small sample. Passing it does not mean a model’s written content is correct, only that it used the tools properly and reached the expected answer. That is why approvals show the content in full.

What went wrong

Five problems that changed the design, recorded as they happened.

None of these was hypothetical. Each was found in a test or in use, and each changed the code or the process.

A guard against loops blocked legitimate work

The repeat-call guard refused an identical call within a run. That broke a normal pattern: a read fails because a file does not exist yet, the model writes the file, then reads it again to check. The second read was refused as a repeat. What changed, 25 September 2026: the guard now resets after any successful change to the folder, while still refusing an immediate identical repeat of the change itself.

A model said it had saved a file when it had not

In a live session on 6 October, qwen2.5-coder:7b replied that a daily note “has been written to Reports/daily-note-2026-10-06.md” without calling any tool. Nothing had been written. What changed, the same day: the loop now spots a final answer that names a file as written or saved when no write succeeded in that run. It sends the model one correction, and if the claim is repeated the reply carries a visible note that no file was written. The same session showed this model drifting in longer conversations, copying the shape of its own earlier “done” replies. That is one reason the recommended model is qwen3:8b.

A downloaded model was invisible to the app

qwen3:8b had been downloaded and evaluated, yet the installed app listed only one model. The model files had been written by a copy of the app started from inside another packaged Windows application, and Windows silently redirects such a process’s writes to AppData into that package’s private storage. The normally installed app reads the real folder, so it never saw them. What changed, 6 October 2026: the files were moved, and the lesson was procedural rather than a code change — test builds are now started the way a user starts them.

A dead engine looked like a model regression

One test run scored qwen3:8b at 9 of 25, against 23 of 25 on 25 September. The pattern gave it away: the first two tasks passed, then every attempt failed. The engine had stopped partway through. The run was discarded and recorded as invalid. What changed, 6 October 2026: the harness now prints every failing attempt with its errors, so a stopped engine shows up as errors rather than as a low score. The valid re-run scored 24 of 25.

The hardware panel reported a guess

The system panel said “Integrated” graphics on a machine with an RTX 4070, because graphics detection had never been written and “Integrated” was the fallback text. It had been wrong in every build since the panel was added. The house rule is that telemetry must be real. What changed, 6 October 2026: the app reads the actual graphics card from Windows, ignoring virtual and software adapters, and says “Not detected” when it cannot tell.

Limits

What it does not do is as deliberate as what it does.

These are positions the product takes, not gaps waiting to be filled.

No web, no screen control

Teammates do not browse the web or click through other programs’ screens. Their tools are the six listed above.

Folder-confined files, honest about commands

File tools cannot leave the teammate’s folder. Commands start there but can reach other files on the PC, and the app says so in the folder setting. That is why commands default to Ask.

Not always on

Routines run only while the PC is on and the app is running, in the window or the tray. Missed scheduled runs are skipped, not caught up.

No guarantee of correct content

A model can misread a document. Approval shows exactly what will be written so a person can catch it; the activity log shows what happened.

No cloud models

The app does not offer a cloud model as a fallback when a local one struggles.

Not yet signed or cross-platform

Windows only, and the installer is not yet code-signed, so Windows shows a warning on install.

Next

The next step is to make it lighter, not cleverer.

The work ahead is about reaching ordinary laptops and one specific industry.

A slimmer bundled engine

Replace the downloaded engine with a small llama.cpp server bundled in the installer, so only the model downloads on first run.

Smaller recommended models

Re-run the same test on two- to four-billion-parameter models to find which keep this reliability on an 8 GB laptop with no dedicated graphics card.

A construction Project Agent pack

Turn a folder of site documents into a weekly report grounded in those documents, using the same approval and logging model.

Code signing

Sign the installer before any public download is offered.

Case study decision record

The commercial case, in one view.

A concise record of what the project set out to prove, the evidence available today, and the next responsible investment step.

Business challenge

Document-heavy teams want AI help with routine reading and reporting, but cannot send project documents to cloud services, and small local models are not reliable enough to trust with files unsupervised.

Why the project mattered

It tests whether useful, supervised document work is possible entirely on an office PC with open models. That is a prerequisite for offering AI assistance to clients who cannot or will not use cloud tools.

What Simam Digital designed and built

A Windows desktop app with a managed local model engine; a Rust agent loop with per-tool permissions and approvals; a measured reliability layer; teammates with notes memory and an activity log; scheduled and folder-triggered routines with tray notifications; and a repeatable agent test harness.

Important product decisions

Writes and commands ask by default. File tools are confined to one folder. Reliability was measured before features were built. No cloud fallback. Missed routine runs are not caught up. Models are recommended only above an 80% test pass rate. What the product does not do is stated publicly.

Measured evidence

qwen3:8b completed 24 of 25 attempts at five everyday file tasks on 6 October 2026, five runs per task, on one machine: an i9-11900KF with an RTX 4070 and 32 GB of RAM. The reliability layer took qwen2.5-coder:7b from 3/15 to 12/15 on the same test.

Credible outcome

A working Windows beta, exercised on a demo construction project of site diaries and an RFI. Approvals, routines, the activity log and the false-claim check all behave as described in the screenshots. It has not yet been used by external testers.

Recommended next engagement

A short pilot with one construction or consultancy team on their own document folder, measuring which tasks the teammates complete unaided, how often approvals are declined, and whether the slimmer engine and smaller models hold up on their hardware.