- Home
- Case studies
- Work automation
A work-memory agent: answering "what did I do yesterday?" in 0.1 seconds
Without keeping a separate log, this tool reads the session records that several AI tools already leave on the computer and answers what was done and what to do next.
- AreaWork automation
- PeriodOctober 2026
- ToolsNode.js (no external packages), Claude Code, Codex
- StatusIn use

Using several AI tools scatters the work. With Claude in a terminal, Codex on the desktop, and a workspace that runs several agents side by side, each on a different project, a few days later you no longer remember where you did what, or how far you got. We built this tool to solve that.
The tool's documentation and source code, plus the query and test results from runs we made while writing this. The figures are the values from those runs.
The problem
The requirement fit in one sentence: "an agent that remembers what I did in which tool, answers straight away when asked, and tells me what to do next."
We ruled out a work journal from the start. A journal is useless the moment you forget to write it. We chose to read the session records each tool already leaves on this computer. The user records nothing extra.
What it reads
| Tool | What is read |
|---|---|
| Claude Code (terminal, editor, desktop) | Session record files per project |
| Codex (terminal, desktop) | Session records and archived sessions |
| A workspace that runs several agents | Sessions it started, and the list of folders it opened |
| Claude desktop's collaboration mode | Title and first request only |
Four things are extracted from each session record: the requests the user typed, the agent's last reply, the files changed, and how the session ended.
A real query
To choose the projects for these case studies, we asked "what I did in the last 60 days". The answer began like this (translated from Korean).
## Done · last 60 days
26 projects · 102 sessions · 949 requests
Below that, sessions are grouped by project. Each one carries its title, the screen it ran on, the time of the last activity, a few lines of the user's requests, the agent's last reply and the files changed. A session still running in another window is marked "in progress", and one that ended without an answer is marked "ended without a reply".
| Item | Value |
|---|---|
| Sessions in the index | 307 |
| Of which automated runs | 166 |
| Projects in the last 60 days | 26 |
| Sessions in the last 60 days (excluding automated runs) | 102 |
| Requests in the last 60 days | 949 |
The other case studies on this site were chosen from this result. Sorting out two months of work across 26 projects from memory would not have been possible.
Design decisions
It answers without a language model
The lookup itself uses no AI model. Reading the records into an index, and reading the period, tool, project name and search terms out of a question to group the results, are all done by ordinary code. That makes it fast, free to run, and consistent: the same question gets the same answer.
| Operation | Time |
|---|---|
| Rebuilding the whole index | about 6 s |
| Later queries (only changed files are re-read) | about 0.1 s |
It is a Node.js script with no external packages, and the source is about 1,000 lines.
The model is used once, at the end
Extracting "what to do next" mechanically lets noise in. So we split it into two steps. Code picks the candidates, and when the question comes through an agent, the model reads each session's last reply and narrows them down to around three that really need attention.
Candidates are chosen by how sessions from the last seven days ended.
- Requests that were interrupted or ended without a reply
- Sessions that ended with the agent asking the user a question
- Things the agent itself wrote that it "could not do" or that "remain"
- Follow-up work the agent suggested, and uncommitted changes
The same question works everywhere
So that nothing has to be learned per tool, it is connected as an extension to each one. In Claude and in Codex alike, you just ask "What did I do yesterday? What should I do next?" It can also be run directly from a terminal.
Records never leave the computer
Session records contain the work itself. So reading happens only on this computer, and the index file is excluded from the repository. Strings that look like keys or tokens are masked before indexing.
Verification
All 9 automated tests pass; we re-ran them while writing this. Right after building it we asked real questions from three places, Claude, Codex and our own work agent screen, and confirmed that answers came back.
Limits
- Web and mobile chats are invisible. Conversations that leave no record on this computer cannot be read.
- Which screen was used is an estimate. A terminal session run in a folder the workspace opened is counted as done in that workspace.
- Question parsing is simple. In this very query, asking "organise it by project" was taken as a search for the words "by project" and returned the wrong thing. We had to ask briefly, as in "what I did in the last 60 days".
- Masking is not perfect. Some formats may slip past the rules, so a person has to check a result once more before it is taken anywhere else.
- "What to do next" includes things already finished. The model filters them, but the final call is a person's.
What we take from this
Automation does not have to mean an AI model. The most useful part of this tool is the index and lookup that run without one; the model is used only at the end, to narrow the results. Doing with code what fixed rules can do, and placing a model only where judgement is needed, is also how we design our work agent.