1. Home
  2. Case studies
  3. Work automation

A work-memory agent: answering "what did I do yesterday?" in 0.1 seconds

Without keeping a separate log, this tool reads the session records that several AI tools already leave on the computer and answers what was done and what to do next.

  • AreaWork automation
  • PeriodOctober 2026
  • ToolsNode.js (no external packages), Claude Code, Codex
  • StatusIn use
A black card box lined with blank index cards, with an orange divider tab rising among them

Using several AI tools scatters the work. With Claude in a terminal, Codex on the desktop, and a workspace that runs several agents side by side, each on a different project, a few days later you no longer remember where you did what, or how far you got. We built this tool to solve that.

What this is based on

The tool's documentation and source code, plus the query and test results from runs we made while writing this. The figures are the values from those runs.

The problem

The requirement fit in one sentence: "an agent that remembers what I did in which tool, answers straight away when asked, and tells me what to do next."

We ruled out a work journal from the start. A journal is useless the moment you forget to write it. We chose to read the session records each tool already leaves on this computer. The user records nothing extra.

What it reads

ToolWhat is read
Claude Code (terminal, editor, desktop)Session record files per project
Codex (terminal, desktop)Session records and archived sessions
A workspace that runs several agentsSessions it started, and the list of folders it opened
Claude desktop's collaboration modeTitle and first request only

Four things are extracted from each session record: the requests the user typed, the agent's last reply, the files changed, and how the session ended.

A real query

To choose the projects for these case studies, we asked "what I did in the last 60 days". The answer began like this (translated from Korean).

## Done · last 60 days
26 projects · 102 sessions · 949 requests

Below that, sessions are grouped by project. Each one carries its title, the screen it ran on, the time of the last activity, a few lines of the user's requests, the agent's last reply and the files changed. A session still running in another window is marked "in progress", and one that ended without an answer is marked "ended without a reply".

ItemValue
Sessions in the index307
Of which automated runs166
Projects in the last 60 days26
Sessions in the last 60 days (excluding automated runs)102
Requests in the last 60 days949

The other case studies on this site were chosen from this result. Sorting out two months of work across 26 projects from memory would not have been possible.

Design decisions

It answers without a language model

The lookup itself uses no AI model. Reading the records into an index, and reading the period, tool, project name and search terms out of a question to group the results, are all done by ordinary code. That makes it fast, free to run, and consistent: the same question gets the same answer.

OperationTime
Rebuilding the whole indexabout 6 s
Later queries (only changed files are re-read)about 0.1 s

It is a Node.js script with no external packages, and the source is about 1,000 lines.

The model is used once, at the end

Extracting "what to do next" mechanically lets noise in. So we split it into two steps. Code picks the candidates, and when the question comes through an agent, the model reads each session's last reply and narrows them down to around three that really need attention.

Candidates are chosen by how sessions from the last seven days ended.

  1. Requests that were interrupted or ended without a reply
  2. Sessions that ended with the agent asking the user a question
  3. Things the agent itself wrote that it "could not do" or that "remain"
  4. Follow-up work the agent suggested, and uncommitted changes

The same question works everywhere

So that nothing has to be learned per tool, it is connected as an extension to each one. In Claude and in Codex alike, you just ask "What did I do yesterday? What should I do next?" It can also be run directly from a terminal.

Records never leave the computer

Session records contain the work itself. So reading happens only on this computer, and the index file is excluded from the repository. Strings that look like keys or tokens are masked before indexing.

Verification

All 9 automated tests pass; we re-ran them while writing this. Right after building it we asked real questions from three places, Claude, Codex and our own work agent screen, and confirmed that answers came back.

Limits

  • Web and mobile chats are invisible. Conversations that leave no record on this computer cannot be read.
  • Which screen was used is an estimate. A terminal session run in a folder the workspace opened is counted as done in that workspace.
  • Question parsing is simple. In this very query, asking "organise it by project" was taken as a search for the words "by project" and returned the wrong thing. We had to ask briefly, as in "what I did in the last 60 days".
  • Masking is not perfect. Some formats may slip past the rules, so a person has to check a result once more before it is taken anywhere else.
  • "What to do next" includes things already finished. The model filters them, but the final call is a person's.

What we take from this

Automation does not have to mean an AI model. The most useful part of this tool is the index and lookup that run without one; the model is used only at the end, to narrow the results. Doing with code what fixed rules can do, and placing a model only where judgement is needed, is also how we design our work agent.

Get in touch.

We are preparing our first client agents. If you have a repetitive task you would like to hand to an agent, tell us about it and we will look at a small trial together.

TopicsTrials · Workshops · Partnerships