← The Archive
Issue #6 5 min read

The inbox kit that found a phone call I'd dropped for twelve days

I tested Kit #3's actual workflows on my own inbox before selling it. One came back almost silent. The other found a commitment I'd forgotten I owed someone for twelve days.

The Workflow

Kit #3 was done. Five workflows drafted, voice-checked, pages built, feedback form wired. One gate stood between “done” and “for sale”: every prompt in it had only been tested against sample inboxes and invented threads. Before I put a price on “this reads your inbox better than you do,” I wanted to know if that claim survived an actual mess.

Step 1. Pull a real batch: Grabbed a real 30-email batch straight out of my own inbox, list view, subjects and previews, nothing curated for the demo. Ran it through Workflow 1, the triage pass, exact prompt, no edits.

Step 2. Pick the messiest thread I had: Not a clean example. A 9-message, five-week negotiation with my landlord and her electrician about installing an EV charger outlet. Pulled it straight from Fastmail and ran it through Workflow 2, the task and commitment extractor.

Step 3. Grade the output myself: Read every line of both results against what I actually knew had happened in that thread. I was the only person alive who could check the AI’s work, which is the whole point of testing on your own life instead of a stranger’s.

The triage pass came back nearly silent. The extractor didn’t.

What Broke

The triage pass looked broken before it looked right. Almost every email in the batch came back tagged IGNORE. My first instinct: the prompt’s too conservative, or it’s tuned wrong, or something upstream is broken.

There was no prompt to fix. My inbox already runs aggressive email aliasing, so most of the noise a triage prompt exists to catch had been filtered out before the AI ever read a subject line. The near-silent output wasn’t a bug. It was the tool correctly declining to manufacture urgency it didn’t find.

The lesson: a quiet AI tool and a broken AI tool produce the identical first five seconds of panic. The only way to tell them apart is knowing what “correct” looks like on your own data before you run the test, not after.

Time Ledger

  • Time saved: Not the axis that matters this week. Testing Kit #3 didn’t save time, it validated it.
  • Time added: None. This was the scheduled pre-publish gate for Kit #3, already budgeted into the build.
  • Net: The real number isn’t a saved/added pair. It’s 12. Twelve days a dropped commitment sat in a thread I wasn’t checking, closed the moment an AI read the whole thing instead of skimming it.

Twelve days doesn’t sound like much until you’re the one who owes somebody a phone call. This week’s ledger isn’t about hours. It’s about the gap between what you think you’re tracking and what you’re actually tracking.

The Prompt File

Extract every task, ask, deadline, and promise from these emails into ledger entries. Catch the buried ones; the ask in paragraph four is the one I'll miss on my own.

THE EMAILS:
[Paste full email text, including threads. The more complete the text, the more this catches.]

TODAY'S DATE: [date]

Produce exactly this format:

**MINE** (things I'm being asked to do, or said I'd do)
- [ ] [task, stated plainly] — due [date, or "no date given"] (from: [sender / subject])

**THEIRS** (things others committed to that I should track)
- [Person]: [what they committed to] — due [date, or "no date given"] (from: [subject])

**DEADLINES MENTIONED**
- [date]: [what's due, and whose it is]

**AMBIGUOUS** (might be an ask, might be thinking out loud, my call)
- "[the line, quoted]" — from [sender]

Rules: pull only what's in the emails. Don't infer tasks that aren't there. A vague "we should sync sometime" goes in AMBIGUOUS, not MINE. If the same ask appears in more than one email, list it once and note the repeat; a repeated ask is usually an ask getting urgent.

Manager’s View

You don’t have a team catching what you drop. What’s the one thing you’d trust an AI to check for you first, before you hand it anything client-facing?

Most solo operators can answer this instantly for their deliverables: spellcheck, a second read, maybe a peer review. Almost nobody has applied it to the plumbing underneath the work, the inbox, the task list, the thread they meant to circle back on. There’s no manager cc’d on those. Nobody notices the gap except the client who’s still waiting.

The honest answer isn’t “run everything through AI.” It’s picking the one channel where a dropped ball costs you the most, and pointing a tool at it before you need it, not after the second reminder email shows up.

Field Notes

  • Alias-filtering already does real triage work upstream. My own 30-email test batch barely had noise left to catch; most inboxes won’t start that clean.
  • “Reads the whole thread instead of skimming it” is the entire pitch of Kit #3. The only way to actually prove that claim is running it on a thread you already lived through.
  • The test used the real Fastmail connection, not a copy-pasted sample, on a thread with three actual people: me, the landlord, and her electrician.
The dispatch

One workflow, every Tuesday morning.

Be among the first subscribers. Real workflows from one person doing the work of a whole team, whether that's your own business or a department of one. Free, forever.

No tracking pixels. No drip campaigns. Unsubscribe anytime.