
At work this thing decides architecture. Which database. Which pattern. Which decision we are going to regret in eighteen months. It does not guess and it does not round up. It reads, it argues with itself, and it shows you where every number came from.
This morning I gave it a river.
01 The premise
Everybody is using this stuff to answer email
Every conversation I have about AI right now is about work. Ship faster. Close the ticket faster. Write the deck in four minutes instead of forty. Fine. Somebody had to say it and now everybody has.
Here is the part nobody mentions. Careful research used to cost a week. Now it costs an afternoon. And careful research was never really a work thing. It is what you would want before you pick a school, or a surgeon, or a used truck, or a river. Almost nobody does it, because who has a week.
I have a Saturday. So I spent this morning writing the question down properly.
02 Two runs
Same machine, two weeks apart
Two weeks ago it spent about thirty five minutes on what an agentic software factory is. Sixty two sources. Three researchers. Nobody dissented, which almost never happens.
This week it gets three rivers and a Saturday. Same skill, same prompt files, same four rounds. I changed the brief. That is the entire difference.
What does “agentic software factory” actually mean in July 2026, and how should the definition change over the next six months?
- Run date
- 2026-07-11
- Researchers
- 3
- Sources cited
- 62
- Uncertainty tags
- 77 raw, 20 kept
- Dissents
- 0
- Start to finish
- about 35 minutes
One day of bank fishing. Three rivers. Eighteen days to pick from. Which river, which stretch, which day, and what would make me stay in bed.
- Brief written
- 2026-07-24
- Researchers
- 4
- Questions
- 10
- Rivers
- Kalama, Cowlitz, Wynoochee
- Window
- 2026-07-24 to 2026-08-10
- History to pull
- up to 20 years
The work run is finished, so those are results. The fishing run is not, so those are the settings going in. What follows is the input. The input is where all the work is.
03 The brief
Ten questions, and not one of them is “is the fishing good”
The machine does not get a question. It gets a brief. What I am deciding, what is in bounds, the rules it has to follow, what finished looks like. A sloppy brief buys you a confident report that is wrong. So this part I still write by hand.
How many fish are in
Real counts on each river this year, with a date attached. Plus how much of a normal season is usually in by now. A big number in July is not the same animal as a big number in September.
What they predicted
What the preseason forecast said, which document said it, how they arrived at it, and whether the fish are running ahead of it or behind it.
Where this year sits
Against the last five, ten, and twenty. Then every reason that comparison might be a lie. They changed how many smolt go in. They changed how they count. The program itself changed. Any one of those and you are comparing two different things.
The rules, exactly
Which stretches are open, by river mile. Season dates. Daily limit. What you can keep. Gear rules. Night closure. Every one cited to the published regulation with a link and the day it was read. Including which markers I have to find on the bank with my own eyes.
What could shut it down
Which emergency rules are live. The exact number that trips each one. Who announces it. How much warning you normally get. And the single page I check at four in the morning before I put the truck in gear.
Water and weather
Flow, temperature, and clarity by gauge number. Dam release schedules. And which hours are honestly fishable, which in late July is a shorter list than anyone wants to admit.
Where I can stand
Holes and drifts you can reach on foot. Where to park. Where you would be trespassing. Where the tribal water starts. Ranked by crowded and by hard. And straight about the spots where a guy in a boat beats me no matter what I do.
What people are saying
Forums, Reddit, the Facebook groups, YouTube, shop reports. Dated and attributed. Real catches separated from guys repeating what they heard. Then held up against the actual creel numbers, because slow is a memory, not a measurement.
What is working
What is producing this year and what is not. Fly and gear. Which water, which hours.
The call
One day. Then two days. Name the river, name the stretch, name the date, hour by hour, drive time from my driveway, and the one thing that calls the whole trip off.
04 The catch
Three rivers, three different machines counting fish
This is the part I would show a friend who does not fish.
The Cowlitz counts fish at a separator below a dam. The Kalama counts them at a hatchery rack. The Wynoochee counts them in a trap. Three numbers. Three different objects.
Kalama
- Basin
- Lower Columbia tributary
- County
- Cowlitz
- Counted at
- Hatchery rack and trap at the falls
Cowlitz
- Basin
- Lower Columbia tributary
- County
- Lewis and Cowlitz, below Barrier Dam
- Counted at
- Barrier Dam separator
Wynoochee
- Basin
- Chehalis
- County
- Grays Harbor
- Counted at
- Wynoochee Dam trap
Put those three numbers in the same column and they look like the same thing. They are not the same thing. That is how you end up driving two hours to the wrong river.
So the brief makes it a contract. Every researcher has to say out loud what their instrument physically counts, and over what stretch of time, and whether it can be compared to the others at all. Cross-river claims get converted to percent of that river's own ten year average for the same date, which is the only common denominator that survives the trip. Anything built on mismatched numbers gets tagged [UNVERIFIED]. And when a number does not exist for one river, the box says [GAP]. It does not quietly get filled in with something close.
Two of these rivers answer to one set of managers. The third answers to another. Different forecasts, different rule authority, different reasons to close. The brief says do not blend them, and it says it twice.
05 The rules
What makes it research instead of a guy talking
Every one of these is in there because the thing it stops is what happens by default. They go into every researcher word for word. Not up for discussion.
There is a fish count at Bonneville Dam that is easy to get and would look terrific in a report. Both of my Columbia rivers come in below Bonneville. Those fish are headed somewhere else entirely. The most convenient number available is the wrong one, so the brief names it and bans it.
Every rule claim needs the published regulation, a link, and the date it was read. A limit somebody posted on a forum gets tagged [UNVERIFIED] and it does not go near the trip plan. Not as a footnote. Not as a maybe.
Blank days count. “Nobody is catching anything” counts. If the guys on the water say one thing and the counts say another, the report writes down that they disagree. It does not get to pick the happier one.
A river that normally gets talked about in the last week of July, and this year, nothing. That is a finding. It goes in the report as a finding.
“Guys are getting one or two” means nothing by itself. One or two per what. The state publishes catch per rod hour going back years. Somebody's mood gets held up against that number before it counts as evidence.
The machine reports what the official forecast said and how the year is running against it. It does not build one of its own. Anything about a day that has not happened yet gets tagged [PROJECTION] and it is barred from the summary.
Each researcher owns a subject, not a river. Every one of them has to name which river their own evidence argues against. If the evidence likes all three, they have to say that out loud instead of inventing a preference.
“Returns look good” is not a finding. The brief spells out what a finding sounds like: 1,842 summer-run through the Cowlitz separator as of July 20, versus a ten year average of 2,610 for the same date. Invented numbers, real shape. Every claim has to land like that.
The tags. [GAP] the number does not exist or could not be reached. [UNVERIFIED] one source, and here is what would settle it. [PROJECTION] a day that has not happened yet. My favorite line in the whole protocol file: a report with no [UNVERIFIED] tags is suspicious, not impressive.
06 The rounds
Four rounds, and the third one is the one that matters
- R0
Everybody gets the assignment
The question, their piece of it, and the rules. Everybody answers with one line saying they got it. That is all. Repeating the instructions back is not free.
- R1
Nobody talks to anybody
Each one writes a near-final draft, a numbered list of sources with the day each was pulled, the claims they would defend if pushed, and the questions they want to put to the others. Alone. That part is deliberate. Four opinions are only worth colliding if they were four opinions to start with.
- R2
Now they read each other
Each one has to find a weak claim in every other draft. Each one has to argue against their own draft as hard as they can, including objections nobody raised. And each one has to name the thing they believe least. All three are required. Leave one out and the chair hands it back.
- R3
The chair writes it
Alone. Where two of them still disagree and neither will fold, the disagreement goes into the report by name, with what would settle it. Everybody agreeing is not the goal. Everybody agreeing is usually a sign that nobody pushed.
On the work run, round two is where it turned. One of them wrote this against their own argument: “I cannot claim factories work AND claim deployable reliability stays coin-flip.” Nobody asked them to. The round made them.
The report has to be willing to tell me to stay home.
Success criteria, brief.md
07 Finished
What finished looks like
Written before the run, not after. Afterward you just grade to whatever you got.
- I can leave the driveway at four in the morning on a named day with a river, a stretch, a plan, and a one minute check that it is still on. Without opening another tab.
- Every count, flow, temperature, and rule carries a source and a date, exact enough that I could go run the same comparison myself.
- A guy who has fished these three rivers for twenty years still learns one thing. A count trend, a program change, an access point, a rule that changed while he was not looking.
- It is willing to say do not bother, this river is a bust this year. And it says it at the top, not on page nine.
08 The point
The machine does not know it is on vacation
It runs the same four rounds on a river that it runs on a build versus buy call. None of the machinery is about software. It is about refusing to answer a question you have not actually looked into.
Which means the interesting move is not getting more done at work. It is looking at all the decisions in your own life you have been making off a forum post and a hunch.
The school
Enrollment trend. Staff turnover. What the report card measures and what it leaves out. What parents say, sitting next to what the numbers say.
The surgeon
How many of your exact procedure they did last year. Published outcomes. Board actions. And why the review sites are built to mislead you.
The truck
Which model years fail and how. What it really costs at your mileage. Which recall in that list is the one that matters.
The move
Not the listicle. Water rights. Where the taxes are headed. Whether anybody will insure it. The commute at the hour you would really be driving it.
I still have to stand in the water and cast. Nobody is doing that part for me. But I am going to be standing in the right water, at the right hour, on a river that is actually holding fish, and that is most of it.
This morning I gave it a river. Three of them, and eighteen days to choose from. I wrote all of this down before it ran. The report landed that afternoon. What it said is the next one.
Everything on this page is the input. Quoted or shortened from the brief and the protocol file, plus results from the work run on 2026-07-11. No fish counts, flows, or regulations are claimed here. Those come out the other end, with sources and dates attached.