Working in Cycles
A practical method for drafting with AI: ask, report, triage, revise. The step everyone skips is the one that decides whether quality rises or drifts.
Most people use an AI assistant the way they'd use a very fast typist. They describe what they want, receive a draft, and then edit it. If the draft is poor they rewrite the request and try again.
That works, sort of, and it plateaus fast. The output keeps arriving at roughly the same quality because nothing in the process is checking it against anything.
There's a better loop. It's slower to describe than to do, and once it's habit you'll run it dozens of times on a single document.
Versions, not edits
Before the loop, one habit that makes the rest work: save a numbered version every lap.
Not because you'll read them all. Because it makes changes reversible, and reversibility is what lets you accept a suggestion you're only 70% sure about. If v7 turns out worse than v6, you go back. Without that, every uncertain change is a small act of faith, and you start refusing good suggestions out of caution.
Versions also let you see drift. Lay v1 beside v40 and you'll sometimes find a carefully hedged sentence has become flatly overstated — not in any single step, but across twenty of them.
Step 1 — Ask a retrieval question, not a writing question
The most common mistake is starting here:
"Write me a section on why the objection was preserved."
You'll get a section. It will be fluent and structurally sound, and you will have no idea which parts are grounded in your record and which the model filled in because the paragraph needed a sentence there.
Ask this instead:
"Find every place in the record where this objection was raised or ruled on. For each, give me the document, the page, and a direct quote. If you cannot find something, say so rather than inferring it."
You get findings, not prose. Each one is separately checkable. And the instruction to admit gaps matters: without it, a model asked to find six things will usually produce six things.
Separating retrieval from composition is the core discipline. Retrieval you can verify. Composition you can judge. Mixed together, you can do neither, because you can't tell which part you're looking at.
What makes a retrieval question good
- Ask for quotes and citations, always. "Summarise what they argued" invites drift. "Quote the two sentences that state their position, with page numbers" does not.
- Ask for the absence. "What does the record NOT say about this?" is often the more valuable question, and models are reluctant to volunteer it.
- Scope it. If you've selected a motion and its responses, say so. A question asked against the whole corpus returns noise from unrelated filings.
- Ask three ways. Retrieval can miss. If something matters, ask for it in the other side's vocabulary as well as your own.
Step 2 — Turn findings into a report
A report is a list of specific, checkable claims about the current version. It is not an essay about quality.
"Here is version 6. For every factual assertion in the argument section, locate its support in the record and give me the page. Produce a list: the assertion, whether support exists, and the citation or the gap. Do not rewrite anything."
That last clause earns its place. Left alone, a model asked to critique will helpfully produce an improved draft, and you'll lose the ability to consider each finding separately — which is the entire point.
Reports worth running repeatedly:
| Report | Asks |
|---|---|
| Support | Does every factual claim trace to a page? |
| Citation | Does every cited case exist, and say what I'm using it for? |
| Unanswered | What did the other side argue that this version doesn't address? |
| Drift | Where have I overstated what the record supports? |
| Procedure | Is each point preserved, and does the record show when it was raised? |
The unanswered-argument report is the one that repays the most, because it's the hardest thing to do unaided. You cannot easily notice the argument you never registered.
Make findings independently acceptable
Structure matters more than it seems. Ask for numbered, atomic findings:
7. Argument §III.B asserts the notice was served on the 14th.
Record support: not located.
Nearest: p.412 references service "that week" without a date.
That can be accepted or rejected on its own. Compare it to "the service timeline section needs work" — which you can only agree with vaguely, and which gives the next step nothing to act on.
Step 3 — Triage. This is the step that matters
Now go through the findings and sort them. Accept. Reject. Defer.
This is where the method lives, and it's the step people skip.
Why accepting everything fails
Accept every suggestion and you have stopped drafting. You're transcribing AI output, and the document drifts toward the model's instincts rather than yours — more hedged, more symmetrical, more generic. Fluent and slightly hollow.
The model has no idea which of your arguments is load-bearing, how this particular court reacts to a particular framing, or what you've deliberately left out for strategic reasons. It cannot know. You can.
Why rejecting out loud matters
Don't just ignore a finding. Say so, and say why:
"Reject 7 — the date is deliberately vague because the record genuinely doesn't establish it, and I don't want to assert more precision than we can support. Don't propose adding a specific date again."
Two reasons this works. It goes into the conversation, so the suggestion stops coming back on subsequent laps. And articulating why forces you to confirm you actually have a reason — occasionally you'll discover you don't, and the finding was right.
Silent rejection is the most expensive habit in this whole method. The same three suggestions return every lap and you re-evaluate them every time.
The three buckets
- Accept — right, and I want it in v+1.
- Reject — wrong, or right but not what I want. Say why.
- Defer — might be right; I need to read the underlying page first. Park it, don't guess.
The defer bucket keeps you honest. The temptation is to accept a plausible finding without opening the cited page. That's how a wrong citation gets laundered into your brief with your own approval on it.
Step 4 — Revise into the next version
Now, and only now, ask for prose:
"Apply findings 2, 5, 9 and 11 only. Leave everything else exactly as it is. Do not reintroduce anything I rejected. Output the full section."
Three things to be firm about:
Accepted items only. Models improve things they weren't asked to improve. Left open, you get a v7 that quietly rewrote a paragraph you were happy with, and you won't spot it without a diff.
Full output, not a patch. "Change the third paragraph" produces ambiguity about what v7 actually is. Take the whole section.
One category per lap, where you can. A lap that fixes citations is easy to evaluate. A lap that fixes citations, restructures the argument and tightens the prose is impossible to evaluate — if v7 is worse, you can't tell which change did it.
Step 5 — Report again, against the new version
Run the report again on v7. Not the old findings — a fresh pass.
Two reasons. Fixes introduce problems: a sentence rewritten for accuracy can break the transition into the next paragraph. And a report against the new text sometimes surfaces something that was always there and didn't surface before, because retrieval isn't deterministic in what it prioritises.
Every few laps, re-run an earlier report too. Citation checks especially. A citation that was fine in v3 can be wrong in v9 because the sentence around it changed while the citation stayed put. That's a regression, and only re-running catches it.
A worked lap
Three cycles, compressed:
Lap 1 — support.
"Version 6 attached. For every factual assertion in §III, find its support in the record. Numbered list: assertion, support or gap, citation. Don't rewrite."
Fourteen findings. Nine supported, three with gaps, two citing the wrong page.
Triage. Accept the two page corrections. Accept one gap — I'd overstated it. Reject one — deliberately vague, said why. Defer one until I read p.412.
Revise. "Apply findings 3, 8 and 12 only. Leave the rest. Full §III." → v7.
Lap 2 — unanswered arguments.
"Here is v7 and their response brief. What do they argue that v7 doesn't address? Quote each, cite it, and say whether v7 engages it."
Four findings. Three I'd covered obliquely; one genuinely unanswered.
Triage. Accept the unanswered one. Reject two — covered deliberately, in the alternative. Defer one.
Revise. → v8.
Lap 3 — drift.
"Compare v8 to v1. Where has a hedged claim become an unqualified one? Quote both versions."
Two findings. Both real. Both my edits, not the model's.
That third report is the one I'd least have thought to run, and it caught a sentence that had been quietly strengthened across four laps until it asserted something the record wouldn't carry.
What goes wrong
Accepting everything. Covered above. You're editing AI output.
Rejecting silently. Suggestions return forever.
Letting the report grade the draft. "Is this argument persuasive?" produces flattery. Ask "which assertions lack record support?" — checkable, not a matter of taste.
Changing too much per lap. You lose the ability to attribute improvement.
Never re-running earlier checks. Regressions accumulate silently.
Treating v50 as necessarily better than v10. Iteration without judgement is drift with extra steps. Every lap must be a decision, not a procedure.
Letting the context get stale. In long sessions the model starts reasoning from its own earlier summaries rather than the documents. If answers feel more confident and less specific, start a fresh conversation with the current version attached.
When to stop
Stop when two consecutive laps produce nothing you want to act on.
Not when it's perfect. A report will always find something — ask a model for problems and it will supply problems. The signal isn't an empty report; it's a report full of findings you keep rejecting, because the remaining suggestions are stylistic preferences rather than defects.
At that point the loop has given you what it has. Everything left is judgement, and judgement was always yours.
Why this works
Nothing here is about the AI being clever. Every step is about making its output checkable, then checking it.
Retrieval separated from composition, so you can verify one and judge the other. Findings atomic, so they can be taken individually. Rejections spoken, so they stick. Versions numbered, so nothing is a one-way door.
What you get isn't a machine that writes well. It's a machine that finds, quickly and specifically, the places where what you wrote and what your record supports have come apart — as many times as you care to ask.
Background: how the retrieval underneath this actually works — embeddings, cosine similarity and pinpoint citation, explained from scratch. Or what this looked like across four days and a hundred and ten versions.