MMarketing Against The Grain
← All frameworks
Strategy

Binary Eval Improvement Flywheel

Use independent pass-fail checks to make AI outputs improve themselves

Difficulty
Moderate
Time to result
~weeks to results
Steps
6
Confidence
99%

The flywheel starts by translating successful examples and explicit requirements into simple pass-fail checks. The agent that generated an output should not be its only evaluator, because self-review introduces bias. A separate agent or context applies the checklist and reports which criteria failed. The generating agent then edits the work and reruns the evaluation until all mandatory checks pass. Numerical scoring is avoided when distinctions such as three versus four out of five cannot be grounded consistently. As people manually review outputs, each recurring correction becomes a candidate for a new binary rule, such as a source-count minimum, prohibited punctuation, or maximum hook length. Over time, the paired skill and eval encode an increasingly precise definition of acceptable work while humans retain responsibility for taste and judgment.

Origin

Peter Yang demonstrated the method on Marketing Against The Grain by drafting a post, evaluating it with a separate pass-fail checklist, and explaining how his newsletter eval evolved from repeated manual editing sessions.

Core principles

  • 01Define quality as observable criteria
  • 02Use a different agent or context to evaluate generated work
  • 03Prefer binary checks over ambiguous numerical scores
  • 04Convert repeated human corrections into durable checks
  • 05Treat skills and evals as living documents
  • 06Reserve subjective taste for human judgment

How to run it

  1. 1

    Define Observable Quality

    Gather strong examples and list characteristics that can be checked consistently. Convert each characteristic into an unambiguous yes-or-no question.

    Pro tip Start with formatting, sourcing, structure, and explicit prohibitions before attempting subjective criteria.

    Watch out If two reviewers could reasonably interpret a check differently, rewrite it.

  2. 2

    Separate Creation and Review

    Have a different agent, context, or process evaluate the output rather than asking the original generator to approve itself. Supply the evaluator with the checklist and necessary evidence.

    Pro tip Keep the evaluation context focused on criteria rather than generation history.

    Watch out Self-evaluation can systematically favor the agent's own wording.

  3. 3

    Run Binary Checks

    Evaluate every required criterion as pass or fail and display the complete result. Avoid unsupported numerical distinctions that create an illusion of precision.

    Pro tip Include evidence or a brief reason beside each failed check.

    Watch out Do not turn subjective taste into a rigid binary rule merely because it is easy to automate.

  4. 4

    Repair Failed Criteria

    Send the failed checks back to the generating agent and ask it to revise the output without breaking checks that already passed. Repeat evaluation after each meaningful revision.

    Pro tip Preserve the original intent while fixing measurable defects.

    Watch out An endless repair loop indicates contradictory checks or missing human judgment.

  5. 5

    Learn From Manual Review

    When a human identifies a recurring problem, ask the agent to propose a new pass-fail check based on that feedback. Approve, revise, or reject the proposed rule before adding it.

    Pro tip Use concrete corrections such as punctuation limits or required source counts.

    Watch out Do not encode a one-off preference as a universal rule without testing it.

  6. 6

    Maintain the Living Eval

    Update the skill and its eval as standards, audiences, and failure patterns change. Periodically remove obsolete or redundant checks.

    Pro tip Version important changes so regressions can be traced.

    Watch out A stale eval can enforce yesterday's quality standard against today's objective.

In the wild

Newsletter Quality Eval

Peter's newsletter eval checks whether the output follows the newsletter format, avoids opening with 'dear subscribers,' keeps the hook short, excludes em dashes, and avoids AI slop. Failed checks trigger another editing pass.

The workflow applies the creator's recurring structural standards to every draft.

Research Coverage Eval

Peter suggested binary research checks such as whether at least ten sources were reviewed and whether recent YouTube interviews from the last thirty days were included.

Research completeness becomes auditable instead of being inferred from fluent prose.

Common mistakes

Using Vague Numerical Scores

An agent may not have a stable distinction between nearby scores, making the result look more precise than it is.

Letting the Writer Grade Itself

The original generating agent can be biased toward approving its own language and decisions.

Automating Taste

Formulaic checks can verify criteria, but human reviewers should decide whether an idea is distinctive, authentic, or creatively strong.

Is it for you?

Best for

It is best for outputs with clear structural, factual, formatting, sourcing, or policy requirements.

Not ideal for

It is not ideal as the sole judge of originality, authenticity, taste, or other deeply subjective qualities.

From the transcript

You don't want the original agent that wrote the draft to run the eval cuz there's some bias there.

Peter Yang · 15:30

So, I just recommend keeping your evals simple. Just do simple pass-fail checks.

Peter Yang · 16:00

Another lesson I think is the skill and the eval that's tied to a skill is like a live working document, right?

Peter Yang · 18:30

From the episode

Automate Boring Tasks With Codex & Claude Code in X Minutes