antiagree
A second opinion that doesn't just agree with you. antiagree reviews your idea, plan, architecture, benchmark or decision the way an outside reviewer with no stake in it would. It checks your claims against the code and the data, looks for what already exists, tries to kill the proposal with the strongest argument it can find, and keeps only what survives. When your plan is sound, it says GO.
01Overview
antiagree is a single Claude skill, shipped as a Claude Code plugin and as an Agent Skill for other agents. Every review ends in one decision you can act on.
Reads first facts
Checks the few claims your plan depends on against the code, data, benchmarks and git history before it judges.
One verdict decision
GO, ITERATE, RETHINK or KILL, with the biggest risk and the cheapest experiment that settles it.
Holds you to it ledger
Records kill criteria in ANTIAGREE.md and reads them first next time.
Audits the session review
With no target, it reviews what you and Claude already decided in the conversation.
Who it is for: people who use Claude to make real decisions about code, products, data or money, and want the decision tested before they commit.
What it is not: a security scanner, a linter, or a tool that criticizes everything. It changes position on new facts or better arguments, not on displeasure.
02Why it exists
Language models lean toward agreeing with the person in front of them. Ask "is this a good plan?" and you often get a polished yes, a list of minor tweaks, and an estimate nobody measured.
The opposite failure is just as common: ask for brutal honesty and you get invented problems, because that is what the role seems to want. Both tell you what you asked to hear instead of what is true.
antiagree tries to prove the proposal wrong and keeps it only if it survives. Its loyalty is to your goal, not to your proposal, your earlier decisions or anything it said earlier in the conversation. A short "this holds up" with reasons is a valid result.
03A review, step by step
Check the ledger
If ANTIAGREE.md exists, it reads it first. A kill criterion that was met, or a deadline that passed, leads the review.
Investigate before judging
It verifies the three to five claims the decision depends on. With no artifacts, it says the review is reasoning-only.
Check whether it already exists
For anything you plan to build, it looks for prior art and asks what is actually different.
Find the real goal
What result is wanted, which variable matters, and whether the question itself is framed wrong.
Attack
Assumptions, activity versus impact, causality, false improvements, the measurement, hidden costs, second-order effects, people and incentives, sunk cost, simpler paths, platform risk.
Try to kill it
It builds the strongest case against the proposal, not a straw man, and says whether the proposal survived.
Rebuild
The bottleneck, the few moves with real leverage, and experiments with GO and KILL thresholds.
04The verdict
Every review opens with a verdict card: what is being reviewed, the verdict, the biggest risk, the next step, and the one fact that would change its mind.
Sound; proceed. The default when the diagnosis is supported, the fix is standard, and it can be verified and rolled back cheaply.
Right goal and approach, with a must-fix the plan doesn't already cover.
Right goal, wrong approach.
Drop it.
Verdict: ITERATE. The benchmark measures a 100% hit rate, which production won't have. The rollout plan ships a correctness risk with no safeguard.
Biggest risk: serving stale product data (price, stock) across 14 endpoints, with user complaints as the only way to detect it.
Next step: replay a sample of real production query logs against the cache and measure hit rate and p50/p95/p99. GO if the improvement holds at the real hit rate.
Output from the planted-flaw-cache eval case on Sonnet 5.5, trimmed.
After the card come the findings, ranked by leverage, and five closing sections: the uncomfortable truth, keep, change, don't build, and still unknown.
05Confidence labels
Every claim that matters carries a label, in your language, and a source when there is one: file:line, a short quote, a URL, or "from your description".
In Spanish: HECHO, FUERTEMENTE SOPORTADO, HIPÓTESIS, ESPECULACIÓN. No invented percentages: if nobody knows whether it's 5% or 30%, the answer is "we don't know yet".
06Reviewing a session
Run /antiagree with nothing after it and it reviews the decisions made in the current conversation, from your first message on.
- It says up front that it took part in those decisions and is anchored to them.
- It opens with a table of every number and claim the session accepted: who introduced it, whether it was verified and how, and what depends on it.
- It checks whether the goal drifted, and separates what got done from what got closer to the goal.
- When the stakes are high, it can offer a blind second opinion from a subagent that sees the goal and the artifacts, not the conclusions.
07Kill criteria
When a review produces experiments with thresholds, antiagree offers to record them in ANTIAGREE.md at the project root. It writes only if you agree, appends, and never edits past entries except their status line.
## 2026-10-05 - What was decided
- Hypothesis: ...
- Test: ... (by YYYY-MM-DD)
- GO if ... / KILL if ...
- Status: open
The next review reads the file first. You set the bar before you were attached to the outcome, so the review doesn't let you move it after the result is in without new facts about the threshold itself.
08Install
Claude Code
/plugin marketplace add antiagree/antiagree
/plugin install antiagree@antiagree
Other agents with Agent Skills
Codex, Cursor, Gemini CLI and others: copy skills/antiagree/ into the agent's skills directory, for example ~/.claude/skills/antiagree/ for Claude Code without the plugin.
09Use
| Command | What it does |
|---|---|
/antiagree | Reviews this session's decisions. |
/antiagree our plan to split the monolith | Reviews anything: an idea in prose, a file, a directory, a PR, a URL. |
/antiagree quick bench/RESULTS.md | Verdict card plus the top three findings. |
Or just ask: "attack my plan", "be brutally honest", "don't just agree with me", "no me des la razón". Once invoked, it keeps the stance for the rest of the session until you tell it to stop.
10How it's checked
The repository ships five cases for claude plugin eval. Four run with and without the plugin; session-review runs with the plugin only.
| Case | What it tests |
|---|---|
planted-flaw-cache | A benchmark with planted flaws: warm-cache measurement, confounded arms, deferred invalidation. |
sound-plan-no-theater | A plan that is actually sound, with "be brutal" in the prompt. Does it say GO or invent problems? |
activity-vs-impact-es | In Spanish: a busy month and a confidence interval that includes zero. |
ledger-holds-you-to-it | A kill criterion in ANTIAGREE.md that the latest results met. |
session-review | A session that agreed its way into a plan with unverified numbers. |
Results from 2026-10-08, three runs per case, Opus 5.5 as the judge:
| Model | With antiagree | Without |
|---|---|---|
| Sonnet 5.5 | 15/15 | 0/12 |
| Opus 5.5 | 15/15 | 3/12 |
| Haiku 4.5 | 2/15 | 0/12 |
Without the plugin, Sonnet and Opus still find the planted flaws, but they invent problems in the sound plan and give no decision rule for the benchmark. Sonnet also misses the ledger's kill criterion.
11What it reads, writes and sends
| What | Details |
|---|---|
| Reads | Files in your project that the review depends on: code, data, benchmarks, git history. |
| Writes | ANTIAGREE.md, and only when you agree. |
| Sends | Nothing of its own. It may ask to run web searches through Claude's WebSearch tool, for example to check prior art, so those searches include terms from your proposal. |
| Context | Like any skill, its short description is always in Claude's context so it can trigger when you ask; the full instructions load only when it runs. |
It is instructions only: no hooks, no MCP servers, no telemetry.
12Limitations
- Small models. On Haiku 4.5 it still over-criticizes sound plans, skips the ledger, and reviews only the latest decision of a session.
- Our own cases. We wrote the eval cases, so treat the numbers as signals, not proof.
- Reasoning-only reviews. Without code or data to read, its conclusions top out at HYPOTHESIS, and it says so.
- Prior art needs search. If web search isn't available or allowed, prior art stays unchecked and any claim of novelty is a HYPOTHESIS.
13FAQ
Will it just say no to everything?
No. It names what your plan gets right before attacking it, credits the safeguards you already built in, and treats GO as the default when the facts support the plan. In the sound-plan eval case, on Sonnet 5.5 and Opus 5.5, it says GO even when the prompt asks it to be brutal.
Does it cost tokens when I'm not using it?
Only its short description, which every skill has so it can be triggered. The full instructions load when it runs.
Does it work in Spanish?
Yes. It answers in your language, including the confidence labels, and phrases like "no me des la razón" or "bajame a tierra" trigger it.
Will it change its mind if I push back?
Only if you bring a new fact or a better argument. Displeasure alone doesn't move it.
Is it a security tool?
No. It reviews ideas, plans and decisions, not systems: no penetration testing and no exploit work.