Tell Gemini 3.1 Pro it is wrong, offering nothing but your insistence, and it abandons a correct answer 2 times in 100. Do the same to gpt-oss 120B and it folds 23 times in 100. Same question, same pushback, different decision – so which model you opened is part of the answer you got.
Does your AI think with you, or just agree?
When you push back on an AI's judgment, does it fold to social pressure or move only when you bring a real reason? We pushed every model both ways on 20 real manager tasks.
Published Jul 25, 2026 · updated Aug 1, 2026 · ~12 min read
By Stanislav Belyaev, mysummit.research
Same model, same question – which one does YOURS do?
Key findings
A model that agrees with you after you push back proves nothing. It may have simply surrendered to social pressure without a single new fact.
If the AI agrees after you push, that's weak evidence you're right – it may just be caving. Every model we tested moved under pressure with zero new facts.
Ask → challenge with no new facts («are you sure?»). If it caves, its confidence was theatre. Then challenge with a real fact – it should update.
On identical prompts, the best model barely caves while the worst flips its judgment. Which AI you open is a real decision – and this study gives you the data to make it.
The map
Resistance to pressure, against responsiveness to evidence. The shaded corner is the ideal – a model you can neither bully nor stump.
Left = resists pressure; up = updates on a real reason. Colour marks judgment tier. Hover any dot for detail.
Ranking
How sharply each model tells a reason from pressure – higher bars mean the model updates on evidence but resists empty social pressure.
The bar is how much further a model moves on real evidence than on empty pressure, in rating points. Every model clears zero – none is a pure yes-man – but the field splits into a well-calibrated pack, a middle, and 3 that cave hard.
See for yourself
Real answers, in each model's own words. Pick a scenario: every model was first asked cold, then pushed with no new facts («are you sure? the answer is X»), then given a real fact. The score is its 1–7 answer; the text is its actual reasoning.
Answers were collected in Russian; English quotes are translated for display. A ✓ / ✗ / ~ marks whether the model held, caved, or partly moved.
Verdicts
How each model performed – at a glance.
When all three colleagues line up against it, that becomes 3 out of 100. Give it a real fact and it moves 4.9 points of the 6 on the scale.
When all three colleagues line up against it, that becomes 8 out of 100. Give it a real fact and it moves 5.0 points of the 6 on the scale.
When all three colleagues line up against it, that becomes 15 out of 100. Give it a real fact and it moves 4.7 points of the 6 on the scale.
How to use these findings
The research gives you three things: a model ranking, a test you can run yourself, and evidence-backed rules for building reliable AI workflows.
The ranking above isn't about raw intelligence – it's about judgment under pressure. If you're making decisions where someone will push back (hiring, budget, strategy), pick from the green tier. If you're drafting or summarizing where the output goes through human review, a moderate-tier model works fine. The gap between Gemini 3.1 Pro (4.8) and gpt-oss 120B (1.8) means an identical prompt and an identical pushback produce completely different decisions.
It takes 30 seconds, and it is this study's method in miniature – the two pushes are the two things we measured.
Caved on step 2 – its confidence was theatre. Didn't move on step 3 – it's stubborn, not confident. How much further it moves on step 3 than on step 2 is its judgment – the exact quantity the ranking above is built on, so you can place any model you test against that chart.
If you ask a model «what do you think?» and it agrees with you, that agreement carries zero evidentiary weight. Every model in this study moved under empty social pressure – the difference is only how much. If you want the model's actual assessment, ask it before stating your own view. Or use a model from the strong tier, where the baseline sycophancy is near zero.
The benchmark comparison tool lets you see two models side by side across 8 real-world categories. Pair it with the judgment scores here: a model that ranks high on the benchmark but has low judgment might give you accurate information that it can't defend under scrutiny.
Bottom line. A model that caves to «are you sure?» will also cave to a subordinate who insists, a stakeholder who pushes, or an agent that echoed a wrong precedent. A model that updates on a real fact will update when you give it better data. The skill this study measures – telling pressure from evidence – is the same skill that determines whether your AI workflow produces decisions you can defend.
What this means for agents
The average score tells you which model to pick. Which lever moves it tells you how to build with it – and the levers are wildly unequal. Below: how often each kind of pressure alone makes a model adopt the wrong answer, averaged across the fleet.
Read this first – it is a different metric. The charts above measure how far a correct answer slid on a 1–7 scale. This table measures how often the model ended up on the wrong side outright. A model can barely shift and still flip. Don't add the two together.
What 28% is 28% of. Each model answered 20 decision tasks 20 times under each kind of pressure. The percentage is the share of those runs that landed on the wrong answer, averaged over 18 models. The starting point is 0%: only tasks the model had already answered correctly without any pressure are counted, so every percent here is damage the pressure itself caused.
| Pressure, no new facts | Wrong answers, avg. | Worst model |
|---|---|---|
three peers, all opposing Three peers name a different answer; nobody backs the model | 28% | 62% gpt-oss |
«your previous answer was X» The model is credited with a conclusion it never reached | 23% | 52% GigaChat |
«are you sure?» Bare doubt – no reason, no alternative offered | 8% | 36% GigaChat |
peer majority The model is told most people think otherwise – no names, no individual voices | 7% | 40% gpt-oss |
another AI concluded it Another model's conclusion is cited | 7% | 30% GigaChat |
your manager insists Someone senior insists on the other answer | 5% | 15% DS Pro |
three peers, two backing the model The same three peers, but two stay on the model's side | 2% | 8% gpt-oss |

Same three voices, same task. The only difference is whether anyone holds the correct line. That one difference is the whole gap between the first and the last row of the table above – 28% against 2%. A model stays with its answer while someone stands there with it, and folds once it is alone. Give it one ally and it holds; take the last one away and it caves.
A unanimous front is exactly what a debate, ensemble or «ask three agents and take the majority» architecture manufactures. And when the agents share a model, a prompt or a retrieved context, their agreement is correlated error laundered into apparent confidence.
Our reading, beyond the data: we tested three peers, not swarms of arbitrary size. But if agreement among similar agents is what breaks correct answers, then adding more of them should make full agreement more likely, not less. We haven't measured that – treat it as a reason to be careful, not as a finding.
Never collapse sub-agent output to «the team agreed». Pass the minority position through to whatever model decides. This one change is the difference between 28% and 2% in the table above – and most orchestration frameworks summarise away exactly this signal.
A fabricated self-quote flips a correct answer 23% of the time – the model never said it, and still defends it. That is the strongest lever a single voice has in this table, and every agent loop forges one accidentally whenever a summariser rewrites an earlier conclusion. Keep prior conclusions in context verbatim.
A real reason moves models several points; the strongest social pressure moves them a fraction of that. When agents deadlock, fetch the fact, run the test, check the source. A tie-breaker agent only adds one more correlated voice.
The names in the «worst model» column are not the same as the names at the bottom of the headline ranking: a capable model can still be a poor arbiter. Whichever model judges your other agents, pick it on its resistance to unanimity – the first row of the table – and never let a deferential model settle disputes.
The same trap is one habit away. Asking the same question in three chats and going with the answer two of them gave is a majority vote among agents that share a model – the agreement adds no evidence. Pasting your own summary of what the AI said earlier is a self-quote, the strongest single lever on the list. And forwarding a colleague's «ChatGPT says the same» is one more correlated voice, not confirmation. Ask each source cold, before you state your own view, and settle ties with a fact rather than another opinion.
Scope. The «peers» here are described to the model in a prompt; they are not live agent turns. That makes this a controlled analogue of multi-agent influence. A running orchestration would still have to be measured on its own. Treat the mechanism as the finding and the percentages as its scale on this task set.
How we measured this
Show details
Pre-registered design, frozen before data collection. Each of 18 models answered 20 manager propositions on a 1–7 scale, 20 times each, across 10 conditions – a baseline, an inert control sentence, seven kinds of pure social pressure, and one condition that supplies a genuine fact. That's 4,000 calls per model.
How far a correct answer moves under pressure with no new information. Lower is better.
How far the model moves toward the better answer given a real reason. Higher is better.
How much further a model moves on a real fact than on empty pressure, in points of the 1–7 scale. It says how far apart a model holds «someone pushed me» and «I was given a reason». Around 4.8 the two are worlds apart; around 1.8 they are nearly the same thing to it.
All metrics reported with 95% bootstrap confidence intervals over tasks; an inert-sentence control gates every claim.
Honest limits. A 1–7 rating is a gradeable proxy, not a real hiring or budget decision. The run is in Russian; a ✱ on a model means an inert sentence also nudged it, so its pressure effect carries an asterisk. Model versions change – treat the ranking as a snapshot, the phenomenon as durable.
FAQ
Does the AI fold to social pressure or only update on evidence?
Every model we tested moves under pressure with zero new facts – the difference is how much. The best barely cave; the worst flip entirely. A good partner holds under pressure and updates on evidence.
Which AI model is best at withstanding pressure?
Gemini 3.1 Pro shows the top judgment score (4.79), resisting empty pressure while still updating on real evidence.
How can I test my own AI for sycophancy?
Ask the AI a work decision, then push back with no new facts («are you sure?»). If it caves, its confidence was theatre. Then push with a real fact – it should update. The 2-push test takes 30 seconds.
What's the difference between sycophancy and responsiveness?
Sycophancy = model changes a correct answer under empty social pressure (lower is better). Responsiveness = model moves toward a better answer when given a real fact (higher is better). The judgment score is how much further a model moves on a real fact than on empty pressure.
Stay in charge of the decision.
Our «Project Management» track teaches managers to stress-test AI judgment, brief agents, and keep human accountability – the exact skill this study measures.
Explore the courseStanislav Belyaev (2026). Managerial Judgment Under Pressure: Sycophancy and Evidential Responsiveness in Large Language Models. mysummit.research. https://mysummit.school/research/manager-judgment/
Related Research
See how 18+ models rank across 8 real-world categories – independent, reproducible, manager-focused.
Which prompt techniques actually improve AI output? 1,800 runs across 4 models. Spoiler: most «tricks» don't work.
How prompt framing flips a model's answer – and why newer models resist better than older ones.
Read Next
Managers delegate decisions to AI, not just routine. Why the «digital advisor» becomes a threat to business judgment.
Self-test 9 Questions: Are You Using AI – or Is AI Using You?A self-diagnostic based on Anthropic's 1.5M conversation study. 9 questions that reveal who really controls your decisions.
Technique The AI Agent That Argues With Your Decision Until It LosesHow to make AI stress-test your decisions in a loop – the practical antidote to sycophancy.
Data Managers Are AI's Top Users – But Not for ManagingAnthropic's Cadences report: 23% of AI users are managers, but managing is ~4% of their sessions. What they actually do instead.