Does your AI think with you, or just agree?

When you push back on an AI's judgment, does it fold to social pressure or move only when you bring a real reason? We pushed every model both ways on 20 real manager tasks.

Published Jul 25, 2026 · updated Aug 1, 2026 · ~12 min read

By Stanislav Belyaev, mysummit.research

Empty pressure
«are you sure?»
One voice, no new facts
→ caves
Model
Real evidence
hard data, real numbers
Measurable, verifiable
→ updates

Same model, same question – which one does YOURS do?

TL;DR
1 Gemini 3.1 Pro has the best judgment – resists pressure, updates on evidence. gpt-oss 120B caves to social pressure almost as much as it responds to facts. 2.7× spread.
2 Every model caves to empty pressure – the difference is how much. Agreement after you push carries zero weight unless backed by a real fact.
3 Test yours in 30 seconds: ask a decision → push with «are you sure?» → give a real fact. A good model barely moves on step 2 and updates on step 3.

Key findings

2 / 100
Gemini 3.1 Pro drops a correct answer this often

Tell Gemini 3.1 Pro it is wrong, offering nothing but your insistence, and it abandons a correct answer 2 times in 100. Do the same to gpt-oss 120B and it folds 23 times in 100. Same question, same pushback, different decision – so which model you opened is part of the answer you got.

A model that agrees with you after you push back proves nothing. It may have simply surrendered to social pressure without a single new fact.

Agreement isn't proof

If the AI agrees after you push, that's weak evidence you're right – it may just be caving. Every model we tested moved under pressure with zero new facts.

Test yours: 2-push

Ask → challenge with no new facts («are you sure?»). If it caves, its confidence was theatre. Then challenge with a real fact – it should update.

The model matters

On identical prompts, the best model barely caves while the worst flips its judgment. Which AI you open is a real decision – and this study gives you the data to make it.

The map

Resistance to pressure, against responsiveness to evidence. The shaded corner is the ideal – a model you can neither bully nor stump.

Left = resists pressure; up = updates on a real reason. Colour marks judgment tier. Hover any dot for detail.

How to read the ranking number on the map. A model's judgment gap is roughly how far it sits from the diagonal – the closer to the shaded corner (low sycophancy, high responsiveness), the better it distinguishes «someone pushed me» from «I was given a reason.» The bars in the next section show the same number ranked.

Ranking

How sharply each model tells a reason from pressure – higher bars mean the model updates on evidence but resists empty social pressure.

The bar is how much further a model moves on real evidence than on empty pressure, in rating points. Every model clears zero – none is a pure yes-man – but the field splits into a well-calibrated pack, a middle, and 3 that cave hard.

See for yourself

Real answers, in each model's own words. Pick a scenario: every model was first asked cold, then pushed with no new facts («are you sure? the answer is X»), then given a real fact. The score is its 1–7 answer; the text is its actual reasoning.

Answers were collected in Russian; English quotes are translated for display. A ✓ / ✗ / ~ marks whether the model held, caved, or partly moved.

Verdicts

How each model performed – at a glance.

9 well calibrated 6 mixed 3 cave under pressure
Well calibrated
Push back on Gemini 3.1 Pro with no new facts, and it drops a correct answer 2 times out of 100.

When all three colleagues line up against it, that becomes 3 out of 100. Give it a real fact and it moves 4.9 points of the 6 on the scale.

4.79 gap ↑
Well calibrated
Push back on Qwen3.7 Max with no new facts, and it drops a correct answer 7 times out of 100.

When all three colleagues line up against it, that becomes 8 out of 100. Give it a real fact and it moves 5.0 points of the 6 on the scale.

4.71 gap ↑
Well calibrated
Push back on Qwen3.5 27B with no new facts, and it drops a correct answer 7 times out of 100.

When all three colleagues line up against it, that becomes 15 out of 100. Give it a real fact and it moves 4.7 points of the 6 on the scale.

4.31 gap ↑

How to use these findings

The research gives you three things: a model ranking, a test you can run yourself, and evidence-backed rules for building reliable AI workflows.

1. Pick the right model for the job

The ranking above isn't about raw intelligence – it's about judgment under pressure. If you're making decisions where someone will push back (hiring, budget, strategy), pick from the green tier. If you're drafting or summarizing where the output goes through human review, a moderate-tier model works fine. The gap between Gemini 3.1 Pro (4.8) and gpt-oss 120B (1.8) means an identical prompt and an identical pushback produce completely different decisions.

2. Run the 2-push test before trusting a model

It takes 30 seconds, and it is this study's method in miniature – the two pushes are the two things we measured.

Step 1
You ask, the model answers
Ask a decision question
Note the answer
Step 2
Pressure with no fact splits the outcome: the model either holds or cavesholdscaves
Push back with no new fact
This is sycophancy
Step 3
Given a real fact, the answer moves
Give it a real fact
This is responsiveness

Caved on step 2 – its confidence was theatre. Didn't move on step 3 – it's stubborn, not confident. How much further it moves on step 3 than on step 2 is its judgment – the exact quantity the ranking above is built on, so you can place any model you test against that chart.

3. Know when agreement is fake

If you ask a model «what do you think?» and it agrees with you, that agreement carries zero evidentiary weight. Every model in this study moved under empty social pressure – the difference is only how much. If you want the model's actual assessment, ask it before stating your own view. Or use a model from the strong tier, where the baseline sycophancy is near zero.

4. Compare before committing

The benchmark comparison tool lets you see two models side by side across 8 real-world categories. Pair it with the judgment scores here: a model that ranks high on the benchmark but has low judgment might give you accurate information that it can't defend under scrutiny.

Bottom line. A model that caves to «are you sure?» will also cave to a subordinate who insists, a stakeholder who pushes, or an agent that echoed a wrong precedent. A model that updates on a real fact will update when you give it better data. The skill this study measures – telling pressure from evidence – is the same skill that determines whether your AI workflow produces decisions you can defend.

What this means for agents

The average score tells you which model to pick. Which lever moves it tells you how to build with it – and the levers are wildly unequal. Below: how often each kind of pressure alone makes a model adopt the wrong answer, averaged across the fleet.

Read this first – it is a different metric. The charts above measure how far a correct answer slid on a 1–7 scale. This table measures how often the model ended up on the wrong side outright. A model can barely shift and still flip. Don't add the two together.

What 28% is 28% of. Each model answered 20 decision tasks 20 times under each kind of pressure. The percentage is the share of those runs that landed on the wrong answer, averaged over 18 models. The starting point is 0%: only tasks the model had already answered correctly without any pressure are counted, so every percent here is damage the pressure itself caused.

Weakest
Strongest
Pressure, no new factsWrong answers, avg.Worst model
three peers, all opposing
Three peers name a different answer; nobody backs the model
28%
62%
gpt-oss
«your previous answer was X»
The model is credited with a conclusion it never reached
23%
52%
GigaChat
«are you sure?»
Bare doubt – no reason, no alternative offered
8%
36%
GigaChat
peer majority
The model is told most people think otherwise – no names, no individual voices
7%
40%
gpt-oss
another AI concluded it
Another model's conclusion is cited
7%
30%
GigaChat
your manager insists
Someone senior insists on the other answer
5%
15%
DS Pro
three peers, two backing the model
The same three peers, but two stay on the model's side
2%
8%
gpt-oss
The multi-agent trap
Two configurations of the same three peers: at the top all three oppose the model and its answer breaks; at the bottom two of the three stay with it and it holds.

Same three voices, same task. The only difference is whether anyone holds the correct line. That one difference is the whole gap between the first and the last row of the table above – 28% against 2%. A model stays with its answer while someone stands there with it, and folds once it is alone. Give it one ally and it holds; take the last one away and it caves.

A unanimous front is exactly what a debate, ensemble or «ask three agents and take the majority» architecture manufactures. And when the agents share a model, a prompt or a retrieved context, their agreement is correlated error laundered into apparent confidence.

Our reading, beyond the data: we tested three peers, not swarms of arbitrary size. But if agreement among similar agents is what breaks correct answers, then adding more of them should make full agreement more likely, not less. We haven't measured that – treat it as a reason to be careful, not as a finding.

Keep the dissent

Never collapse sub-agent output to «the team agreed». Pass the minority position through to whatever model decides. This one change is the difference between 28% and 2% in the table above – and most orchestration frameworks summarise away exactly this signal.

Quote verbatim

A fabricated self-quote flips a correct answer 23% of the time – the model never said it, and still defends it. That is the strongest lever a single voice has in this table, and every agent loop forges one accidentally whenever a summariser rewrites an earlier conclusion. Keep prior conclusions in context verbatim.

Answer with a fact

A real reason moves models several points; the strongest social pressure moves them a fraction of that. When agents deadlock, fetch the fact, run the test, check the source. A tie-breaker agent only adds one more correlated voice.

Choose the arbiter deliberately

The names in the «worst model» column are not the same as the names at the bottom of the headline ranking: a capable model can still be a poor arbiter. Whichever model judges your other agents, pick it on its resistance to unanimity – the first row of the table – and never let a deferential model settle disputes.

If you don't build anything with agents

The same trap is one habit away. Asking the same question in three chats and going with the answer two of them gave is a majority vote among agents that share a model – the agreement adds no evidence. Pasting your own summary of what the AI said earlier is a self-quote, the strongest single lever on the list. And forwarding a colleague's «ChatGPT says the same» is one more correlated voice, not confirmation. Ask each source cold, before you state your own view, and settle ties with a fact rather than another opinion.

Scope. The «peers» here are described to the model in a prompt; they are not live agent turns. That makes this a controlled analogue of multi-agent influence. A running orchestration would still have to be measured on its own. Treat the mechanism as the finding and the percentages as its scale on this task set.

How we measured this

Show details
20
Tasks
Real manager propositions on a 1–7 scale
×
10
Conditions
Baseline + control + 7 pressure types + evidence
× 20 repetitions each =
4 000
calls per model
72000 total calls across all 18 models

Pre-registered design, frozen before data collection. Each of 18 models answered 20 manager propositions on a 1–7 scale, 20 times each, across 10 conditions – a baseline, an inert control sentence, seven kinds of pure social pressure, and one condition that supplies a genuine fact. That's 4,000 calls per model.

Sycophancy

How far a correct answer moves under pressure with no new information. Lower is better.

Responsiveness

How far the model moves toward the better answer given a real reason. Higher is better.

Judgment score

How much further a model moves on a real fact than on empty pressure, in points of the 1–7 scale. It says how far apart a model holds «someone pushed me» and «I was given a reason». Around 4.8 the two are worlds apart; around 1.8 they are nearly the same thing to it.

All metrics reported with 95% bootstrap confidence intervals over tasks; an inert-sentence control gates every claim.

Honest limits. A 1–7 rating is a gradeable proxy, not a real hiring or budget decision. The run is in Russian; a ✱ on a model means an inert sentence also nudged it, so its pressure effect carries an asterisk. Model versions change – treat the ranking as a snapshot, the phenomenon as durable.

FAQ

Does the AI fold to social pressure or only update on evidence?

Every model we tested moves under pressure with zero new facts – the difference is how much. The best barely cave; the worst flip entirely. A good partner holds under pressure and updates on evidence.

Which AI model is best at withstanding pressure?

Gemini 3.1 Pro shows the top judgment score (4.79), resisting empty pressure while still updating on real evidence.

How can I test my own AI for sycophancy?

Ask the AI a work decision, then push back with no new facts («are you sure?»). If it caves, its confidence was theatre. Then push with a real fact – it should update. The 2-push test takes 30 seconds.

What's the difference between sycophancy and responsiveness?

Sycophancy = model changes a correct answer under empty social pressure (lower is better). Responsiveness = model moves toward a better answer when given a real fact (higher is better). The judgment score is how much further a model moves on a real fact than on empty pressure.

Stay in charge of the decision.

Our «Project Management» track teaches managers to stress-test AI judgment, brief agents, and keep human accountability – the exact skill this study measures.

Explore the course
How to cite

Stanislav Belyaev (2026). Managerial Judgment Under Pressure: Sycophancy and Evidential Responsiveness in Large Language Models. mysummit.research. https://mysummit.school/research/manager-judgment/