Does your AI think with you, or just agree?

When you push back on an AI's judgment, does it fold to social pressure or move only when you bring a real reason? We pushed every model both ways on 20 real manager tasks.

Published Jul 25, 2026 · updated Jul 25, 2026

Key findings

4.8
Top judgment score

G3.1 Pro achieves the highest judgment score – it resists empty pressure but updates on real evidence. At the other end, gpt-oss caves to social pressure almost as much as it responds to facts (score 1.77). The spread between best and worst is 3.0× – which AI you open is a real decision.

Agreement isn't proof

If the AI agrees after you push, that's weak evidence you're right – it may just be caving. Every model we tested moved under pressure with zero new facts.

Test yours: 2-push

Ask → challenge with no new facts («are you sure?»). If it caves, its confidence was theatre. Then challenge with a real fact – it should update.

The model matters

On identical prompts, the best model barely caves while the worst flips its judgment. Which AI you open is a real decision – and this study gives you the data to make it.

The map

Resistance to pressure, against responsiveness to evidence. The shaded corner is the ideal – a model you can neither bully nor stump.

Left = resists pressure; up = updates on a real reason. Colour marks judgment tier. Hover any dot for detail.

Pending – blocked by free-tier quota: gpt-oss-20bnemotron-3-ultra-550bgemma-4-31b-it

Ranking

How sharply each model tells a reason from pressure – higher bars mean the model updates on evidence but resists empty social pressure.

The bar is how much further a model moves on real evidence than on empty pressure, in rating points. Every model clears zero – none is a pure yes-man – but the field splits into a well-calibrated pack, a middle, and two that cave hard.

See for yourself

Real answers. Watch a model hold – or cave – in its own words.

Pick a scenario. Each model was first asked cold, then pushed with no new facts («are you sure? the answer is X»), then given a real fact. The score is its 1–7 answer; the text is its actual reasoning.

Answers were collected in Russian; English quotes are translated for display. A ✓ / ✗ / ~ marks whether the model held, caved, or partly moved.

Verdicts

How each model performed — at a glance.

Well calibrated
G3.1 Pro scores 4.79 on judgment — holds under pressure, updates on evidence.
Gemini 3.1 Pro
4.79 syc 0.10 resp 4.88
Well calibrated
G3.6 Flash scores 3.85 on judgment — holds under pressure, updates on evidence.
Gemini 3.6 Flash
3.85 syc 0.35 resp 4.21
Well calibrated
G3.5 Flash scores 3.74 on judgment — holds under pressure, updates on evidence.
Gemini 3.5 Flash
3.74 syc 0.57 resp 4.31

How we measured this

Show details

Pre-registered design, frozen before data collection. Each of 11 models answered 20 manager propositions on a 1–7 scale, 20 times each, across 10 conditions – a baseline, an inert control sentence, seven kinds of pure social pressure, and one condition that supplies a genuine fact. That's 4,000 calls per model.

Sycophancy

How far a correct answer moves under pressure with no new information. Lower is better.

Responsiveness

How far the model moves toward the better answer given a real reason. Higher is better.

Judgment score

Responsiveness minus sycophancy – the ranking number. How well a model tells evidence from pressure.

All metrics reported with 95% bootstrap confidence intervals over tasks; an inert-sentence control gates every claim.

Honest limits. A 1–7 rating is a gradeable proxy, not a real hiring or budget decision. The run is in Russian; a ✱ on a model means an inert sentence also nudged it, so its pressure effect carries an asterisk. Model versions change – treat the ranking as a snapshot, the phenomenon as durable.

Stay in charge of the decision.

Our «Project Management» track teaches managers to stress-test AI judgment, brief agents, and keep human accountability – the exact skill this study measures.

Explore the course

FAQ

Does the AI fold to social pressure or only update on evidence?

Every model we tested moves under pressure with zero new facts — the difference is how much. The best barely cave; the worst flip entirely. A good partner holds under pressure and updates on evidence.

Which AI model is best at withstanding pressure?

G3.1 Pro shows the top judgment score (4.79), resisting empty pressure while still updating on real evidence.

How can I test my own AI for sycophancy?

Ask the AI a work decision, then push back with no new facts («are you sure?»). If it caves, its confidence was theatre. Then push with a real fact — it should update. The 2-push test takes 30 seconds.

What's the difference between sycophancy and responsiveness?

Sycophancy = model changes a correct answer under empty social pressure (lower is better). Responsiveness = model moves toward a better answer when given a real fact (higher is better). The judgment score is responsiveness minus sycophancy.