Tool Comparison

40 GigaChat Case Studies vs the Benchmark: Checking Sber's Numbers

23 min read

Sber, Russia’s largest bank and the company behind GigaChat, released a sponsored showcase: forty business cases from companies that deployed GigaChat and reported the results. EdTech, MedTech, HRTech, cybersecurity, PropTech. Polished cards, concrete numbers, real startups.

Sber’s promotional project

On the image: the “One step ahead” promo slide from the Sber500×GigaChat accelerator – 40 startups across 9 industries. Claimed effects: business processes up to x16 faster, costs down by up to 90%, up to 95% task automation, and revenue up by up to 30%.

We have a benchmark of our own: 29 models, 4,308 independent evaluations on managerial tasks. In it, GigaChat sits dead last – 29th out of 29 after the second wave of testing. That creates an interesting situation.

Not because Sber is lying. The cases are real, the startups exist, the automation works. The question is different: was this the optimal model for the tasks they were solving?

Read more
40 GigaChat Case Studies vs the Benchmark: Checking Sber's Numbers
How to Get the Most Out of YandexGPT: What Works and What Doesn't
13 min

How to Get the Most Out of YandexGPT: What Works and What Doesn't

Millions of people in Russia use Alice every day – not because they choose to, but because it’s free, built into Yandex Browser, and works without a VPN. YandexGPT, the model under Alice’s hood, is the best Russian model in our benchmark, but it’s still a long way behind GPT-5.4.

Can you get answers from it that come close to GPT, if you learn how to ask the right way? We tested exactly that in an experiment: ten prompting techniques, six management tasks, two independent LLM judges. The short answer: yes, you can – but not every technique works, and some make things worse.

Below are the concrete templates you can copy into the chat right now, and the anti-patterns to steer clear of.

GigaChat Ultra Thinking: Thinks Longer – Answers Worse?
7 min

GigaChat Ultra Thinking: Thinks Longer – Answers Worse?

GigaChat Ultra Thinking takes longer to think and uses more compute. It solves management tasks 3.3% worse than the version without reasoning. This is not a bug or a fluke – it’s a pattern documented in academic papers over the past two years.

This week, Sber unveiled GigaChat Ultra – a new flagship model with a reasoning mode (Thinking). The model is available for free via web, mobile apps, and a Telegram bot. We immediately added both variants to our AI model research for managers: ran them through all 32 scenarios using our unified methodology, scored them with both LLM judges, and compared against the other 52 models.

Chat Z.AI (GLM-5) Review 2026: Pricing, Benchmarks & Agent Mode
15 min

Chat Z.AI (GLM-5) Review 2026: Pricing, Benchmarks & Agent Mode

On February 6, 2026, an anonymous model called “Pony Alpha” appeared on OpenRouter – free, with zero details about its creators. The AI community immediately set about identifying it. Its coding abilities came remarkably close to Claude Opus 4.5. When asked “who are you?”, the model responded: “I am GLM.” But when prompted to write a web page describing itself – it wrote: “I am Claude, created by Anthropic.”