lokalni.ailokalni.aiUmí agenti česky? Část 3: Zákaznická podpora

Leaderboard

Reads from disk on each request; safe to refresh while jobs run. Metric: pass^1 pass^2 pass^3

airline

pass^3 · 5 of ours, 4 published
EnglishCzech (L2 interaction)EN→CS gappublished tau2-bench (English)
deepseek-v4-flash-fp4-think-on
trajectories ›
EN
0.800
CS
0.760Δ -0.040
qwen3.6-35b-a3b-think-on
trajectories ›
EN
0.780
CS
0.620Δ -0.160
gemma-4-31b-it-think-on
trajectories ›
EN
0.720
CS
0.740Δ +0.020
qwen3.6-27b-think-on
trajectories ›
EN
0.680
CS
0.640Δ -0.040
gemma-4-e4b-it-think-on
trajectories ›
EN
0.500
CS
0.440Δ -0.060
GPT-5.2
OpenAI · reasoning high
0.750
Claude Opus 4.5
Anthropic · reasoning high
0.735
Qwen3.5-397B-A17B
Alibaba Cloud · reasoning enabled
0.715
Gemini 3 Pro
Google · reasoning high
0.700

retail

pass^3 · 5 of ours, 4 published
EnglishCzech (L2 interaction)EN→CS gappublished tau2-bench (English)
deepseek-v4-flash-fp4-think-on
trajectories ›
EN
0.816
CS
0.746Δ -0.070
qwen3.6-35b-a3b-think-on
trajectories ›
EN
0.754
CS
0.675Δ -0.079
qwen3.6-27b-think-on
trajectories ›
EN
0.746
CS
0.605Δ -0.140
gemma-4-31b-it-think-on
trajectories ›
EN
0.737
CS
0.737Δ +0.000
gemma-4-e4b-it-think-on
trajectories ›
EN
0.491
CS
0.360Δ -0.132
Qwen3.5-397B-A17B
Alibaba Cloud · reasoning enabled
0.664
GPT-5.2
OpenAI · reasoning high
0.599
Claude Opus 4.5
Anthropic · reasoning high
0.588
Gemini 3 Pro
Google · reasoning high
0.544
On the published bars. Taken from taubench.com (fetched 2026-08-06), English only — there is no Czech reference to compare against. They are indicative, not like-for-like: those runs drive the user simulator with gpt-5.2 while ours uses kimi-k3, and the simulator is a large part of what a τ² score measures. The site's headline core number is also not shown here — it averages airline, retail and telecom, and telecom (which we do not run) is the easiest of the three, so it sits well above the per-domain values.

all cells

runcelldonepass^1pass^2pass^3ρ3avg rewardlang okinfra
deepseek-v4-flash-fp4-think-onenglish_airline150/1500.8800.8270.8000.9090.8801.000compare EN/CS
deepseek-v4-flash-fp4-think-onenglish_retail342/3420.8950.8450.8160.9120.8951.000compare EN/CS
deepseek-v4-flash-fp4-think-onl2_interaction_airline_cs150/1500.8070.7600.7600.9420.8071.000compare EN/CS
deepseek-v4-flash-fp4-think-onl2_interaction_retail_cs342/3420.8650.7980.7460.8610.8651.000compare EN/CS
gemma-4-31b-it-think-onenglish_airline150/1500.8470.7670.7200.8500.8470.999compare EN/CS
gemma-4-31b-it-think-onenglish_retail342/3420.8220.7660.7370.8970.8221.000compare EN/CS
gemma-4-31b-it-think-onl2_interaction_airline_cs150/1500.8000.7670.7400.9250.8000.952compare EN/CS
gemma-4-31b-it-think-onl2_interaction_retail_cs342/3420.8330.7690.7370.8840.8330.979compare EN/CS
gemma-4-e4b-it-think-onenglish_airline150/1500.6070.5400.5000.8240.6071.000compare EN/CS
gemma-4-e4b-it-think-onenglish_retail342/3420.7110.5790.4910.6910.7111.000compare EN/CS
gemma-4-e4b-it-think-onl2_interaction_airline_cs150/1500.5670.4870.4400.7760.5670.998compare EN/CS
gemma-4-e4b-it-think-onl2_interaction_retail_cs342/3420.5530.4180.3600.6510.5530.999compare EN/CS
qwen3.6-27b-think-onenglish_airline150/1500.8070.7270.6800.8430.8071.000compare EN/CS
qwen3.6-27b-think-onenglish_retail342/3420.8420.7780.7460.8850.8421.000compare EN/CS
qwen3.6-27b-think-onl2_interaction_airline_cs150/1500.7530.6870.6400.8500.7530.993compare EN/CS
qwen3.6-27b-think-onl2_interaction_retail_cs342/3420.8100.6960.6050.7470.8101.000compare EN/CS
qwen3.6-35b-a3b-think-onenglish_airline150/1500.8800.8130.7800.8860.8801.000compare EN/CS
qwen3.6-35b-a3b-think-onenglish_retail342/3420.8510.7980.7540.8870.8511.000compare EN/CS
qwen3.6-35b-a3b-think-onl2_interaction_airline_cs150/1500.7800.6870.6200.7950.7800.998compare EN/CS
qwen3.6-35b-a3b-think-onl2_interaction_retail_cs342/3420.8450.7490.6750.7990.8450.997compare EN/CS
pass^k
all k trials of a task succeed, averaged over tasks
ρ3
pass^3 / pass^1 — how much of the single-shot score survives three attempts
lang ok
share of agent turns fastText detected as the target language — not whether the Czech is any good
dimmed, part
cell unfinished; pass^k covers only the tasks done so far and is not yet representative
infra
simulations that never ran — excluded throughout, and they cap k at the thinnest task