lokalni.ailokalni.aiUmí agenti česky? Část 3: Zákaznická podpora

Leaderboard

Reads from disk on each request; safe to refresh while jobs run. Metric: pass^1 pass^2 pass^3

airline

pass^2 · 5 of ours, 4 published
EnglishCzech (L2 interaction)EN→CS gappublished tau2-bench (English)
deepseek-v4-flash-fp4-think-on
trajectories ›
EN
0.827
CS
0.760Δ -0.067
qwen3.6-35b-a3b-think-on
trajectories ›
EN
0.813
CS
0.687Δ -0.127
gemma-4-31b-it-think-on
trajectories ›
EN
0.767
CS
0.767Δ +0.000
qwen3.6-27b-think-on
trajectories ›
EN
0.727
CS
0.687Δ -0.040
gemma-4-e4b-it-think-on
trajectories ›
EN
0.540
CS
0.487Δ -0.053
GPT-5.2
OpenAI · reasoning high
0.783
Claude Opus 4.5
Anthropic · reasoning high
0.777
Qwen3.5-397B-A17B
Alibaba Cloud · reasoning enabled
0.757
Gemini 3 Pro
Google · reasoning high
0.747

retail

pass^2 · 5 of ours, 4 published
EnglishCzech (L2 interaction)EN→CS gappublished tau2-bench (English)
deepseek-v4-flash-fp4-think-on
trajectories ›
EN
0.845
CS
0.798Δ -0.047
qwen3.6-35b-a3b-think-on
trajectories ›
EN
0.798
CS
0.749Δ -0.050
qwen3.6-27b-think-on
trajectories ›
EN
0.778
CS
0.696Δ -0.082
gemma-4-31b-it-think-on
trajectories ›
EN
0.766
CS
0.769Δ +0.003
gemma-4-e4b-it-think-on
trajectories ›
EN
0.579
CS
0.418Δ -0.161
Qwen3.5-397B-A17B
Alibaba Cloud · reasoning enabled
0.744
GPT-5.2
OpenAI · reasoning high
0.696
Claude Opus 4.5
Anthropic · reasoning high
0.674
Gemini 3 Pro
Google · reasoning high
0.635
On the published bars. Taken from taubench.com (fetched 2026-08-06), English only — there is no Czech reference to compare against. They are indicative, not like-for-like: those runs drive the user simulator with gpt-5.2 while ours uses kimi-k3, and the simulator is a large part of what a τ² score measures. The site's headline core number is also not shown here — it averages airline, retail and telecom, and telecom (which we do not run) is the easiest of the three, so it sits well above the per-domain values.

all cells

runcelldonepass^1pass^2pass^3ρ3avg rewardlang okinfra
deepseek-v4-flash-fp4-think-onenglish_airline150/1500.8800.8270.8000.9090.8801.000compare EN/CS
deepseek-v4-flash-fp4-think-onenglish_retail342/3420.8950.8450.8160.9120.8951.000compare EN/CS
deepseek-v4-flash-fp4-think-onl2_interaction_airline_cs150/1500.8070.7600.7600.9420.8071.000compare EN/CS
deepseek-v4-flash-fp4-think-onl2_interaction_retail_cs342/3420.8650.7980.7460.8610.8651.000compare EN/CS
gemma-4-31b-it-think-onenglish_airline150/1500.8470.7670.7200.8500.8470.999compare EN/CS
gemma-4-31b-it-think-onenglish_retail342/3420.8220.7660.7370.8970.8221.000compare EN/CS
gemma-4-31b-it-think-onl2_interaction_airline_cs150/1500.8000.7670.7400.9250.8000.952compare EN/CS
gemma-4-31b-it-think-onl2_interaction_retail_cs342/3420.8330.7690.7370.8840.8330.979compare EN/CS
gemma-4-e4b-it-think-onenglish_airline150/1500.6070.5400.5000.8240.6071.000compare EN/CS
gemma-4-e4b-it-think-onenglish_retail342/3420.7110.5790.4910.6910.7111.000compare EN/CS
gemma-4-e4b-it-think-onl2_interaction_airline_cs150/1500.5670.4870.4400.7760.5670.998compare EN/CS
gemma-4-e4b-it-think-onl2_interaction_retail_cs342/3420.5530.4180.3600.6510.5530.999compare EN/CS
qwen3.6-27b-think-onenglish_airline150/1500.8070.7270.6800.8430.8071.000compare EN/CS
qwen3.6-27b-think-onenglish_retail342/3420.8420.7780.7460.8850.8421.000compare EN/CS
qwen3.6-27b-think-onl2_interaction_airline_cs150/1500.7530.6870.6400.8500.7530.993compare EN/CS
qwen3.6-27b-think-onl2_interaction_retail_cs342/3420.8100.6960.6050.7470.8101.000compare EN/CS
qwen3.6-35b-a3b-think-onenglish_airline150/1500.8800.8130.7800.8860.8801.000compare EN/CS
qwen3.6-35b-a3b-think-onenglish_retail342/3420.8510.7980.7540.8870.8511.000compare EN/CS
qwen3.6-35b-a3b-think-onl2_interaction_airline_cs150/1500.7800.6870.6200.7950.7800.998compare EN/CS
qwen3.6-35b-a3b-think-onl2_interaction_retail_cs342/3420.8450.7490.6750.7990.8450.997compare EN/CS
pass^k
all k trials of a task succeed, averaged over tasks
ρ3
pass^3 / pass^1 — how much of the single-shot score survives three attempts
lang ok
share of agent turns fastText detected as the target language — not whether the Czech is any good
dimmed, part
cell unfinished; pass^k covers only the tasks done so far and is not yet representative
infra
simulations that never ran — excluded throughout, and they cap k at the thinnest task