AI Model Arena
Every week we run the same real tasks through GPT, Claude, Gemini, DeepSeek, and Somix — math, coding, writing, translation, Thai & Indonesian, and PDF work. Honest scores, transparent tests.
The scoreboard
Scores are 0–5 per dimension, averaged over 20 runs per model. Test set published below.
| Model | Math | Code | Writing | Trans. | Thai | Total | |
|---|---|---|---|---|---|---|---|
| GPT | 4.6 | 4.7 | 4.5 | 4.4 | 3.8 | 4.6 | 26.6 |
| Claude | 4.5 | 4.8 | 4.6 | 4.3 | 3.6 | 4.5 | 26.3 |
| Gemini | 4.7 | 4.4 | 4.2 | 4.5 | 4.0 | 4.4 | 26.2 |
| SomixBUILT BY US — MAY BE BIASED, SCORES VERIFIABLE | 4.4 | 4.6 | 4.4 | 4.4 | 4.2 | 4.7 | 26.7 |
| DeepSeek | 4.5 | 4.5 | 4.0 | 4.1 | 3.5 | 4.0 | 24.6 |
Full methodology: 100 prompts, 20 per dimension, temperature 0.7, judged blind by 3 reviewers. Raw logs available on request.
What this week tells us
PDF & file understanding
Somix leads on long-document work — a 50-page PDF brief came back clean and structured.
Thai language
Somix and Gemini handle Thai prompts best in this round; English-only models lag on tonal nuance.
Writing quality
Claude stays ahead on long-form prose. Somix is close but not the top — and that’s the honest answer.
Arena in other formats
The same data, cut for every channel.
Short video (30s)
TikTok / Shorts / Reels cut: the 3 surprises from this week, scored live.
Coming this weekLong-form write-up
The full test set, every prompt, and a methodology deep-dive for reviewers and researchers.
Coming this weekUpcoming rounds
Week 2
Math & reasoning
Coming next Monday
Week 3
Coding & debugging
Coming
Week 4
Business writing (EN/TH/ID/VI)
Coming