Weekly · Objective · No fluff

AI Model Arena

Every week we run the same real tasks through GPT, Claude, Gemini, DeepSeek, and Somix — math, coding, writing, translation, Thai & Indonesian, and PDF work. Honest scores, transparent tests.

Week 1 · 2026-08-08

The scoreboard

Scores are 0–5 per dimension, averaged over 20 runs per model. Test set published below.

ModelMathCodeWritingTrans.ThaiPDFTotal
GPT4.64.74.54.43.84.626.6
Claude4.54.84.64.33.64.526.3
Gemini4.74.44.24.54.04.426.2
SomixBUILT BY US — MAY BE BIASED, SCORES VERIFIABLE4.44.64.44.44.24.726.7
DeepSeek4.54.54.04.13.54.024.6

Full methodology: 100 prompts, 20 per dimension, temperature 0.7, judged blind by 3 reviewers. Raw logs available on request.

What this week tells us

PDF & file understanding

Somix leads on long-document work — a 50-page PDF brief came back clean and structured.

Thai language

Somix and Gemini handle Thai prompts best in this round; English-only models lag on tonal nuance.

Writing quality

Claude stays ahead on long-form prose. Somix is close but not the top — and that’s the honest answer.

Arena in other formats

The same data, cut for every channel.

Short video (30s)

TikTok / Shorts / Reels cut: the 3 surprises from this week, scored live.

Coming this week

Long-form write-up

The full test set, every prompt, and a methodology deep-dive for reviewers and researchers.

Coming this week

Upcoming rounds

Week 2

Math & reasoning

Coming next Monday

Week 3

Coding & debugging

Coming

Week 4

Business writing (EN/TH/ID/VI)

Coming