What LiveCodeBench Actually Measures
Algorithmic problem solving, complex mathematical reasoning, and edge-case handling on brand-new competitive programming problems published after model training cutoff dates.
⚠ Known Boundaries & Harness Sensitivity
Measures competitive programming and algorithmic puzzles rather than real-world software maintenance or multi-file application development.
Complete Verified Leaderboard
Fully server-side rendered table with verified primary source attributions and direct repository links.
No verified results loaded yet — ingestion pending
We only publish benchmark scores with verified primary source provenance and public methodology. Evaluations for this benchmark are currently in queue.
In compliance with upstream dataset licenses, all evaluated telemetry rows are credited to their primary sources:
-
Manual / Direct Evaluation Ingest
—
Licence:
Editorial
What This Benchmark Does Not Measure
Does not measure real-world repo navigation or git workflows.
How to Reproduce This Evaluation
All evaluations published by Vibecoder Journal follow publicly documented harness specifications. You can execute this test suite against any local or API model:
# Standardized containerized evaluation execution
git clone https://livecodebench.github.io/ && cd evaluation
docker build -t vcj-harness .
docker run --rm -v $(pwd)/results:/results vcj-harness --benchmark livecodebench
In compliance with upstream dataset licenses, all evaluated telemetry rows are credited to their primary sources:
-
Aider Polyglot Leaderboard
—
Licence:
Apache-2.0 -
LMSYS Chatbot Arena (WebDev)
—
Licence:
CC-BY-4.0 -
Manual / Direct Evaluation Ingest
—
Licence:
Editorial -
Models.dev API
—
Licence:
MIT -
OpenRouter Models Catalogue
—
Licence:
Platform-ToS -
SWE-bench Official Leaderboard
—
Licence:
MIT -
Terminal-Bench Submissions
—
Licence:
Apache-2.0