jev-lite
Choices, binary judgments, and ordered scores. Returns native candidate probabilities.
231 cases · serial baseline
jev-lite
AN OPEN LOOK AT MODEL PERFORMANCE
Real evaluations. Clear methodology.
Find the right model, then
put it to the test.
Public results. Reproducible by design.
01 / THE BENCHMARK
JevBench evaluates choices, binary decisions, and ordered scores.
182 correct decisions across
231 public evaluation cases.
One fixed serial baseline.
All published public cases
included.
| Difficulty | Correct / cases | Accuracy | Accuracy · 0–100% |
|---|
One full pass of the same 231 cases at each setting.
40.82 requests/s over 2,310 requests at concurrency 64. Zero HTTP failures in this 56.6-second run.
SOURCE Instinct evaluation record, 21 Sep 2026·JevBench a1799db ↗
02 / THE MODELS
Use the model that fits your task, with the same API key.
Choices, binary judgments, and ordered scores. Returns native candidate probabilities.
jev-lite
Text generation and multi-turn workflows through the Chat Completions API.
DeepSeek-V4.1-Flash
A 27B model for text tasks, available through the same model API.
Qwen/Qwen3.8-27B
03 / THE METHOD
A useful benchmark is more than a number. You should be able to inspect how it was measured and run it yourself.
Full evaluation notes ↗
JevBench commit a1799db. All 48 easy, 72
original, and 111 hard public cases. The private cases are
outside this evaluation.
Original instructions, rubrics, and candidate order are preserved. Ground truth stays in the local scorer and is never sent to the model.
Use the upstream scorer and a fixed serial baseline. Failures count as incorrect. No retries, prompt truncation, or synthetic probabilities.
Latency includes the public HTTPS round trip. The load test reuses inputs; caching and concurrency can change performance and close decisions.
04 / YOUR TURN
Use your separately provided API key. No website account required.
Set INSTINCT_API_KEY in your local environment. Keys
are shared separately and are never embedded in this website.
Download the client, pin the dataset, and score responses locally. The script saves a manifest, per-request evidence, and aggregate results.
Download jev-lite benchmark client ↓
Good models deserve
good evidence.