instinct.

AN OPEN LOOK AT MODEL PERFORMANCE

Intelligence,
measured.

Real evaluations. Clear methodology.
Find the right model, then put it to the test.

Public results. Reproducible by design.

EVALUATION / 001Public subset
JEV-LITE / JEVBENCH
78.79%
Accuracy
182 / 231 correct
231public evaluation cases
3difficulty tiers
3models on one API
100%serial requests succeeded

01 / THE BENCHMARK

Every result has a receipt.

JevBench evaluates choices, binary decisions, and ordered scores.

Download results ↓
jev-lite · 1 × H200 · BF16
Overall accuracy
78.79%

182 correct decisions across
231 public evaluation cases.

COMPLETE RUN

One fixed serial baseline.
All published public cases included.

Evaluated 21 September 2026
Jev-lite accuracy by JevBench difficulty tier
Difficulty Correct / cases Accuracy Accuracy · 0–100%
Scope: 231 public cases, not the full 534-case leaderboard. No official rank or composite score is claimed.Read the method ↗

SOURCE Instinct evaluation record, 21 Sep 2026·JevBench a1799db ↗

02 / THE MODELS

One endpoint. Different strengths.

Use the model that fits your task, with the same API key.

Catalog verified · 22 Sep 2026
GENERAL PURPOSE

DeepSeek V4.1 Flash

Text generation and multi-turn workflows through the Chat Completions API.

Not evaluated on this
published JevBench run
DeepSeek-V4.1-Flash
GENERAL PURPOSE

Qwen3.8 27B

A 27B model for text tasks, available through the same model API.

Not evaluated on this
published JevBench run
Qwen/Qwen3.8-27B

03 / THE METHOD

Make the result
reproducible.

A useful benchmark is more than a number. You should be able to inspect how it was measured and run it yourself.

Full evaluation notes ↗
01

Pin the evaluation.

JevBench commit a1799db. All 48 easy, 72 original, and 111 hard public cases. The private cases are outside this evaluation.

02

Keep the question intact.

Original instructions, rubrics, and candidate order are preserved. Ground truth stays in the local scorer and is never sent to the model.

03

Count the entire run.

Use the upstream scorer and a fixed serial baseline. Failures count as incorrect. No retries, prompt truncation, or synthetic probabilities.

04

Report the conditions.

Latency includes the public HTTPS round trip. The load test reuses inputs; caching and concurrency can change performance and close decisions.

04 / YOUR TURN

From evidence to your own experiment.

Use your separately provided API key. No website account required.

Ask a structured question.

Set INSTINCT_API_KEY in your local environment. Keys are shared separately and are never embedded in this website.

POST /v1/chat/completionsBearer authentication
REPRODUCE THE PUBLISHED RUN

Your key. The same 231 cases.

Download the client, pin the dataset, and score responses locally. The script saves a manifest, per-request evidence, and aggregate results.

Download jev-lite benchmark client ↓
TERMINAL

Good models deserve
good evidence.