Aggregate methods for the September 19, 2026 lead-scoring test. This is not an independent replication or a public raw dataset.

Scope and reference labels

The supplied run used 83 contacts, selected toward decision-band boundaries, each scored for two B2B brands: a design subscription (Brand A) and an AI marketing tool (Brand B). The gold reference is adjudicated consensus of three independent model agents, as explicitly confirmed in the originating session. It is not human ground truth or an outcome label such as a purchase. “Send” means a fit decision under the reference rubric, not permission to send outreach.

Scores of 7 or more map to send, 5–6 to maybe, and below 5 to no. Band agreement is the fraction matching the reference band. Send precision is the fraction of predicted sends that the reference labels send. Send recall is the fraction of reference sends recovered. Send↔no flips skip the intermediate maybe band. Self-agreement compares Brand A bands across two runs where available.

Comparison conditions

The short arm uses the same roughly 150-word brief and five-level rubric, with Jev's typed-question interface and each baseline's own response schema. One request per row, four concurrent requests, generative temperature zero and provider-default reasoning. Full arms use an 11.6k-character prompt tuned on this gold set; they are not a fair held-out test against the short arm. The GPT-5.6 Sol reference ran in a batch agent CLI, not the same per-row API setup. Its total tokens and duration are not a per-row price or latency estimate.

For the short generative arms, five-level scores map 0–1 to no, 2 to maybe, and 3–4 to send. Jev returns probabilities over those five levels: the evaluator sums levels 0+1 and 3+4, then selects the band with greatest probability mass. It does not threshold the probability-weighted score. This interface difference is part of the comparison.

Costs are provider-reported where available, otherwise list price multiplied by token counts. Jev uses input tokens × $0.042 per million. Reported figures are rounded. They exclude source acquisition, enrichment, review, integration and downstream campaigns. Latency is per-call p50 unless explicitly labeled batch total.

The table identifies requested models as recorded in the run configuration. The saved provider receipts omit returned model identifiers, so the labels do not establish immutable backend revisions. Short and full prompt arms are listed separately. GPT-5.6 Sol via Codex is a separate batch reference, not a like-for-like per-record API comparator. Brand identities and raw contact data remain private. Naming the models improves interpretation, but this small, private sample cannot establish general vendor rankings. These are the supplied aggregate results, not a new inference run.

The historical v2 pipeline retained three votes for 48 contacts and one vote for 35 contacts (179 retained votes across 83 rows). The retained rows do not establish the model identity of that historical execution. We therefore keep the historical v1 and v2 systems unnamed. Retained vote counts are not a complete billed-call ledger.

Aggregate table

Model (prompt) Brand A band A send prec / rec A send↔no flips Brand B band B send prec / rec Cost, 83 rows p50 latency Self-agree (Brand A band, 2 runs)
jev-1.13.0 (short) 77% 82 / 60 6 76% 100 / 44 $0.0055 0.95s 98%
deepseek-v4-flash (short) 72% 52 / 73 4 73% 62 / 62 $0.0109 8.4s 86%
deepseek-v4-flash (full) 75% 73 / 73 3 73% 100 / 69 $0.0356 20.6s —
deepseek-v4.1-flash (short) 81% 80 / 80 1 77% 77 / 62 $0.1834 30.5s —
qwen3.8-flash (short) 77% 92 / 80 0 67% 73 / 50 $0.1463 52.3s —
mercury-2.5 (short) 73% 80 / 53 3 70% 82 / 56 $0.0293 5.4s 84%
mercury-2.5 (full) 75% 87 / 87 0 60% 71 / 75 $0.0376 8.3s —
gpt-5.6-luna (short) 67% 68 / 87 3 76% 100 / 44 $0.0351 4.2s 87%
gpt-5.6-luna (full) 65% 75 / 80 1 57% 87 / 81 $0.1022 11.4s —
glm-5.3-flash (short) 65% 58 / 93 2 80% 71 / 75 $0.0583 25.4s —
separate batch reference: gpt-5.6-sol via codex (full, batch) 77% 83 / 100 1 55% 72 / 81 120k tokens 7m42s total —
existing: v2 LLM adaptive 1-or-3 votes (full) 75% 65 / 73 5 81% 88 / 88 179 retained votes — 80% single-call
existing: v1 LLM single 66% 42 / 67 7 80% 79 / 69 — — —

Additional observations and their limits

The supplied run reports Brand A Jev confidence at least 0.6 covering 48% of rows with 92% band agreement. At least 0.7 covered 33% with 100% agreement. Both are in-sample observations, not calibrated probabilities or validated operating thresholds. A post-hoc gate recovered 93% of reference sends while dropping 53% of rows in the same sample. It needs held-out validation before automatic decisions.

Small atomic checks were correct for five of five competitor-studio cases and fifteen of fifteen function/budget-owner checks on reference sends. Those counts do not estimate general task accuracy. The fixed combining rule reached 66% band agreement, which shows why correct component answers do not guarantee a useful business decision.

The supplied report states two-run Jev band self-agreement of 98%, with raw scores drifting by up to 0.24. It reports 84–87% self-agreement for the three repeated short baseline arms. These are descriptive observations from two runs, not a reliability guarantee. Unrepeated arms have no self-agreement estimate.

Limits

There are 83 contacts, not 166 independent contacts. Both brand decisions reuse the same people and evidence. Boundary-focused sampling prevents direct inference about production accuracy. No formal paired significance test or equivalence test is reported; a few percentage points cannot establish a reliable winner here. The model-agent reference can share biases with the tested systems. The full prompt was tuned on the sample, and the Jev question author had seen earlier model errors. There was no reasoning-disabled baseline arm, no held-out operational gate, and no business-outcome experiment.

Raw contacts, source snapshots and prompts containing private commercial context remain private. The source scripts and frozen input hashes are retained for the owner, but are not linked as a downloadable public reproduction package. A public methods release would require a separately reviewed, de-identified test set.

Return to the Jev field guide.