Systems Lab

Agent skill

scoring-analysis-and-distribution

Audit a TAM scoring model against its underlying data, simulate revised rubrics, and recommend rebalanced bands + weights + multipliers that widen the score distribution so accounts visibly stack-rank instead of piling at the same Final %.

activeSelf-containedActs undeclared4,337 words

Filed under Prospecting and list building.

From OneGTM/gtm-skills · 9 skill entries · 0 · pushed 2026-09-30

What it does when it runs

Audit a TAM scoring model against its underlying data, simulate revised rubrics, and recommend rebalanced bands + weights + multipliers that widen the score distribution so accounts visibly stack-rank instead of piling at the same Final %. Use whenever a client hands over a Clay / Deepline / spreadsheet TAM with a working scoring model and the symptom is "everyone scores the same," "distribution is compressed," "top of the list looks wrong," or "we want a clean DQ cutoff that means something." Triggers on phrases like 'rescore this list', 'fix the scoring', 'why is everyone scoring the same', 'distribution is bad', 'compressed', 'rebuild this scoring model', 'spread out the scores', 'twin-firm test', or any CSV export + scoring rubric.

Automated analysis of the skill and the 8 files bundled beside it. A skill’s own description is written to be selected by an agent, so it describes the job and not the dependencies.

Keys and connectors you must supply
None found.
Hosts it reaches
No third-party host appears in the skill or its bundled files.
Tool permissions it declares
No allowed-tools in the frontmatter. It does act, so it runs under whatever permissions your session already grants.
Actions present in the files
shellwrites files

Ask about scoring-analysis-and-distribution

Opens your assistant with this page's verified links already in the prompt.

Is this safe to install?ClaudeChatGPT
Adapt it to my stackClaudeChatGPT
What else do I need for it to workClaudeChatGPT
Rather ask a human? Talk to Cheetah
git clone --depth 1 --filter=blob:none --sparse https://github.com/OneGTM/gtm-skills.git /tmp/gtm-skills
git -C /tmp/gtm-skills sparse-checkout set "skills/scoring-analysis-and-distribution"
mkdir -p ~/.claude/skills/scoring-analysis-and-distribution
cp -R "/tmp/gtm-skills/skills/scoring-analysis-and-distribution/." ~/.claude/skills/scoring-analysis-and-distribution/

Picked up without a restart. A project skill of the same name is shadowed by your personal one. For one repository only, swap ~/.claude/skills for .claude/skills. Claude Code docs ↗

Or take the whole library

This repo ships a .claude-plugin manifest, so Claude Code can install all 9 skills at once. Plugin skills are invoked as /<plugin>:<skill>, so they never collide with your own.

/plugin marketplace add OneGTM/gtm-skills
/plugin

The folder is the same in every client that implements the format — 46 of them — so if yours is not above, only the destination changes.

Reproduced in full from OneGTM/gtm-skills/blob/36da828338f8062c90f6a4a8ea8b918e4160f1e5/skills/scoring-analysis-and-distribution/SKILL.md, which is licensed MIT (repository). 4,337 words, 39 headings.

Scoring Analysis & Distribution Skill

End-to-end workflow for auditing a scoring rubric, simulating revisions on the actual data, and recommending changes that widen the spread so accounts stack-rank usefully. Output is a single boss-forwardable doc set plus a before/after comparison.

Before you start: read references/scoring-tables-primer.md for the full conceptual model (mechanics, levers, architectures, worked example). Everything in this SKILL.md is the enforcement layer on top of that primer.

The non-negotiable workflow order

Do these phases in this order. Skipping or reordering creates rework.

  1. Phase 0, Architecture decision (FIRST, before any band design). Ask the client one question: "In your scoring tool, does DQ stop enrichment from running on a row?"

    • Yes (Clay default) → use pure-positive + smart Hard DQ architecture. This is the default for any Clay GTM engagement.
    • No → ask if they'll tolerate cross-axis formula columns. If yes, use with-multipliers. If no, use pure-card with negatives.

    This decision determines which levers you can pull. Don't propose bands or weights until this is settled.

  2. Phase 1, Profile + bug-audit (BEFORE touching weights). Coverage per column. Distribution stats on existing Final %. All 4 bug detectors (default-when-null, normalization-cap, band-gap, empty enrichment). A single bug fix often moves stdev more than any weight change.

  3. Phase 2, Apply levers within the chosen architecture. See lever table below. Pull only the ones available to your architecture.

  4. Phase 3, Simulate side-by-side. Original vs proposed, on the same row pool. Report stdev, histogram, twin-firm spread per axis, tier counts.

  5. Phase 4, Validate against the 6 gates. Fail any gate → iterate within Phase 2.

  6. Phase 5, Write up. Boss-forwardable. TL;DR table at top. Paste-ready bands for the client's tool.

Input contract

The skill expects three things from the client (or the working directory):

InputRequiredForm
Accounts databaseyesCSV / Clay export / DB query: one row per company, all enrichment columns + current scoring columns
Per-criterion scoresyesAlready-computed score columns per row (the *_Score, *_Golden, Grand Total, Final % columns Clay produces, or equivalent)
Scoring rubric definitionyesThe bands per card + per-criterion weight + Grand Total formula. Either documented (table) or inferable from the score columns + screenshots
Client contextrecommendedWhat they sell, ICP definition, named good-fit firms
DQ threshold targetrecommendedThe Final % below which they want to disqualify (e.g. 40%). Default: pick from histogram.

Output contract

The skill produces these artifacts in <workspace>/scoring/analysis/:

FileContents
RECOMMENDATIONS.mdTL;DR table (original vs proposed stdev + twin-firm spread + DQ-threshold check), list of bugs found, full revised rubric (bands + weights), 4-lever rationale
simulate_v[N].pyPer-row simulation script
scored_v[N].csvPer-row score + breakdown columns for the revised model
summary_v[N].txtHistogram, percentiles, top-K, twin-firm spread, tier counts
display_score_lookup.txtGT → Display Score mapping table + paste-ready Clay IF formula (only when bell curve normalization is applied)
(optional) COST_PLAN.mdOnly if new enrichment is required to close real data gaps

Every recommendation must include:

  1. Proposed new distribution (stdev + histogram shape + twin-firm spread)
  2. Quantified improvement over the original
  3. The new bands/scores/weights/multiplier formulas in copy-paste form for the client's tool (Clay scoring cards or formula columns)
  4. Impact on the Final % for a small set of named firms (top 5, bottom 5, 3 known-good targets)
  5. If bell curve normalization is applied: the GT → Display Score lookup formula, the Display Score histogram, and the recalibration trigger criteria

The two metrics that matter

MetricWhat it measuresTarget depends on architecture
stdev of Final % (within DQ=False pool)Overall spread across the listPure-positive: ≥ 0.20 / Pure-card w/negs: ≥ 0.27 / With-multipliers: ≥ 0.35
Twin-firm spreadΔ between two identical firms differing only in axis X. Run on every weighted axis.≥ 10 pts on the top 3 axes regardless of architecture

Optimize both. stdev alone can be flattered by a long penalty tail without giving you adjacent-row resolution. Twin-firm spread alone can be flattered by one outsized axis weight that compresses the rest.

The stdev target scales with architecture choice, don't compare a pure-positive 0.22 to a with-multipliers 0.37 and call it failure. Compare to the right baseline.

The five levers that actually widen distribution

These are the ONLY five levers. Anything else is a permutation of them.

LeverWhat it doesFree in Clay?Pure-card OK?
1. More bands per cardIncreases resolution within an axis. A 3-band card can emit 3 values; a 9-band card can emit 9. Required prerequisite for any weight adjustment to matter.YesYes
2. Balance weights across axesFlatter weight distribution → more independent contributors summing → bell-shaped histogram. Don't pile weight onto correlated axes, it creates 5%-of-list pile-ups.YesYes
3. Deep negative band values within cardsA card with bands going from −25 to +22 has more variance than one going 0 to 22. The negative tail extends Grand Total downward and the misfit pile separates cleanly from the candidate pile. In pure-card architecture, this is the primary spread driver.YesYes
4. Widen Fit_Multiplier swing (cross-axis)Compounds top-tier fit. Range 0.4-1.8 stretches the top of the histogram. Only works if client allows a separate Multiplier formula column.Yes (one formula column)No, breaks the architecture
5. Deepen Negative_Penalty as separate column (cross-axis)Like Lever 3 but cross-axis. Catches misfits that span multiple axes (e.g. DTC consumer brands via name regex).Yes (one formula column)No

What does NOT widen distribution (common traps)

Tempting moveWhy it fails
Push more weight into the top axes (Industry/Reps/PE/etc.)Those axes correlate, firms hitting all of them already top out together. More weight just creates pile-ups at the joint maximum. Tested: pushed Industry 22→28, stdev moved +0.007 but largest bucket grew from 3.15% to 5.93%.
Raise a card's raw max from 1 to 10 without adding bandsMathematically identical to original, Golden = score/max → same 0-1 output. Resolution is set by band COUNT, not max value.
Add weight to an axis with low coverageIf only 5% of rows have data for axis X, bumping its weight from 5 to 15 only differentiates that 5%. Doesn't widen the overall histogram.
Add weight to an axis derived from another already-weighted axis (e.g. Fragmentation as Industry lookup, Rev/Rep as Revenue ÷ Reps)Double-counts the same signal. Keep redundant-but-useful axes as tie-breakers (weight 2-5).

Architecture choice: three options, asked first

Ask the client which architecture they want before designing bands. The choice is operational, not mathematical:

ArchitectureWhat you can useWhat you give upRealistic stdev gain
Pure-positive + smart Hard DQ (RECOMMENDED for Clay-based GTM teams where DQ stops enrichment)All scores ≥ 0. Expanded Hard DQ catches objectively-wrong-fit categories. Soft-fit firms get low positive scores.Negative scores. Cross-axis multiplier.+15% to +35% over baseline (example engagement 1, v10d est: 0.178 → ~0.22)
Pure-card with negativesEach card max = weight. Negative bands within cards for misfit categories. No multipliers. Grand Total can go negative.Cross-axis compounding. Cross-axis name regex. Operator must accept negative scores in the UI.+30% to +60% over baseline (example engagement 1, v10c: 0.178 → 0.279)
With-multipliersPer-card weights + Fit_Multiplier formula column + Completeness_Multiplier + Negative_Penalty column. Most aggressive.3 new formula columns. Most complex Grand Total formula. Cross-axis multiplier is harder to explain to a non-technical reviewer.+50% to +110% over baseline (example engagement 1, v9: 0.178 → 0.367)

The deciding question: does DQ stop enrichment in the client's setup?

If yes (Clay default behavior): pure-positive + smart Hard DQ is correct. Negative scores feel like "soft-DQ" but actually keep the row in the enrichment loop. That's not what DQ semantically means in Clay. Use Hard DQ for "objectively wrong, never spend another credit," score-low-positive for "weak fit, keep enriching, might surface later via signals."

If no (custom workflow where DQ is just a filter for the active outbound list, not a stop-enrichment flag): pure-card with negatives gives more stdev.

How to decide which categories belong in Hard DQ vs score-low-positive

  • Hard DQ: the firm's industry/type is fundamentally wrong (software company, bank, school, restaurant). No future enrichment refinement makes them a fit.
  • Score-low-positive: the firm might be misclassified, weak-fit, or have wrong data that better enrichment could refine. Keep in pool with score 0-2.

Within whichever architecture you pick, push every lever

To maximize stdev within the chosen architecture:

  • Wide band gaps between tiers (e.g. Tier 1 = 22, Tier 2 = 16, Tier 3 = 8, 6+ pt gaps do the discrimination)
  • Tight bullseyes (e.g. Sales Reps peak only at 60-89, with sharp taper outside)
  • Intra-card tie-breakers when an axis has a clear primary signal + secondary modifier (e.g. PE = 10 base for any PE firm + 0-4 bonus from Fundraised $ tier; max 14)
  • Pure-card with negatives only: deep negative bands (e.g. Industry "Other" = −25)
  • Pure-positive only: expanded Hard DQ list catches the misfits; soft-fit categories score 0-2

Operational tradeoff to surface in the writeup

The architecture choice has real operating consequences:

ArchitectureEnrichment cost trajectoryStack-rank cleanness
Pure-positive + smart Hard DQLower (no enrichment on wrong-fit)Cleanest (0-100% range, no negatives to explain)
Pure-card with negativesHigher (everything gets enriched, even -25-scoring rows)Negative scores in UI
With-multipliersHighest + 3 formula columns to maintainMultiplier values in UI; harder to QA

Clay column types: scoring cards vs formula columns

Clay has two distinct column types for scoring. Use the right one.

Scoring cards (band-based UI)

Used for: individual criterion scores (Engineer Count Score, FTE Score, etc.).

Clay's scoring card UI defines bands declaratively, no code. Each card has one or more Scoring Criteria, and their scores sum within the card.

Each scoring criteria has:

  • Values to Score, the input column
  • Comparison Type, Contains (Text) or Between (Number)
  • Keywords, comma-separated values/ranges to match
  • Scores, comma-separated scores, one per keyword

Between (Number) rules:

  • Every keyword must be a range, not a bare number. Use 0-0.9 not 0.
  • Open-ended upper bounds: 251-999999 not 251+.
  • Ranges are inclusive on both ends.

Contains (Text) rules:

  • Keywords match if the cell value contains the keyword text (case-insensitive).
  • First matching keyword wins. Order keywords from most-specific to least.
  • Empty cells match no keyword and score 0 by default.
  • Exhaustive keyword lists required. Any value not matching a keyword silently scores 0. This is especially dangerous for Country/Geography cards, if Germany isn't in the keyword list, all German firms score 0 with no error. Always validate coverage (% nonzero) after creating or editing a Contains (Text) card.

Multi-criteria cards: A single scoring card can have multiple criteria (e.g., Cloud Specialization has Criteria 1 on Cloud Practice Classification and Criteria 2 on Cloud Service Count). The scores from all criteria in the same card sum.

Documenting scoring cards in RECOMMENDATIONS.md: Use this format for each card so the operator can paste directly into Clay:

> **Scoring Criteria 1**
> Values to Score: `Column Name`
> Comparison Type: Between (Number)
> Keywords: `0-0.9, 1-5, 6-10, 11-999999`
> Scores: `0, 3, 6, 10`

Formula columns (Clayscript)

Used for: DQ logic, Grand Total (sum of card scores), Display Score (bell curve lookup), Tier labels, and any cross-column computation.

Clayscript is JavaScript. Column references use {{Column Name}} syntax. Guard against null with || 0 for numbers, || "" for strings.

Critical: Clay strips const/let/var declarations and inlines the expression. This breaks operator precedence. Never write const gt = {{Col}} || 0; gt >= 50 ? ..., Clay will render it as {{Col}} || 0 >= 50 which parses as {{Col}} || (0 >= 50) due to >= binding tighter than ||. Always wrap the null guard in parentheses: ({{Col}} || 0) >= 50 ? ...

Scoring cards should NEVER be written as Clayscript formulas. Clayscript is only for columns that need cross-column logic (sums, lookups, conditionals across multiple scoring card outputs).

Process

Phase 0: Workspace

Create <workspace>/scoring/analysis/. TaskCreate the phases: profile → audit bugs → simulate → validate → write up.

Phase 1: Profile coverage and existing distribution

Use templates/profile_data.py adapted to the client's columns. Compute:

  • Per-column coverage (% non-null)
  • Distribution stats on Final %: mean, median, stdev, P10/25/50/75/90/95/99
  • Histogram (0.05-pct bins)
  • Largest 0.01-pct bucket (the pile-up metric)
  • Per-axis Golden column: min, max, mode, mode share

Phase 1.1: Bug-hunt the existing scoring (always run all 4)

  1. Default-when-null detector, for each derived score column, compare its nonzero rate to the coverage of its source columns. If derived nonzero >> source coverage, a default-when-null bug is feeding constant values to most rows.
  2. Normalization-cap check, for each *_Golden column, find the observed max. If it caps well below 1.0, the denominator is larger than realistically achievable raw max, the top of that axis is unreachable.
  3. Band-gap detector, inspect keyword bands (1-5, 11-25, 26-50, ...). Firms whose raw value falls in a gap silently score 0.
  4. Empty enrichment check, any column referenced in the rubric but <2% populated is dead weight.

Bug fixes are Phase 0, execute these before iterating on weights. A single default-when-null bug often does more damage to the distribution than any weight imbalance. See references/common_pitfalls.md for worked examples.

Phase 2: Translate the rubric and apply the 4 levers

For each card and each weight, decide which lever applies:

Symptom from Phase 1LeverConcrete fix
Card has ≤4 bands but covers wide value range1 (bands)Add bands; log-shape for right-tailed numerics
Two firms identical except axis X differ by <5 pts2 (weights)Bump axis X weight (only if uncorrelated with top axes)
Largest 0.01-pct bucket > 5%2 (weights)Trim top-axis weight; spread to mid-tier axes
Top-K firms tied at the cap3 (multiplier)Widen Fit_Multiplier swing; remove cap on Final %
DQ rows pile at exactly 0 (clamp)4 (clamp)Drop MAX(0, ...) on Grand Total
Bad-fit firms scoring same as low-data correct-fit4 (penalty)Add Negative_Penalty formula column; make penalties ≥-50 for hard misfits
Histogram is bimodal or step-shaped2 (weights)More balanced weights → smoother curve
Many firms at exactly the mediansourcingRe-check the underlying TAM filter; rubric can't fix sourcing gaps

See references/distribution_techniques.md for the formulas behind each lever. See references/levers.md for the empirical results from past engagements showing which lever produced which delta.

Phase 3: Simulate the revised rubric

Use templates/simulate_scoring.py adapted to the client's columns. Always run at least two versions side-by-side: v[N] (current state, baseline) and v[N+1] (proposed). Report:

  • stdev (target ≥ 0.30)
  • Median, P25, P75
  • Histogram (5-pt bins on raw final pts, 0.05 bins on Final %)
  • Largest 0.01-pct bucket (target ≤ 5%)
  • Tier distribution (A/B/C/D/F)
  • Twin-firm spread for the top 3 axes by weight (synthesize identical firms varying one axis)
  • Top-15 and bottom-10 named lists

Phase 3a: Bell curve normalization (when raw stdev hits ceiling)

When the raw rubric's stdev plateaus below target due to structural data-coverage limits (common when >30% of rubric weight sits on axes with <10% coverage), apply bell curve normalization as a display layer on top of the raw scores.

This does NOT replace rubric iteration, it runs AFTER the rubric is as good as the data allows. The raw Grand Total remains the source of truth for debugging.

Process:

  1. Export non-DQ pool's Grand Total values
  2. Compute percentile rank per row → inverse-normal CDF → scale to mean=50, std=20
  3. Average Display Score per GT value to build a lookup table
  4. Implement as a nested IF formula in Clay (see references/distribution_techniques.md §5a)
  5. Sync Display Score (not raw GT) to HubSpot as the rep-facing "Fit Score"

When to use: Raw stdev is 10-18 on 0-100 scale (0.10-0.18 on 0-1) and further rubric iteration yields diminishing returns (<+1 stdev per iteration). The bell curve transform lifts effective stdev to ~20 and centers the distribution at 50.

When NOT to use: Raw stdev is already ≥20, or the client needs absolute scores ("hit 70% of the rubric criteria") rather than relative ranking. In those cases, keep iterating the rubric.

Recalibration trigger: The GT → Display lookup must be recomputed whenever DQ criteria, card bands, or card weights change. A structural rubric change invalidates the mapping.

Phase 3b: Raw Grand Total max is typically 85-95, not 100 (this is correct)

A well-designed rubric with mutually-exclusive bullseyes (Sales Reps 60-89, FTE 400-799, Revenue $1B-$5B, etc.) produces a structural raw max of ~85-95 because no real firm hits every card's bullseye simultaneously. This is intentional, not a bug.

For example: a 60-rep firm with $3B revenue would need ~50K FTE per rep (impossible). A 400-FTE firm with $3B revenue has $7.5M/employee (unrealistic for sales-led orgs). So the structural raw max for real firms is ~90.

Do not "fix" this by loosening bullseyes, that reduces discrimination and collapses the top tier. The right answer is:

  • Use the Display Score (Phase 3a bell curve) as the 0-100 rep-facing metric
  • Tier on Display Score, not raw Grand Total
  • Treat raw Grand Total as a debugging metric only

If a client insists on raw 0-100 with no bell curve, use one of:

  • A. Accept raw max ≈ 90 (most honest)
  • B. Rescale to observed max, Final % = (GT / observed_max) × 100. Warn that this makes scoring population-dependent (one new high-scoring firm shifts everyone).
  • C. Use a "perfect-fit firm" denominator, define an explicit max raw possible for the optimal firm (e.g. 92 in the sales-led B2B case) and divide by that. Honest but requires explaining the denominator to the operator.

Default to A + Display Score. B and C are escape hatches if the client refuses bell curve normalization.

Phase 3c: Tier scheme: 20 tiers at 5% intervals on Display Score

The recommended tiering for the rep-facing UX is 20 tiers at 5% intervals on the Display Score. This gives finer stack-rank granularity than the A/B/C/D/Low scheme while still being scannable.

Two valid scales for Display Score, pick one and match the Tier formula:

ScaleWhen to useDisplay valuesTier formula (T1 = best, T20 = worst)
0-1 (recommended)Display Score column type = Number, matches Final Score Percentage convention0.000 to 1.000"T" + (20 - Math.floor(Math.min(0.99, ({{Display Score}} || 0)) / 0.05))
0-100Display Score column type = Number, want literal integer-ish values0 to 100"T" + (20 - Math.floor(Math.min(99, ({{Display Score}} || 0)) / 5))

Both produce T1 (best) through T20 (worst) with identical population distributions. Pick the scale based on what reads naturally in the client's existing column conventions.

Tier direction convention: T1 = top (best fit), T20 = bottom (worst fit). This matches how people naturally read rankings ("we have 200 T1 accounts to call this week"). If a client prefers T20 = best, swap (20 - Math.floor(...)) for (Math.floor(...) + 1) in the Tier formula.

In templates/bell_curve.py, set DISPLAY_SCALE = 1 for 0-1 output or DISPLAY_SCALE = 100 for 0-100. The lookup file generates both the Display Score formula AND the matching Tier formula automatically (defaulting to T1 = best).

Letter tiers (alternative), 15-tier A+/A/A- ... F-/F+ scheme:

Tier_Letter =
  (({{Display Score}} || 0) >= 95) ? "A+" :
  (({{Display Score}} || 0) >= 90) ? "A"  :
  (({{Display Score}} || 0) >= 85) ? "A-" :
  (({{Display Score}} || 0) >= 80) ? "B+" :
  (({{Display Score}} || 0) >= 75) ? "B"  :
  ... (continue through F-)

Why 5% tiers over the old 5-tier scheme:

  • 5 tiers (A/B/C/D/Low) hide a lot of stack-rank signal, 50% of the pool ends up in "C" or "D" with no further differentiation
  • 20 tiers map cleanly to the bell-curve Display Score; each tier's expected population is predictable (T20 ≈ 2.4%, T11 ≈ 9.2%, T1 ≈ 2.5%)
  • Operator can filter "T17 and up" or sort descending, both work without further thresholding

Old 5-tier scheme is still fine for coarse summaries (boss memos, "we have 3.4% A-tier accounts"). Keep both columns: a 20-tier Tier for filtering, a 5-tier Tier_Coarse for reporting.

Phase 4: Validate against gates

GateTargetIf fails
Raw stdev (DQ=False pool)≥ 0.15 (raw), or ≥ 0.20 (after bell curve)Apply more aggressive levers or bell curve normalization
Largest 0.01-pct bucket≤ 5%Weights too top-heavy, flatten across mid-tier axes
Twin-firm spread on top 3 axes≥ 10 pts eachWeight too low on the failing axis; bump it (only if uncorrelated)
DQ threshold separationDQ rows median ≤ threshold − 10 pts; real candidates median ≥ threshold + 10 ptsNegative_Penalty too shallow; deepen
Top-K named good-fit firms≥ 80% appear in A or B tierSourcing gap OR specific bug, re-run Phase 1.1
Histogram shape (Display Score)Bell-shaped around 50; 20 tiers at 5% intervals with predictable populations (T20 ≈ 2.4%, T11 ≈ 9.2%, T1 ≈ 2.5%)Rubric iteration or recalibrate bell curve lookup
Raw Grand Total max~85-95 (NOT 100) for a well-designed rubric. Don't loosen bullseyes to reach 100, that reduces discrimination. Use Display Score for the 0-100 UX.If raw max < 80, bullseyes may be too tight. If raw max = 100 on many rows, bullseyes are too loose.
Rows at Display Score = 100≤ 3% of non-DQ pool (the true top tail).If >3% pile at exactly 100, suspect the inv_norm_cdf function has the operator-precedence bug (numerator polynomial not wrapped in outer parens, c[5] / denom parses as a separate term). See templates/bell_curve.py for the correct implementation.
Rows at Display Score = 0≤ 3% of non-DQ pool.Same bug as above (tail branch broken). Or DQ list is letting too many true zero-fit rows through.

Phase 5: Write up

RECOMMENDATIONS.md template:

# TL;DR: table comparing original vs proposed on:
#   stdev | mean | median | largest 0.01-pct bucket | A-tier % | F-tier %
#   twin-firm spread for top 3 axes
#   DQ threshold (e.g. "all 5,362 F-tier firms below -10 pts, all real candidates above 25 pts")

# Bugs found in current model (Phase 1.1 output)
# Revised rubric (paste-ready bands + weights + formulas)
# 4-lever rationale: which lever was pulled, why, expected stdev delta
# Validation gate results
# Open questions (sourcing gaps, data-gap fixes that need budget: defer to COST_PLAN.md)

Key conventions

  • Always run the twin-firm test. Synthesize one row from your top-15, clone it 8-10 times varying ONE axis, score all of them, report the spread between #1 and #10. Repeat for each axis with weight ≥ 5.
  • Two versions side-by-side, minimum. Never report a proposed model without showing the baseline numbers next to it.
  • Bugs before weights. A default-when-null bug or normalization-cap bug undoes any weight tuning. Find them first.
  • Bands first, weights second, shape third. Number of bands sets the upper bound on resolution. Weights determine each axis's share. Multiplier/penalty stretch the top and bottom. Always in that order.
  • The bottom matters as much as the top. Negative_Penalty + no-clamp on the Grand Total formula is responsible for ~half of any meaningful stdev improvement. Don't skip it.
  • Boss-forwardable means one number at the top. TL;DR has the headline (stdev gain + twin-firm spread gain). Detail goes below.
  • No emojis in docs. Per OneGTM CLAUDE.md.

When NOT to use this skill

  • Building a scoring model from scratch with no existing data → use scoring-model skill
  • Pure outreach copy → use instantly-copywriter
  • Pure enrichment with no scoring question → use deepline-gtm
  • Client wants a Clay → Deepline migration of the workflow → use clay-to-deepline

References

  • references/scoring-tables-primer.md, READ FIRST. Full conceptual model: mechanics, 6 levers, 3 architectures, validation gates, example engagement 1 empirical track (v4 through v10d). Everything in this SKILL.md is the enforcement layer on top of that primer.
  • references/levers.md, empirical results table for all 6 levers + architectural decision flowchart
  • references/distribution_techniques.md, formulas: log transforms, multiplicative gates, type-aware scoring, percentile normalization, completeness multipliers
  • references/common_pitfalls.md, failure modes and their fixes
  • references/cost_benchmarks.md, provider rates if new enrichment is genuinely needed

Templates

  • templates/profile_data.py, coverage + distribution profiler + 4 bug detectors (run in Phase 1)
  • templates/simulate_scoring.py, simulation harness with ARCHITECTURE toggle, twin-firm test, DQ-threshold validation, side-by-side comparison report. Defaults to pure-positive architecture.
  • templates/bell_curve.py, Phase 3a Display Score normalization. Reads the scored CSV, builds the GT→Display lookup, emits paste-ready Clayscript IF formula. Includes built-in saturation gates (warns if >3% of rows pile at Display=100 or =0, which indicates the inv_norm_cdf operator-precedence bug).

Files bundled with it

These load only when the skill asks for them, so they cost nothing until it runs.

Other skills for the same job

Different authors, same problem. Matched on the words in the skill name, across every library in the catalogue except this one.

Need help setting it up?

This page tells you what scoring-analysis-and-distribution does and what it needs. Cheetah builds the agent setup it runs inside: data, CRM, sequencing and the guardrails.

Book a call →

The directory stays free. There is nothing gated behind this.