Systems Lab

Agent skill

perfect-company-list

Build a near-complete company list for any ICP by learning how your real customers label themselves, pulling every industry and keyword they use, judging every single row with a cheap AI model (Jev or gpt-5-nano), and running a Claygent lookalike loop until it runs dry.

activeNeeds a keyActs undeclared1,168 words

Filed under Prospecting and list building.

From growthenginenowoslawski/coldoutboundskills · 52 skill entries · 739 · pushed 2026-10-05

What it does when it runs

Build a near-complete company list for any ICP by learning how your real customers label themselves, pulling every industry and keyword they use, judging every single row with a cheap AI model (Jev or gpt-5-nano), and running a Claygent lookalike loop until it runs dry. The method from the SCULPT 2026 talk "Build the perfect company list". Use when a database filter feels too narrow or too noisy, when someone says "find every company like our customers", "our TAM looks too small", or "why is this list full of junk".

Automated analysis of the skill and the 2 files bundled beside it. A skill’s own description is written to be selected by an agent, so it describes the job and not the dependencies.

Keys and connectors you must supply
  • TYPESAFE_API_KEY
Hosts it reaches
  • api.typesafe.ai
Tool permissions it declares
No allowed-tools in the frontmatter. It does act, so it runs under whatever permissions your session already grants.
Actions present in the files
shellwrites files

Ask about perfect-company-list

Opens your assistant with this page's verified links already in the prompt.

Is this safe to install?ClaudeChatGPT
Adapt it to my stackClaudeChatGPT
What else do I need for it to workClaudeChatGPT
Rather ask a human? Talk to Cheetah
git clone --depth 1 --filter=blob:none --sparse https://github.com/growthenginenowoslawski/coldoutboundskills.git /tmp/coldoutboundskills
git -C /tmp/coldoutboundskills sparse-checkout set "skills/perfect-company-list"
mkdir -p ~/.claude/skills/perfect-company-list
cp -R "/tmp/coldoutboundskills/skills/perfect-company-list/." ~/.claude/skills/perfect-company-list/

Picked up without a restart. A project skill of the same name is shadowed by your personal one. For one repository only, swap ~/.claude/skills for .claude/skills. Claude Code docs ↗

The folder is the same in every client that implements the format — 46 of them — so if yours is not above, only the destination changes.

Before you install: this skill will not complete its job on a bare agent. It needs TYPESAFE_API_KEY, which you have to obtain separately.

Reproduced in full from growthenginenowoslawski/coldoutboundskills/blob/25c5d85fbb5dd3efec97b0ca457cbbc286547476/skills/perfect-company-list/SKILL.md, which is licensed MIT (repository). 1,168 words, 10 headings.

Perfect company list

Database filters fail in two directions at once.

  • They miss real fits. Industry labels are self-reported, or guessed: many LinkedIn company pages carry the banner "This listing was automatically created by LinkedIn", which means nobody at the company picked the label. We found an FDIC bank listed as Construction, an SEC-registered wealth manager listed as Semiconductor Manufacturing, and a New Jersey school district listed as Retail Luxury Goods and Jewelry.
  • They pull in junk. A plain Banking + United States filter returned 16,895 companies. About 10,500 of them were not banks: 3,679 mortgage companies and lenders, 2,414 credit unions, 2,210 dead sites or unrelated businesses (a Toyota dealer, a pawn and gun shop, a youth soccer club), 839 vendors that sell to banks, and the Federal Reserve Bank of St. Louis.

The fix is not a better filter. It is: use filters to cast a wide net, and use AI to decide who is in it. Judging is now cheap enough that pulling five times too many companies costs almost nothing.

The workflow

1. Learn    your customers  -> the industries and keywords they actually use
2. Pull     every one of those industries + keywords (keep size and geo tight)
3. Judge    every single row with one ICP question
4. Loop     Claygent finds lookalikes of your fits -> dedupe -> same judge -> repeat
5. Stop     when 20 runs in a row add one company or fewer

1. Learn from your customers

Put your customer list (or dream accounts) through Clay's Enrich Company and record, per customer:

  • the LinkedIn industry they self-selected
  • keywords in their description and job posts
  • headcount and HQ geography (the fields databases get right)

Then tally the industries. In our bank demo, 50 known banks used as the "customer list" came back 43 Banking, 4 Financial Services, 2 blank and 1 Telecommunications. You pull all four.

2. Pull every industry and keyword

If even one customer shows up under an industry, pull every company in that industry inside your size and geography band. Add keyword searches over name and description ("bank", "school district", "wealth management") with no industry filter, which is how blank-industry records get in.

Keep headcount and geography tight. Keep industries wide open. You will pull a lot of junk. That is the point.

3. Judge every row

One question per company, yes or no. Write it so the edge cases are decided in the criteria, not left to the model:

Question: Is this company a bank or savings institution that takes deposits and would be
          FDIC-insured (commercial bank, community bank, savings bank, thrift, trust bank)?
Yes:      A deposit-taking bank or thrift (FDIC-insured type)
No:       Anything else: credit union, mortgage lender, fintech, broker, wealth manager,
          insurance, payments, consultancy, bank software vendor, bank holding company with
          no bank, association

Run it with scripts/jev_judge.py (Jev, standard library only):

TYPESAFE_API_KEY=... python scripts/jev_judge.py pulled.csv judged.csv \
  --question "Is this company a bank or savings institution that takes deposits?" \
  --yes "A deposit-taking bank or thrift" \
  --no  "Anything else: credit union, lender, fintech, vendor, association"

It sends 20 companies per request as one shared state with one question per company, which keeps you under Jev's default 1,200 requests per minute. On a 60-row canary, batched and one-at-a-time answers agreed on 51 of 54 rows. Canary 50 to 100 rows before you run the whole pull.

Any cheap model works for this step. What it costs to judge 10 million companies at ~350 input tokens each, minimal reasoning:

ModelCost for 10M
GPT-6 Luna$375 to $450 (batch: $188 to $225)
gpt-5-nano$195 to $255 (batch: $98 to $128)
Jevabout $147 (output tokens are free)

Leave reasoning on and output tokens dominate: $1,000 to $3,000 for the same job.

4. Loop with Claygent

Take companies you have judged as fits and ask Claygent: "these companies fit our ICP, find more like these that we might have missed." For each result:

  1. Check the domain against everything you have already judged. Skip it if you have seen it.
  2. Run anything new through the same judge question.
  3. New fits become the next seeds.

5. Stop when it runs dry

Stop when 20 runs in a row add one company or fewer. In our demos the loop was still adding about one company per run at 150 runs, so also set a budget cap.

Checking yourself

When a public registry exists, score against it. Registries are complete lists you can match on:

RegistryCount (Sept 2026)Has websites
FDIC BankFind (active insured banks)4,23198%
SEC investment adviser roster (US)14,86360% list a real website
NCES district directory (New Jersey)668 districtsyes

Results from our runs, normal filter = every LinkedIn industry label that fits, industry only:

ListNormal filterThis process
FDIC banks3,356 (79%)3,718 (88%)
SEC advisers4,636 (31%)13,621 (92%)*
NJ school districts360 (54%)532 (80%)

* The SEC number counts every firm found anywhere in the data stack, including exact-name matches in a second database, because 2,729 advisers list no usable website on the SEC roster.

Gotchas we hit

  • Match on every website a registry lists. The SEC roster CSV shows only the first URL, which is often an Instagram or LinkedIn link. The full IAPD XML feed has all of them; switching added 1,263 matches that were already in our list.
  • Enrichment providers mis-join. One provider mapped mariner.com to a nursing home's LinkedIn page and tagstonecapital.com to a nail salon's. Verify any label you plan to quote on the live page.
  • Registries go stale too. Some "missing" companies had a parked domain on the registry but a live site elsewhere. Search the name before you call something unfindable.
  • Some entities are not companies. Separate bank charters of big brands, fund GP shells and tiny banks with no website exist in registries and nowhere else. No database will have them.

Files

FileWhat it is
SKILL.mdthis method
clay-workflow.mdthe three Clay workflows (learn, judge, loop) built from the clay CLI
scripts/jev_judge.pybatch judge for a CSV with Jev

Verification: the method, numbers and gotchas above come from runs on 2026-09-27 against FDIC, SEC and NCES registries (more than 900,000 company judgments across the demos, about $12 of Jev in total). scripts/jev_judge.py was run on a 40-row sample (40 in / 2 fit). The Clay workflow in clay-workflow.md is a specification: it has not been built and published end to end.

Related: list-expander (seed fingerprinting and filter mining), list-builder (multi-source lanes), icp-prompt-builder.

Files bundled with it

These load only when the skill asks for them, so they cost nothing until it runs.

Other skills for the same job

Different authors, same problem. Matched on the words in the skill name, across every library in the catalogue except this one.

Need help setting it up?

This page tells you what perfect-company-list does and what it needs. Cheetah builds the agent setup it runs inside: data, CRM, sequencing and the guardrails.

Book a call →

The directory stays free. There is nothing gated behind this.