Agent skill
claygent-builder
Generate production-ready Claygent configurations (system prompt + task prompt + JSON schema + examples), test them via Clay webhook, and iterate until quality hits 8.0+
Filed under Prospecting and list building.
From jurjen-gtm-engineer/gtmskills · 55 skill entries · 0 · pushed 2026-10-04
What it does when it runs
Generate production-ready Claygent configurations (system prompt + task prompt + JSON schema + examples), test them via Clay webhook, and iterate until quality hits 8.0+
Automated analysis of the skill and the 3 files bundled beside it. A skill’s own description is written to be selected by an agent, so it describes the job and not the dependencies.
- Keys and connectors you must supply
- None found.
- Hosts it reaches
- No third-party host appears in the skill or its bundled files.
- Tool permissions it declares
- No
allowed-toolsin the frontmatter. It does act, so it runs under whatever permissions your session already grants. - Actions present in the files
- shellwrites files
Install it
View source on GitHub ↗git clone --depth 1 --filter=blob:none --sparse https://github.com/jurjen-gtm-engineer/gtmskills.git /tmp/gtmskills git -C /tmp/gtmskills sparse-checkout set "skills/claygent-builder" mkdir -p ~/.claude/skills/claygent-builder cp -R "/tmp/gtmskills/skills/claygent-builder/." ~/.claude/skills/claygent-builder/
Picked up without a restart. A project skill of the same name is shadowed by your personal one. For one repository only, swap ~/.claude/skills for .claude/skills. Claude Code docs ↗
The folder is the same in every client that implements the format — 46 of them — so if yours is not above, only the destination changes.
The skill
Source on GitHub ↗Reproduced in full from jurjen-gtm-engineer/gtmskills/blob/77dc0b3112dbf6cf906dfc3d526b6f7031bf964c/skills/claygent-builder/SKILL.md, which is licensed MIT (repository). 1,813 words, 30 headings.
Skill: Claygent Builder
Purpose
Generate production-ready Claygent configurations (system prompt + task prompt + JSON schema + examples), test them via Clay webhook, and iterate until quality hits 8.0+. Replaces manual prompt-test-rewrite cycles with a structured loop.
This is a standalone skill: not part of any sequential workflow. Invoke independently.
Inputs
- Task description (required): What the Claygent should do (e.g., "Classify whether a company is B2B SaaS from their website")
- Clay webhook URL (required for testing): The webhook URL from a Clay table set up for testing
- Input variables (required): What data Clay sends per row (e.g., company domain, company name)
- Output fields (required): What structured data the Claygent should return
- Test domains/data (optional): If not provided, auto-generated in Step 8
- Model preference (optional): Claygent Neon (default), GPT-4, or Claude
Process
Step 1: Webhook Setup
This step is BLOCKING: cannot proceed to testing without it.
Ask the user for their Clay webhook URL. Explain the setup if needed:
- User creates a Clay table with Webhook as the source
- User copies the webhook URL from Clay
- User pastes the webhook URL into Claude Code
Start the local webhook listener for receiving Clay's callback results (webhook-listener.py lives in this skill's folder):
python3 webhook-listener.py &
If ngrok is available, start a tunnel to expose the listener:
ngrok http 8765
The public ngrok URL becomes the callback_url sent to Clay in each test batch.
Fallback (no ngrok or tunnel): Poll for results via Clay's API/MCP tools, or ask the user to manually copy results from the Clay UI.
Reference: clay-test-template.md (in this skill's folder) for detailed Clay table setup instructions.
Step 2: Native Integration Check
This step is BLOCKING: if a native integration solves the task, stop here.
Check if Clay's 150+ native integrations or existing Claybooks already solve this task without needing a Claygent. Native integrations are:
- Cheaper (no AI credits)
- Faster (direct API calls)
- More reliable (structured APIs vs web scraping)
Decision:
- If a native integration solves it completely: recommend that instead, stop here
- If a native integration solves part of it: recommend hybrid approach (native + Claygent for the gap)
- If no native integration applies: proceed to Step 3
Step 3: Classify Complexity
Classify the task:
| Complexity | Criteria | Example Count |
|---|---|---|
| Simple | Binary yes/no, single field extraction, one page visit | 3 examples |
| Medium | Multi-field extraction, classification with reasoning, 2-3 page visits | 5 examples |
| Advanced | Multi-step research, conditional logic, multiple sources, chain of thought | 10 examples |
Also determine:
- Recommended model: Claygent Neon (default), GPT-4 (complex reasoning), Claude (nuanced classification)
- Estimated credits per row: based on complexity and expected page visits
- Expected fields: list all output fields with types
Step 4: Generate System Prompt
Build the system prompt with exactly 6 sections:
Section 1: Identity + Scope
- Specific role statement (not generic "helpful AI")
- Bounded capabilities (what this agent does and nothing else)
- 2-3 sentences maximum
Section 2: Definitions
- Every classification value defined with concrete criteria
- No ambiguous terms: if two values could overlap, specify the tiebreaker
- Use plain language, not jargon
Section 3: Decision Tree
- Exhaustive if/then logic
- Ordered by priority (most common/clear cases first)
- Every possible input hits exactly one leaf
- Always ends with a default/unknown case
Section 4: Good vs Bad Examples
- Paired correct/incorrect examples with reasoning
- Cover: happy path, edge cases, missing data
- Show the WRONG answer and explain WHY it's wrong
Section 5: Output Format Contract
- Field-by-field spec matching the JSON schema
- Valid values for each enum field with when-to-use criteria
- Confidence level definitions (high/medium/low)
Section 6: Scope Constraints
- What NOT to do (max pages, no form submission, no assumptions from domain name alone)
- Classify vs unknown boundary
- "If unsure, return unknown": false confidence is worse than admitting uncertainty
Step 5: Generate Task Prompt
Keep SHORT: under 200 words. All intelligence lives in the system prompt.
Structure:
- Task statement: one sentence describing what to do
- Input variables:
{{Company Website}},{{Company Name}}, etc. (Clay template syntax) - Instructions: 3-5 numbered, specific steps
- Fallback behavior: one sentence for when things go wrong
- Output instruction: "Return JSON matching the schema."
Step 6: Generate JSON Schema
Build a Clay-compliant JSON schema following OpenAI Structured Outputs rules. Validate against ALL 10 rules before outputting:
- Root =
{ "type": "object" } "additionalProperties": falseon every object- All properties in
"required"array "anyOf": [{"type": "string"}, {"type": "null"}]for optional fields (NOT oneOf/allOf/not)- No validation keywords (no minLength, maxLength, minimum, maximum, pattern, format)
"enum"for fixed value sets- Arrays need explicit
"items"definition - Nested objects also need
additionalProperties: falseand fullrequired - No
$refor$defs: inline everything - Maximum 5 levels of nesting
The schema must be directly pasteable into Clay's "Click JSON Schema" field: no wrapping, no modification.
Step 7: Generate Examples
Generate input/output pairs based on complexity (3/5/10 from Step 3).
Each example must include:
- Input: The exact data Clay would send (domain, company name, etc.)
- Expected output: The exact JSON matching the schema
- Reasoning: Why this is the correct output (2-3 sentences)
- Edge case tested: What boundary this example covers
Coverage requirements:
- Clear positives (unambiguous matches)
- Clear negatives (unambiguous non-matches)
- Edge cases (ambiguous inputs that test decision tree boundaries)
- Missing data (what happens when key info is unavailable)
- Adversarial (inputs designed to confuse, e.g., marketing copy that mimics a different category)
Step 8: Source Test Domains
If the user provided test data, use that. Otherwise, auto-generate a balanced test batch:
- 3-4 clear positives: companies that obviously match the classification
- 2-3 clear negatives: companies that obviously don't match
- 2-3 edge cases: ambiguous companies that test decision boundaries
Use web search to find appropriate companies if the task domain is niche.
For each test domain, note the expected classification (ground truth) so results can be scored.
Step 9: Test via Clay Webhook
This step is BLOCKING: must achieve 8.0+ quality score.
Send test batch to Clay webhook via curl:
curl -X POST "CLAY_WEBHOOK_URL" \
-H "Content-Type: application/json" \
-d '{
"domain": "example.com",
"prompt": "[full task prompt with variables resolved]",
"prompt_version": "v1",
"callback_url": "NGROK_URL/results",
"json_schema": { ... }
}'
Send one request per test domain. Wait for Clay to process (Claygent takes 1-5 minutes per row).
Read results from ./clay-results/*.json (written by the webhook listener).
Score each result on the 7-criterion rubric:
| # | Criterion | Weight | Score |
|---|---|---|---|
| 1 | Accuracy | 25% | Compare to expected ground truth |
| 2 | Consistency | 20% | Same logic applied across all test rows? |
| 3 | Completeness | 15% | All fields populated correctly? |
| 4 | Edge case handling | 15% | Graceful handling of ambiguous/missing data? |
| 5 | Reasoning quality | 10% | Does reasoning match classification? |
| 6 | Schema compliance | 10% | Output matches JSON schema exactly? |
| 7 | Cost efficiency | 5% | Minimal pages visited? |
Calculate weighted average. Decision:
- < 8.0: MUST iterate (proceed to Step 10)
- 8.0 - 8.9: SHOULD iterate (optional, recommend if easy wins exist)
- 9.0+: STOP, production-ready
Step 10: Iterate
Analyze the Clay output, focusing on agent steps (where it navigated, what it read, where it went wrong).
Common failure patterns and fixes:
| Failure | Fix |
|---|---|
| Wrong page visited | Add navigation instructions: "Visit /pricing first, then /about" |
| Timeout on JS-heavy sites | Add constraint: "If homepage doesn't load in 30s, classify from domain + meta tags" |
| Misclassification on edge case | Add specific example for that edge case in Section 4 |
| Missing field | Add explicit instruction for that field in the decision tree |
| Inconsistent confidence | Tighten confidence definitions in Section 5 |
| Non-English site mishandled | Add guardrail: "For non-English sites, use visual signals (pricing layout, business imagery)" |
Bump prompt version and update changelog:
v1: Initial prompt, generated from task description
v2: Fixed: [specific issue found in test results]
v3: Added: [specific improvement based on agent step analysis]
Re-send to Clay webhook with updated prompt. Max 3 iterations.
If after 3 iterations the score is still below 8.0, report the persistent issues and recommend:
- Breaking the task into simpler sub-tasks (multiple Claygent columns)
- Switching models (e.g., Neon to GPT-4 for complex reasoning)
- Adding pre-enrichment steps (native integrations) to reduce what Claygent must figure out
Output
Save final deliverable to outputs/claygent-[task-slug].md (e.g., outputs/claygent-b2b-saas-qualification.md).
Deliverable Structure
# Claygent: [Task Name]
## Built with Claygent Builder
◆
## Native Integration Check
[Result: Claygent required / hybrid approach / native sufficient]
[If hybrid: which native integrations supplement the Claygent]
◆
## Configuration Summary
- **Complexity:** [Simple / Medium / Advanced]
- **Model:** [Claygent Neon / GPT-4 / Claude]
- **Credits/row:** [Estimate]
- **Quality score:** [Final rubric score] (v[N], [N] iterations)
- **Test results:** [X/Y correct on test batch]
◆
## System Prompt
[Copy-paste ready for Clay's system prompt field]
◆
## Task Prompt
[Copy-paste ready for Clay's Claygent prompt field]
◆
## JSON Schema
[Copy-paste ready for Clay's "Click JSON Schema" field, exact Clay-compliant format]
◆
## Examples ([3/5/10])
### Example 1: [Description]
**Input:** [domain / company data]
**Expected Output:**
```json
{ ... }
Reasoning: [Why this is correct] Edge case tested: [What boundary this covers]
[Repeat for all examples]
◆
Test Results
Rubric Scores (v[final])
| Criterion | Weight | Score | Notes |
|---|---|---|---|
| Accuracy | 25% | X.X | ... |
| Consistency | 20% | X.X | ... |
| Completeness | 15% | X.X | ... |
| Edge case handling | 15% | X.X | ... |
| Reasoning quality | 10% | X.X | ... |
| Schema compliance | 10% | X.X | ... |
| Cost efficiency | 5% | X.X | ... |
| Weighted Total | 100% | X.X |
Per-Domain Results
| Domain | Expected | Got | Correct? | Notes |
|---|---|---|---|---|
| ... | ... | ... | ... | ... |
Version History
[Changelog from v1 to final version]
◆
Deployment Notes
- Test protocol: Run on 10 new rows before scaling. Check accuracy > 80%.
- Conditional logic: [Any Clay formula conditions to add, e.g., skip rows without website]
- Known limitations: [Edge cases that remain, cost considerations]
- Scaling notes: [Credit estimates at 100/1000/10000 rows]
Files bundled with it
These load only when the skill asks for them, so they cost nothing until it runs.
Other skills for the same job
Different authors, same problem. Matched on the words in the skill name, across every library in the catalogue except this one.
- workflows-claygent by clay-run · 129
- claygent by Frontal-so · 6
Need help setting it up?
This page tells you what claygent-builder does and what it needs. Cheetah builds the agent setup it runs inside: data, CRM, sequencing and the guardrails.
Book a call →The directory stays free. There is nothing gated behind this.