Systems Lab

Agent skill

yc-jobs-scraper

Scrape daily job listings from YCombinator's Workatastartup platform without duplicates.

activeReaches the webActs undeclared361 words

From Varnan-Tech/opendirectory · 57 skills · 626 · pushed 2026-08-16

What it does when it runs

Scrape daily job listings from YCombinator's Workatastartup platform without duplicates. Use this skill when asked to scrape YC jobs, update the YC companies list, or retrieve the latest startup jobs. It handles authentication, extracts company slugs via Inertia.js JSON payloads, falls back to public YC job pages when necessary, and maintains a local SQLite database to track historical jobs and prevent duplicates.

Read from the skill and the 7 files bundled beside it. A skill’s own description is written to be selected by an agent, so it describes the job and not the dependencies.

Keys and connectors you must supply
None found.
Hosts it reaches
  • account.ycombinator.com
  • dotenvx.com
  • feross.org
  • registry.npmjs.org
  • www.patreon.com
  • www.workatastartup.com
  • www.ycombinator.com
Tool permissions it declares
No allowed-tools in the frontmatter. It does act, so it runs under whatever permissions your session already grants.
Actions present in the files
shellwrites files

Ask about yc-jobs-scraper

Opens your assistant with this page's verified links already in the prompt.

Is this safe to install?ClaudeChatGPT
Adapt it to my stackClaudeChatGPT
What else do I need for it to workClaudeChatGPT
Rather ask a human? Talk to Cheetah
git clone --depth 1 --filter=blob:none --sparse https://github.com/Varnan-Tech/opendirectory.git /tmp/opendirectory
git -C /tmp/opendirectory sparse-checkout set "skills/yc-intent-radar-skill/yc-jobs-scraper"
mkdir -p ~/.claude/skills/yc-jobs-scraper
cp -R "/tmp/opendirectory/skills/yc-intent-radar-skill/yc-jobs-scraper/." ~/.claude/skills/yc-jobs-scraper/

Picked up without a restart. A project skill of the same name is shadowed by your personal one. For one repository only, swap ~/.claude/skills for .claude/skills. Claude Code docs ↗

Or take the whole library

This repo ships a .claude-plugin manifest, so Claude Code can install all 57 skills at once. Plugin skills are invoked as /<plugin>:<skill>, so they never collide with your own.

/plugin marketplace add Varnan-Tech/opendirectory
/plugin

The folder is the same in every client that implements the format — 46 of them — so if yours is not above, only the destination changes.

Reproduced in full from Varnan-Tech/opendirectory/blob/62e437ab13408171805a87d16f5cb0151f96ea3c/skills/yc-intent-radar-skill/yc-jobs-scraper/SKILL.md, which is licensed MIT (repository). 361 words, 7 headings.

YC Jobs Scraper

This skill provides a robust architecture for scraping jobs from YCombinator and workatastartup.com. It is designed to run automatically, bypass login bottlenecks, and maintain state to never scrape duplicate jobs.

Architecture

The scraper uses a hybrid approach to maximize reliability and minimize bot detection:

  1. Authentication: scripts/auth.js uses Playwright to let a human log in once and saves the session to scripts/state.json.
  2. Database: scripts/db.js uses better-sqlite3 to manage scripts/jobs.db. It tracks every company_slug and job_id ever seen.
  3. Primary Extraction: scripts/scraper.js loads state.json, visits YC query URLs, and extracts company slugs from the hidden Inertia.js data-page JSON payload.
  4. Job Extraction (JSON): It then visits the authenticated company pages (/companies/[slug]) to extract jobs from the backend JSON payload to ensure we get the real job_id for accurate deduplication.
  5. Job Extraction (Fallback): If the JSON extraction fails, it falls back to parsing public HTML job cards from ycombinator.com/companies/[slug]/jobs.

Workflows

1. First-Time Setup

If this is the first time running the scraper in an environment, or if node_modules is missing:

cd @path/scripts
npm install
npx playwright install

2. Authentication (Manual Step)

If scripts/state.json is missing or expired, the scraper will fail. You must instruct the human user to run the authentication script manually:

cd @path/scripts
node auth.js

Tell the user a browser will open, and they must log in. Playwright will automatically save the cookies/tokens to state.json.

3. Running the Daily Scraper

To scrape for new companies and jobs:

cd @path/scripts
node scraper.js

This script will output exactly how many new companies and new jobs were found. Because of jobs.db, running it multiple times consecutively will result in 0 new jobs found.

4. Querying the Database

If you need to analyze the scraped data or view the companies/jobs, you can query scripts/jobs.db directly using better-sqlite3.

Example: Count Companies

cd @path/scripts
node -e "const db = require('better-sqlite3')('jobs.db'); console.log('Companies:', db.prepare('SELECT COUNT(*) as count FROM companies').get().count);"

Example: View Recent Jobs

cd @path/scripts
node -e "const db = require('better-sqlite3')('jobs.db'); const jobs = db.prepare('SELECT title, company_slug, location FROM jobs ORDER BY created_at DESC LIMIT 5').all(); console.table(jobs);"

Files bundled with it

These load only when the skill asks for them, so they cost nothing until it runs.

Other skills for the same job

Different authors, same problem. Matched on the words in the skill name, across every library in the catalogue except this one.

Need help setting it up?

This page tells you what yc-jobs-scraper does and what it needs. Cheetah builds the agent setup it runs inside: data, CRM, sequencing and the guardrails.

Book a call →

The directory stays free. There is nothing gated behind this.