Another week, another shake-up. The big story: a single-file behavioral skill born from one viral Karpathy post is now the most-starred Claude Code skill on GitHub, and an enterprise giant (Feishu/Lark) quietly planted a 16-million-install foothold on skills.sh. Coding skills dominated; the security shelf is nearly empty after last week's buying spree. Six categories, eighteen podiums, zero repeats from the last 30 days โ except one champion defending its crown and one major-update returner.
Methodology: Ranked from live skills.sh leaderboard data (all-time + 24h trending), the September 8 Firecrawl developer roundup, and cross-checked community sources. No skill qualifies if reviewed here in the past 30 days โ unless it's defending a #1 spot or shipped a major update.
| Rank | Skill | Developer | Stars / Installs | Format | Verdict |
|---|---|---|---|---|---|
| ๐ฅ | andrej-karpathy-skills | Forrest Chang (multica-ai) | 210k+ โญ | SKILL.md + CLAUDE.md | The behavioral guardrail everyone was waiting for |
| ๐ฅ | Superpowers | Jesse Vincent (obra) | 280k+ โญ, 25k+ forks | SKILL.md suite + commands + hooks | The full-SDLC heavyweight |
| ๐ฅ | Caveman | Julius Brussee | 100k+ โญ | SKILL.md + install.sh (4 modes) | Cheapest tokens on the market |
Why these ranked here
๐ฅ andrej-karpathy-skills โ Four hard rules distilled from Karpathy's January 2026 viral critique of AI coding, turned into the most-starred behavioral skill on GitHub (144k stars within weeks, now 210k+). It directly kills the three failure modes every agent user knows: silent wrong assumptions, 50-line solutions bloated into 500, and edits to code it was never supposed to touch. One file, zero dependencies, works in Claude Code, Cursor, and Copilot. This week's runaway #1.
๐ฅ Superpowers โ At 280k+ stars it's the biggest community-built skill library in the ecosystem, and nothing else chains the whole lifecycle: brainstorm โ worktree setup โ implementation plan โ fresh subagent per task with two-stage review โ TDD โ merge review. The RED-GREEN-REFACTOR discipline (it deletes code written before a failing test exists) is the strictest quality gate we've seen in any skill collection. It defends silver on maturity alone.
๐ฅ Caveman โ A 65% average output-token cut (range 22โ87%) with every technical fact preserved byte-for-byte, plus /caveman-compress shrinking your CLAUDE.md by ~46% permanently. Compatible with 30+ agents, and a March 2026 paper found brevity-constrained models actually improved accuracy by 26 points on some benchmarks. Last reviewed August 9, so it steps back into the ring fully eligible.
| Rank | Skill | Developer | Stars / Installs | Format | Verdict |
|---|---|---|---|---|---|
| ๐ฅ | lark-doc | Feishu (open.feishu.cn) | 704.3K installs (suite: 16.1M) | SKILL.md + Lark API connectors | Enterprise docs just got an agent lane |
| ๐ฅ | grill-me | Matt Pocock | 1.2M installs | SKILL.md (collection: 255k+ โญ) | Interviews you until your plan stops lying |
| ๐ฅ | teach | Matt Pocock | 670.9K installs | SKILL.md | Turns any codebase into a course |
Why these ranked here
๐ฅ lark-doc โ The sleeper of the year: Feishu's official skills occupy 22+ leaderboard slots totaling 16.1M installs, with lark-doc alone at 704.3K and sitting at #14 all-time. Create, edit, and analyze Lark/Feishu docs natively from your agent, and your agent suddenly speaks enterprise. First appearance on our leaderboard โ and it lands with gold.
๐ฅ teach โ 670.9K installs and climbing: point it at any codebase or concept and it produces structured teaching material โ a course-building engine inside your agent. It rounds out a Pocock family double on this podium and is quietly one of the highest-install non-Anthropic, non-Vercel skills in the ecosystem.
๐ฅ ask-matt โ Meta, but earned: a skill that routes your agent's hardest architecture questions to Pocock's documented decision frameworks, with 575.1K installs. Between grill-me, teach, ask-matt, code-review, and handoff, his collection holds five separate top-65 slots on skills.sh โ the most influential indie skill author in the ecosystem.
| Rank | Skill | Developer | Stars / Installs | Format | Verdict |
|---|---|---|---|---|---|
| ๐ฅ | frontend-design ๐ก๏ธ | Anthropic | 897.1K installs (110k+/week) | SKILL.md (official) | Defending champion, still untouchable |
| ๐ฅ | webapp-testing | Anthropic | skills.sh top-50 club | SKILL.md + Playwright scripts | UI bugs that static analysis can't see |
| ๐ฅ | artifacts-builder | Anthropic | anthropics/skills core | SKILL.md + build scripts | Complex React/Tailwind/shadcn interfaces |
Why these ranked here
๐ฅ frontend-design โ Defending its #1 from last week's leaderboard per the champion exception: 897.1K all-time installs, 110k+ weekly across Claude Code, Codex, and Gemini CLI. The banned-fonts list and commit-to-a-direction-first methodology remain the single best cure for purple-gradient AI slop. Until something beats the standard-setter at its own game, the crown stays put.
๐ฅ webapp-testing โ Official Anthropic skill that hands Claude a real Playwright-driven browser to test your local app: auth flows, JS-rendered content, form validation โ while you watch. It catches the exact class of bugs (timing, JS errors, dead interactions) that code review misses. The weekly-install data and first-party maintenance earn silver.
๐ฅ artifacts-builder โ The workhorse inside anthropics/skills for building complex HTML interfaces with React/Tailwind/shadcn scaffolding done right. Less flashy than frontend-design, but it's the skill agents reach for when a design needs to become a working component. Solid bronze on reliability.
| Rank | Skill | Developer | Stars / Installs | Format | Verdict |
|---|---|---|---|---|---|
| ๐ฅ | hyperframes-cli ๐ | HeyGen (heygen-com) | 589.7K installs, #6 on 24h trending (19.8K) | SKILL.md + CLI + registry | Major update โ now a full media framework |
| ๐ฅ | ai-music | GenMedia Labs | 16.9K in 24h, top-20 trending | SKILL.md (multi-model suite) | Text-to-music at agent speed |
| ๐ฅ | yt-dlp-ffmpeg-media-stack | Open Claude Workshop | Trending 24h | SKILL.md + pipeline scripts | The boring-but-bulletproof backbone |
Why these ranked here
๐ฅ hyperframes-cli โ Major-update exception granted: since our August 28 review, HeyGen expanded HyperFrames into a multi-skill framework (cli, core, registry, keyframes, media-use) that now dominates the 24h trending board with five simultaneous entries. 589.7K all-time installs and climbing fast. From "one video skill" to "an agentic media studio" โ that's a rebuild worth gold.
๐ฅ ai-music โ GenMedia Labs' suite is all over this week's trending: ai-music at 16.9K installs in 24h, alongside wan-3-0-prime-reference-to-video (16.8K) and seedance-2-5-reference-to-video (16.6K). Multi-model, schema-driven generation makes it the strongest pure music entry we've tested this quarter.
๐ฅ yt-dlp-ffmpeg-media-stack โ No hype, all plumbing: a SKILL.md wrapper that gives agents competent download-transcode-subtitle pipelines through yt-dlp and FFmpeg. Media agents die without this layer, and this is the cleanest implementation we've found trending this week.
| Rank | Skill | Developer | Stars / Installs | Format | Verdict |
|---|---|---|---|---|---|
| ๐ฅ | deep-research | samber | Installs not published | SKILL.md + parallel-search scripts | The most rigorous research workflow in a folder |
| ๐ฅ | deep-research | affaan-m (everything-claude-code) | 4.6โ (38 ratings) | SKILL.md + subagent team config | Parallel agents, publication-grade output |
| ๐ฅ | deep-research-agent | Qodex AI | LobeHub featured | SKILL.md + workflow scripts | Credibility scoring built in |
Why these ranked here
๐ฅ samber/deep-research โ Broad parallel web searches, multi-source validation, confidence tracking, and a cited Markdown report at the end โ with 11 structured research types from market analysis to tech evaluation. It's the skill most closely aligned with how serious research should actually work: evidence-first, sources attached, uncertainty flagged. Never reviewed here before; instant podium.
๐ฅ deep-research (everything-claude-code) โ Spawns a team of agents that each search, read, and return findings for competitive analysis and market/tech evaluations, with a 4.6 rating across 38 community reviews. The multi-agent split gives it real throughput on big questions. Silver on parallelism.
๐ฅ deep-research-agent (Qodex AI) โ End-to-end automation covering planning, multi-channel source gathering, and โ its differentiator โ explicit credibility evaluation before synthesis. Featured on LobeHub's skills marketplace. Bronze for bringing source-quality discipline to the party.
| Rank | Skill | Developer | Stars / Installs | Format | Verdict |
|---|---|---|---|---|---|
| ๐ฅ | grill-me ๐ก๏ธ | Matt Pocock | 1.2M installs | SKILL.md | Defending its crown โ plan stress-testing as a sport |
| ๐ฅ | andrej-karpathy-skills | Forrest Chang (multica-ai) | 210k+ โญ | SKILL.md + CLAUDE.md | Behavioral rules double as acceptance criteria |
| ๐ฅ | code-review | Matt Pocock | 571.1K installs | SKILL.md | The quality gate every agent needs |
Why these ranked here
๐ฅ grill-me โ Defending champion from its era at the top of the Testing & Eval conversation: at 1.2M installs it's the highest-install quality skill in the ecosystem, and it went viral again on X this month for fixing agentic coding's most expensive failure โ charging ahead on wrong assumptions. It interrogates your plan until shared understanding is reached, reads the codebase first when it can, and surfaces dependency chains before they become bugs. A testing skill for the only test that matters: does anyone actually know what they're building?
๐ฅ andrej-karpathy-skills โ The behavioral four-rule set earns a second podium this week because it doubles as an acceptance harness: silent assumptions blocked, scope creep cut, orthogonal edits refused. That's regression prevention for agent behavior โ the closest thing eval engineering has to a smoke test that runs on every prompt. Crossover appeal from its Coding gold, and it costs 210k+ stars' worth of community validation.
๐ฅ code-review โ Matt Pocock's structured review pass checks simplification opportunities, extraction candidates, and test gaps with 571.1K installs of quiet adoption. It's the "measure twice" skill for everything agents write โ deterministic checklist behavior in a category full of vibes.
npx skills add anything from an org you don't recognize.Ranked by AgentSkillReview. Data sourced from skills.sh, the Firecrawl developer skills roundup (Sep 8, 2026), and community forums. Dedup window: 30 days against 84 previously reviewed skills.