Scraping

Web scraping and data extraction

Showing 25-48 of 708 skills
github

pdftk-server

by github

'Skill for using the command-line tool pdftk (PDFtk Server) for working with PDF files. Use when asked to merge PDFs, split PDFs, rotate pages, encrypt or decrypt PDFs, fill PDF forms, apply watermarks, stamp overlays, extract metadata, burst documents into pages, repair corrupted PDFs, attach or extract files, or perform any PDF manipulation from the command line.'

CLI Tools 38.3K 6mo ago
alchaincyf

huashu-design

by alchaincyf

设计哲学顾问,从20种风格中推荐3个方向并生成视觉Demo和AI提示词。当用户提到"设计风格"、"设计方向"、"配色方案"、"视觉风格"、"设计评审"、"推荐风格"时使用。

Agents 1.4K 6mo ago
alirezarezvani

senior-qa

by alirezarezvani

This skill should be used when the user asks to "generate tests", "write unit tests", "analyze test coverage", "scaffold E2E tests", "set up Playwright", "configure Jest", "implement testing patterns", or "improve test quality". Use for React/Next.js testing with Jest, React Testing Library, and Playwright.

Code Gen 25.2K 7mo ago
tavily-ai

tavily-best-practices

by tavily-ai

"Build production-ready Tavily integrations with best practices baked in. Reference documentation for developers using coding assistants (Claude Code, Cursor, etc.) to implement web search, content extraction, crawling, and research in agentic workflows, RAG systems, or autonomous agents."

Academic 462 5mo ago
tavily-ai

tavily-best-practices

by tavily-ai

"Build production-ready Tavily integrations with best practices baked in. Reference documentation for developers using coding assistants (Claude Code, Cursor, etc.) to implement web search, content extraction, crawling, and research in agentic workflows, RAG systems, or autonomous agents."

Embeddings 462 7mo ago
tavily-ai

tavily-map

by tavily-ai

Discover and list all URLs on a website without extracting content, via the Tavily CLI. Use this skill when the user wants to find a specific page on a large site, list all URLs, see the site structure, find where something is on a domain, or says "map the site", "find the URL for", "what pages are on", "list all pages", or "site structure". Faster than crawling — returns URLs only. Essential when you know the site but not the exact page. Combine with extract for targeted content retrieval.

CLI Tools 462 5mo ago
tavily-ai

extract

by tavily-ai

"Extract content from specific URLs using Tavily's extraction API. Returns clean markdown/text from web pages. Use when you have specific URLs and need their content without writing code."

Processing 462 7mo ago
tavily-ai

crawl

by tavily-ai

"Crawl any website and save pages as local markdown files. Use when you need to download documentation, knowledge bases, or web content for offline access or analysis. No code required - just provide a URL."

Processing 462 7mo ago
bitwize-music-studio

document-hunter

by bitwize-music-studio

Searches and retrieves documents from free public sources using automated browser navigation. Use when research needs primary source documents like court filings, government reports, or public records.

File Ops 454 6mo ago
bitwize-music-studio

setup

by bitwize-music-studio

Detects your Python environment and guides you through installing plugin dependencies. Use on first-time setup or when MCP server fails to start.

CLI Tools 454 7mo ago
proffesor-for-testing

compatibility-testing

by proffesor-for-testing

"Cross-browser, cross-platform, and cross-device compatibility testing ensuring consistent experience across environments. Use when validating browser support, testing responsive design, or ensuring platform compatibility."

Responsive 466 6mo ago
xstongxue

paper-write

by xstongxue

本科与硕士学位论文全流程撰写辅助。支持大纲审核(理工科/文科)、结构仿写(通用章节/实验章节/绪论/摘要)、参考文献获取、融合、润色、缩写、扩写、防 AIGC、中英互译、结构化信息提取。当用户提到论文撰写、大纲审核、论文章节仿写、参考文献、论文润色、防 AIGC、论文翻译时使用。

Code Review 2.6K 6mo ago
BasedHardware

self-improvement

by BasedHardware

"Meta-skill for analyzing PRs, issues, and user interactions to improve Cursor rules and skills automatically"

Code Gen 13.4K 7mo ago
Arize-ai

phoenix-playwright-tests

by Arize-ai

Write Playwright E2E tests for the Phoenix AI observability platform. Use when creating, updating, or debugging Playwright tests, or when the user asks about testing UI features, writing E2E tests, or automating browser interactions for Phoenix.

Scraping 11.2K 6mo ago
pbakaus

extract

by pbakaus

Extract and consolidate reusable components, design tokens, and patterns into your design system. Identifies opportunities for systematic reuse and enriches your component library.

Code Gen 63K 6mo ago
cyberkaida

ctf-rev

by cyberkaida

Solve CTF reverse engineering challenges using systematic analysis to find flags, keys, or passwords. Use for crackmes, binary bombs, key validators, obfuscated code, algorithm recovery, or any challenge requiring program comprehension to extract hidden information.

Processing 818 10mo ago
Galaxy-Dawn

kaggle-learner

by Galaxy-Dawn

This skill should be used when the user asks to "learn from Kaggle", "study Kaggle solutions", "analyze Kaggle competitions", or mentions Kaggle competition URLs. Provides access to extracted knowledge from winning Kaggle solutions across NLP, CV, time series, tabular, and multimodal domains.

Processing 5.2K 7mo ago
Galaxy-Dawn

webapp-testing

by Galaxy-Dawn

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

Automation 5.2K 6mo ago
smallnest

summarize

by smallnest

Summarize or extract text/transcripts from URLs, podcasts, and local files (great fallback for “transcribe this YouTube/video”).

Processing 601 6mo ago
HKUDS

summarize

by HKUDS

Summarize or extract text/transcripts from URLs, podcasts, and local files (great fallback for “transcribe this YouTube/video”).

Processing 47.4K 7mo ago
comet-ml

playwright-e2e

by comet-ml

Playwright E2E test generation workflow for Opik. Use when generating, fixing, or planning automated tests in tests_end_to_end/.

Agents 21.6K 6mo ago
NeverSight

extract-transcripts

by NeverSight

Extract readable transcripts from Claude Code and Codex CLI session JSONL files

Auth 203 7mo ago
NeverSight

data-processing

by NeverSight

"Process JSON with jq and YAML/TOML with yq. Filter, transform, query structured data efficiently. Triggers on: parse JSON, extract from YAML, query config, Docker Compose, K8s manifests, GitHub Actions workflows, package.json, filter data."

Processing 203 7mo ago
Project-N-E-K-O

pdf

by Project-N-E-K-O

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI Tools 2.7K 7mo ago