Scraping

Web scraping and data extraction

Showing 481-504 of 708 skills
QuestForTech-Investments

pdf

by QuestForTech-Investments

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI Tools 6 9mo ago
QuestForTech-Investments

Playwright Browser Automation

by QuestForTech-Investments

Complete browser automation with Playwright. Auto-detects dev servers, writes clean test scripts to /tmp. Test pages, fill forms, take screenshots, check responsive design, validate UX, test login flows, check links, automate any browser task. Use when user wants to test websites, automate browser interactions, validate web functionality, or perform any browser-based testing.

Automation 6 9mo ago
thechandanbhagat

pdf

by thechandanbhagat

Work with PDF files - read, extract text/images/tables, create PDFs, merge, split, and convert PDFs. Use when the user asks to read, create, modify, or analyze PDF documents.

CLI Tools 8 7mo ago
DUZ1287

web-auto-form

by DUZ1287

JSON 驱动的浏览器表单自动化工具,为 AI Agent 提供原生 function-calling 集成,支持表单填写、条件分支、数据提取与 PII 脱敏

Debugging 2 3mo ago
alfredang

pdf

by alfredang

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI Tools 2 6mo ago
tankpkg

@tank/bdd-e2e-testing

by tankpkg

"BDD end-to-end testing against real systems. Covers web apps (Playwright), libraries (pytest-bdd + Docker), APIs, CLIs, message queues. Gherkin writing, step definitions, Page Objects, Screenplay, 3-layer architecture, CI/CD, multi-language (TypeScript, Python, Java, .NET). Triggers: BDD test, Gherkin, Cucumber, feature file, Given When Then, playwright-bdd, pytest-bdd, Behave, Cucumber-JVM, Serenity BDD, Reqnroll, Example Mapping, Three Amigos, living documentation, BDD setup, BDD architecture."

API Dev 1 6mo ago
Victory-Hugo

pdf

by Victory-Hugo

用于提取文本与表格、创建新 PDF、合并/拆分文档以及处理表单的综合 PDF 操作工具包。当需要填写 PDF 表单,或以编程方式批量处理、生成或分析 PDF 文档时使用。

CLI Tools 1 7mo ago
manastalukdar

e2e-generate

by manastalukdar

Generate end-to-end tests with Playwright browser automation

Code Gen 1 7mo ago
rarestg

cf-browser

by rarestg

Browse and scrape websites using Cloudflare's Browser Rendering REST API. Use when the agent needs to fetch rendered web content, extract structured data from pages, take screenshots, or scrape specific elements via CSS selectors. Triggers on tasks like "scrape this site", "get listings from this page", "extract data from this URL", "take a screenshot of this page", "browse this website", or any task requiring headless browser access to read, crawl, or extract information from live web pages. Also use when WebFetch is insufficient (JS-heavy sites, SPAs, pages requiring cookies, or when structured extraction is needed).

API Dev 1 7mo ago
ryanhudson

tapestry

by ryanhudson

This skill should be used when the user says "tapestry <URL>", "weave <URL>", "help me plan <URL>", "extract and plan <URL>", "make this actionable <URL>", or wants to extract content from a URL and create an action plan. Automatically detects content type (YouTube video, article, PDF) and orchestrates the full extract-to-plan workflow.

CLI Tools 7 7mo ago
ryanhudson

article-extractor

by ryanhudson

This skill should be used when the user wants to "download article", "extract article", "save blog post", "get article text", or provides a web URL and asks to extract the main content without ads, navigation, or clutter. Saves clean, readable text from web articles and blog posts.

CLI Tools 7 7mo ago
sumik5

developing-react

by sumik5

React 19.x development guide covering internals (rendering, reconciliation, Fiber), performance optimization (47+ react-doctor rules, memoization, bundle size), UI animation patterns (CSS transitions, easing, hover/touch), and React Testing Library (RTL queries, interactions, TDD patterns). Use when package.json contains 'react' (without 'next'), or when working on React-specific concerns in any framework. For Next.js-specific features (App Router, Server Components, Cache Components), use developing-nextjs instead. For E2E testing with Playwright, use testing-e2e-with-playwright. For general testing methodology, use testing-code.

Scraping 1 6mo ago
famaoai-creator

browser-navigator

by famaoai-creator

Automates browser actions using Playwright CLI. Can record, replay, and generate browser automation scenarios stored in the knowledge base. Useful for UI testing, data extraction, and visual auditing.

Code Gen 1 6mo ago
michelg10

PDF Processing

by michelg10

Comprehensive PDF manipulation toolkit for extracting text and tables,

CLI Tools 6 11mo ago
timequity

airflow-workflows

by timequity

Apache Airflow DAG design, operators, and scheduling best practices.

Processing 6 9mo ago
Nymbo

music-downloader

by Nymbo

This skill should be used when users need to download audio or music from online platforms like YouTube, SoundCloud, Spotify, or other streaming services. It provides yt-dlp and spotdl command templates for high-quality audio extraction, playlist downloads, metadata embedding, and multi-platform support.

CLI Tools 6 8mo ago
deletexiumu

x-ai-digest

by deletexiumu

Scrape AI-related posts from X platform's "For You" feed, generate daily digest with reply suggestions. Features real-time scraping, AI topic filtering, share card generation, multi-language output (ZH/EN/JA).

Code Gen 3 7mo ago
vibery-studio

pdf

by vibery-studio

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI Tools 3 8mo ago
Crumbgrabber

pdf-processing

by Crumbgrabber

Extract text and tables from PDF files, fill forms, merge documents.

Processing 3 8mo ago
lihaoze123

nix-packaging-best-practices

by lihaoze123

Best practices for packaging pre-compiled binaries (.deb, .rpm, .tar.gz, AppImage) for NixOS, handling library dependencies, or facing "library not found" errors with binary distributions

Debugging 3 7mo ago
Crumbgrabber

pdf

by Crumbgrabber

Comprehensive PDF manipulation toolkit for extracting text and tables,

CLI Tools 3 8mo ago
vibery-studio

webapp-testing

by vibery-studio

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

Automation 3 8mo ago
liauw-media

playwright-frontend-testing

by liauw-media

"Use when testing frontend applications. AI-assisted browser testing with Playwright MCP. Fast, deterministic, no vision models needed."

Scraping 3 9mo ago
breverdbidder

website-to-vite-scraper

by breverdbidder

Multi-provider website scraper that converts any website (including CSR/SPA) to deployable static sites. Uses Playwright, Apify RAG Browser, Crawl4AI, and Firecrawl for comprehensive scraping. Triggers on requests to clone, reverse-engineer, or convert websites.

Embeddings 5 8mo ago