抓取

网页抓取与数据提取

显示 481-504 / 共 708 个技能
QuestForTech-Investments

pdf

QuestForTech-Investments

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI 工具 6 9个月前
QuestForTech-Investments

Playwright Browser Automation

QuestForTech-Investments

Complete browser automation with Playwright. Auto-detects dev servers, writes clean test scripts to /tmp. Test pages, fill forms, take screenshots, check responsive design, validate UX, test login flows, check links, automate any browser task. Use when user wants to test websites, automate browser interactions, validate web functionality, or perform any browser-based testing.

自动化 6 9个月前
thechandanbhagat

pdf

thechandanbhagat

Work with PDF files - read, extract text/images/tables, create PDFs, merge, split, and convert PDFs. Use when the user asks to read, create, modify, or analyze PDF documents.

CLI 工具 8 7个月前
DUZ1287

web-auto-form

DUZ1287

JSON 驱动的浏览器表单自动化工具,为 AI Agent 提供原生 function-calling 集成,支持表单填写、条件分支、数据提取与 PII 脱敏

调试 2 3个月前
alfredang

pdf

alfredang

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI 工具 2 6个月前
tankpkg

@tank/bdd-e2e-testing

tankpkg

"BDD end-to-end testing against real systems. Covers web apps (Playwright), libraries (pytest-bdd + Docker), APIs, CLIs, message queues. Gherkin writing, step definitions, Page Objects, Screenplay, 3-layer architecture, CI/CD, multi-language (TypeScript, Python, Java, .NET). Triggers: BDD test, Gherkin, Cucumber, feature file, Given When Then, playwright-bdd, pytest-bdd, Behave, Cucumber-JVM, Serenity BDD, Reqnroll, Example Mapping, Three Amigos, living documentation, BDD setup, BDD architecture."

API 开发 1 6个月前
Victory-Hugo

pdf

Victory-Hugo

用于提取文本与表格、创建新 PDF、合并/拆分文档以及处理表单的综合 PDF 操作工具包。当需要填写 PDF 表单,或以编程方式批量处理、生成或分析 PDF 文档时使用。

CLI 工具 1 7个月前
manastalukdar

e2e-generate

manastalukdar

Generate end-to-end tests with Playwright browser automation

代码生成 1 7个月前
rarestg

cf-browser

rarestg

Browse and scrape websites using Cloudflare's Browser Rendering REST API. Use when the agent needs to fetch rendered web content, extract structured data from pages, take screenshots, or scrape specific elements via CSS selectors. Triggers on tasks like "scrape this site", "get listings from this page", "extract data from this URL", "take a screenshot of this page", "browse this website", or any task requiring headless browser access to read, crawl, or extract information from live web pages. Also use when WebFetch is insufficient (JS-heavy sites, SPAs, pages requiring cookies, or when structured extraction is needed).

API 开发 1 7个月前
ryanhudson

tapestry

ryanhudson

This skill should be used when the user says "tapestry <URL>", "weave <URL>", "help me plan <URL>", "extract and plan <URL>", "make this actionable <URL>", or wants to extract content from a URL and create an action plan. Automatically detects content type (YouTube video, article, PDF) and orchestrates the full extract-to-plan workflow.

CLI 工具 7 7个月前
ryanhudson

article-extractor

ryanhudson

This skill should be used when the user wants to "download article", "extract article", "save blog post", "get article text", or provides a web URL and asks to extract the main content without ads, navigation, or clutter. Saves clean, readable text from web articles and blog posts.

CLI 工具 7 7个月前
sumik5

developing-react

sumik5

React 19.x development guide covering internals (rendering, reconciliation, Fiber), performance optimization (47+ react-doctor rules, memoization, bundle size), UI animation patterns (CSS transitions, easing, hover/touch), and React Testing Library (RTL queries, interactions, TDD patterns). Use when package.json contains 'react' (without 'next'), or when working on React-specific concerns in any framework. For Next.js-specific features (App Router, Server Components, Cache Components), use developing-nextjs instead. For E2E testing with Playwright, use testing-e2e-with-playwright. For general testing methodology, use testing-code.

抓取 1 6个月前
famaoai-creator

browser-navigator

famaoai-creator

Automates browser actions using Playwright CLI. Can record, replay, and generate browser automation scenarios stored in the knowledge base. Useful for UI testing, data extraction, and visual auditing.

代码生成 1 6个月前
michelg10

PDF Processing

michelg10

Comprehensive PDF manipulation toolkit for extracting text and tables,

CLI 工具 6 11个月前
timequity

airflow-workflows

timequity

Apache Airflow DAG design, operators, and scheduling best practices.

数据处理 6 9个月前
Nymbo

music-downloader

Nymbo

This skill should be used when users need to download audio or music from online platforms like YouTube, SoundCloud, Spotify, or other streaming services. It provides yt-dlp and spotdl command templates for high-quality audio extraction, playlist downloads, metadata embedding, and multi-platform support.

CLI 工具 6 8个月前
deletexiumu

x-ai-digest

deletexiumu

Scrape AI-related posts from X platform's "For You" feed, generate daily digest with reply suggestions. Features real-time scraping, AI topic filtering, share card generation, multi-language output (ZH/EN/JA).

代码生成 3 7个月前
vibery-studio

pdf

vibery-studio

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI 工具 3 8个月前
Crumbgrabber

pdf-processing

Crumbgrabber

Extract text and tables from PDF files, fill forms, merge documents.

数据处理 3 8个月前
lihaoze123

nix-packaging-best-practices

lihaoze123

Best practices for packaging pre-compiled binaries (.deb, .rpm, .tar.gz, AppImage) for NixOS, handling library dependencies, or facing "library not found" errors with binary distributions

调试 3 7个月前
Crumbgrabber

pdf

Crumbgrabber

Comprehensive PDF manipulation toolkit for extracting text and tables,

CLI 工具 3 8个月前
vibery-studio

webapp-testing

vibery-studio

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

自动化 3 8个月前
liauw-media

playwright-frontend-testing

liauw-media

"Use when testing frontend applications. AI-assisted browser testing with Playwright MCP. Fast, deterministic, no vision models needed."

抓取 3 9个月前
breverdbidder

website-to-vite-scraper

breverdbidder

Multi-provider website scraper that converts any website (including CSR/SPA) to deployable static sites. Uses Playwright, Apify RAG Browser, Crawl4AI, and Firecrawl for comprehensive scraping. Triggers on requests to clone, reverse-engineer, or convert websites.

向量嵌入 5 8个月前