抓取

网页抓取与数据提取

显示 313-336 / 共 708 个技能
guia-matthieu

pdf-extractor

guia-matthieu

"Extract text, tables, and images from PDFs. Use when: extracting data from reports; converting PDF tables to CSV; pulling images from presentations; processing research papers; batch converting PDFs to text"

CLI 工具 145 6个月前
guia-matthieu

web-scraper

guia-matthieu

"Extract structured data from websites. Use when: collecting competitor pricing; scraping product listings; extracting contact information; gathering research data; monitoring website changes"

数据处理 145 6个月前
phamquiluan

pdf

phamquiluan

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI 工具 15 7个月前
tao3k

crawl4ai

tao3k

Use when crawling web pages, extracting markdown content, or scraping website data with intelligent chunking and skeleton planning. Use when the user provides a URL or link to fetch or crawl.

文档生成 15 6个月前
oakoss

e2e-testing

oakoss

'E2E test architecture and patterns with Playwright. Use when designing test suites, structuring Page Object Models, planning CI sharding strategies, setting up authentication flows, or organizing tests with tags and annotations. Use for test architecture, accessibility auditing with axe-core, network mocking strategies, visual regression workflows, HAR replay, and storageState authentication patterns. For Playwright API details, browser automation, or web scraping, use the playwright skill instead.'

无障碍 14 6个月前
samhvw8

chrome-devtools

samhvw8

"Browser automation via Puppeteer CLI scripts (JSON output). Capabilities: screenshots, PDF generation, web scraping, form automation, network monitoring, performance profiling, JavaScript debugging, headless browsing. Actions: screenshot, scrape, automate, test, profile, monitor, debug browser. Keywords: Puppeteer, headless Chrome, screenshot, PDF, web scraping, form fill, click, navigate, network traffic, performance audit, Lighthouse, console logs, DOM manipulation, element selector, wait, scroll, automation script. Use when: taking screenshots, generating PDFs from web, scraping websites, automating form submissions, monitoring network requests, profiling page performance, debugging JavaScript, testing web UIs."

CLI 工具 14 7个月前
LeastBit

webapp-testing

LeastBit

使用 Playwright 与本地 Web 应用程序交互及进行测试的工具包。支持验证前端功能、调试 UI 行为、捕获浏览器截图以及查看浏览器日志。

CI/CD 566 7个月前
Bbeierle12

pdf

Bbeierle12

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI 工具 8 8个月前
parhumm

qa-test-generate

parhumm

Generate runnable Vitest and Playwright test files from BDD test cases and scaffold code. Use when generating test implementations.

代码生成 25 6个月前
blacktop

ipsw

blacktop

Apple firmware and binary reverse engineering with the ipsw CLI tool. Use when analyzing iOS/macOS binaries, disassembling functions in dyld_shared_cache, dumping Objective-C headers from private frameworks, downloading IPSWs or kernelcaches, extracting entitlements, analyzing Mach-O files, or researching Apple security. Triggers on requests involving Apple RE, iOS internals, kernel analysis, KEXT extraction, or vulnerability research on Apple platforms.

CLI 工具 81 8个月前
kachkaev

fractal-tree-file-structure

kachkaev

"Guides file and folder organization using fractal tree structure: self-similar directories, encapsulation, shared/ folders, mini-library pattern, and kebab-case naming. Use when creating, moving, or renaming files, extracting modules into subdirectories, deciding where shared code belongs, or checking import boundaries."

文件操作 7 6个月前
op7418

video-wrapper

op7418

为访谈视频添加综艺特效(花字、卡片、人物条、章节标题等)。支持 4 种视觉主题,先分析字幕内容生成建议供用户审批,再渲染视频。

数据处理 332 7个月前
JinFanZheng

data-base

JinFanZheng

Data acquisition for web scraping and data collection. Use when user needs "爬取数据/抓取网页/scrape data". Outputs structured JSON/CSV for analysis.

数据处理 79 7个月前
salvo-rs

salvo-data-extraction

salvo-rs

Extract and validate data from requests including JSON, forms, query parameters, and path parameters. Use for handling user input and API payloads.

数据处理 20 7个月前
ils15

playwright-e2e-testing

ils15

Write end-to-end tests with Playwright for web applications. Includes fixtures, page objects, test templates, visual regression testing, and accessibility audits.

无障碍 10 7个月前
run-llama

Extract structured data from unstructured files (PDF, PPTX, DOCX...)

run-llama

Invoke this skill BEFORE implementing any structured data extraction from documents to learn the correct llama_cloud_services API usage. Required reading before writing extraction code. Requires llama_cloud_services package and LLAMA_CLOUD_API_KEY as an environment variable.

API 开发 178 10个月前
mgd34msu

precision-mastery

mgd34msu

"ALWAYS load before starting any task. Maximizes token efficiency for all file operations, searches, and command execution. Covers extract modes (content, outline, symbols, ast, lines), verbosity tuning, multi-file batching, and discover tool orchestration. Ensures agents spend tokens on outcomes, not overhead."

文件操作 6 6个月前
buainoai

remotion-best-practices

buainoai

Remotion 最佳实践 - 使用 React 创建视频

动画 101 7个月前
oldwinter

firecrawl

oldwinter

Official Firecrawl CLI skill for web scraping, search, crawling, and browser automation. Returns clean LLM-optimized markdown. USE FOR: - Web search and research - Scraping pages, docs, and articles - Site mapping and bulk content extraction - Browser automation for interactive pages Must be pre-installed and authenticated. See rules/install.md for setup, rules/security.md for output handling.

数据处理 6 6个月前
v1-io

e2e-testing

v1-io

Use when implementing E2E tests, debugging flaky tests, testing web applications with Playwright, or establishing E2E testing standards. Triggers on "e2e test", "end-to-end", "Playwright", "flaky test", "browser test".

认证鉴权 6 7个月前
kernel

kernel-typescript-sdk

kernel

Build browser automation scripts using the Kernel TypeScript SDK with Playwright, CDP, and remote browser management.

代码生成 9 6个月前
kernel

kernel-python-sdk

kernel

Build browser automation scripts using the Kernel Python SDK with Playwright and remote browser management.

代码生成 9 6个月前
shiqkuangsan

tooyoung:ink-reader

shiqkuangsan

"Intelligently read any URL content with auto platform detection and fallback strategies. Supports WeChat, Zhihu, Bilibili, Toutiao, Weibo, Xiaohongshu, Douyin, X/Twitter, and generic websites. Trigger words: read url, read link, read this, fetch url, grab content, ink-reader"

抓取 17 7个月前
ImGoodBai

pdf

ImGoodBai

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI 工具 196 7个月前