抓取

网页抓取与数据提取

显示 25-48 / 共 708 个技能
github

pdftk-server

github

'Skill for using the command-line tool pdftk (PDFtk Server) for working with PDF files. Use when asked to merge PDFs, split PDFs, rotate pages, encrypt or decrypt PDFs, fill PDF forms, apply watermarks, stamp overlays, extract metadata, burst documents into pages, repair corrupted PDFs, attach or extract files, or perform any PDF manipulation from the command line.'

CLI 工具 3.8万 6个月前
alchaincyf

huashu-design

alchaincyf

设计哲学顾问,从20种风格中推荐3个方向并生成视觉Demo和AI提示词。当用户提到"设计风格"、"设计方向"、"配色方案"、"视觉风格"、"设计评审"、"推荐风格"时使用。

智能体 1408 6个月前
alirezarezvani

senior-qa

alirezarezvani

This skill should be used when the user asks to "generate tests", "write unit tests", "analyze test coverage", "scaffold E2E tests", "set up Playwright", "configure Jest", "implement testing patterns", or "improve test quality". Use for React/Next.js testing with Jest, React Testing Library, and Playwright.

代码生成 2.5万 7个月前
tavily-ai

tavily-best-practices

tavily-ai

"Build production-ready Tavily integrations with best practices baked in. Reference documentation for developers using coding assistants (Claude Code, Cursor, etc.) to implement web search, content extraction, crawling, and research in agentic workflows, RAG systems, or autonomous agents."

学术 462 5个月前
tavily-ai

tavily-best-practices

tavily-ai

"Build production-ready Tavily integrations with best practices baked in. Reference documentation for developers using coding assistants (Claude Code, Cursor, etc.) to implement web search, content extraction, crawling, and research in agentic workflows, RAG systems, or autonomous agents."

向量嵌入 462 7个月前
tavily-ai

tavily-map

tavily-ai

Discover and list all URLs on a website without extracting content, via the Tavily CLI. Use this skill when the user wants to find a specific page on a large site, list all URLs, see the site structure, find where something is on a domain, or says "map the site", "find the URL for", "what pages are on", "list all pages", or "site structure". Faster than crawling — returns URLs only. Essential when you know the site but not the exact page. Combine with extract for targeted content retrieval.

CLI 工具 462 5个月前
tavily-ai

extract

tavily-ai

"Extract content from specific URLs using Tavily's extraction API. Returns clean markdown/text from web pages. Use when you have specific URLs and need their content without writing code."

数据处理 462 7个月前
tavily-ai

crawl

tavily-ai

"Crawl any website and save pages as local markdown files. Use when you need to download documentation, knowledge bases, or web content for offline access or analysis. No code required - just provide a URL."

数据处理 462 7个月前
bitwize-music-studio

document-hunter

bitwize-music-studio

Searches and retrieves documents from free public sources using automated browser navigation. Use when research needs primary source documents like court filings, government reports, or public records.

文件操作 454 6个月前
bitwize-music-studio

setup

bitwize-music-studio

Detects your Python environment and guides you through installing plugin dependencies. Use on first-time setup or when MCP server fails to start.

CLI 工具 454 7个月前
proffesor-for-testing

compatibility-testing

proffesor-for-testing

"Cross-browser, cross-platform, and cross-device compatibility testing ensuring consistent experience across environments. Use when validating browser support, testing responsive design, or ensuring platform compatibility."

响应式 466 6个月前
xstongxue

paper-write

xstongxue

本科与硕士学位论文全流程撰写辅助。支持大纲审核(理工科/文科)、结构仿写(通用章节/实验章节/绪论/摘要)、参考文献获取、融合、润色、缩写、扩写、防 AIGC、中英互译、结构化信息提取。当用户提到论文撰写、大纲审核、论文章节仿写、参考文献、论文润色、防 AIGC、论文翻译时使用。

代码评审 2640 6个月前
BasedHardware

self-improvement

BasedHardware

"Meta-skill for analyzing PRs, issues, and user interactions to improve Cursor rules and skills automatically"

代码生成 1.3万 7个月前
Arize-ai

phoenix-playwright-tests

Arize-ai

Write Playwright E2E tests for the Phoenix AI observability platform. Use when creating, updating, or debugging Playwright tests, or when the user asks about testing UI features, writing E2E tests, or automating browser interactions for Phoenix.

抓取 1.1万 6个月前
pbakaus

extract

pbakaus

Extract and consolidate reusable components, design tokens, and patterns into your design system. Identifies opportunities for systematic reuse and enriches your component library.

代码生成 6.3万 6个月前
cyberkaida

ctf-rev

cyberkaida

Solve CTF reverse engineering challenges using systematic analysis to find flags, keys, or passwords. Use for crackmes, binary bombs, key validators, obfuscated code, algorithm recovery, or any challenge requiring program comprehension to extract hidden information.

数据处理 818 10个月前
Galaxy-Dawn

kaggle-learner

Galaxy-Dawn

This skill should be used when the user asks to "learn from Kaggle", "study Kaggle solutions", "analyze Kaggle competitions", or mentions Kaggle competition URLs. Provides access to extracted knowledge from winning Kaggle solutions across NLP, CV, time series, tabular, and multimodal domains.

数据处理 5214 7个月前
Galaxy-Dawn

webapp-testing

Galaxy-Dawn

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

自动化 5214 6个月前
smallnest

summarize

smallnest

Summarize or extract text/transcripts from URLs, podcasts, and local files (great fallback for “transcribe this YouTube/video”).

数据处理 601 6个月前
HKUDS

summarize

HKUDS

Summarize or extract text/transcripts from URLs, podcasts, and local files (great fallback for “transcribe this YouTube/video”).

数据处理 4.7万 7个月前
comet-ml

playwright-e2e

comet-ml

Playwright E2E test generation workflow for Opik. Use when generating, fixing, or planning automated tests in tests_end_to_end/.

智能体 2.2万 6个月前
NeverSight

extract-transcripts

NeverSight

Extract readable transcripts from Claude Code and Codex CLI session JSONL files

认证鉴权 203 7个月前
NeverSight

data-processing

NeverSight

"Process JSON with jq and YAML/TOML with yq. Filter, transform, query structured data efficiently. Triggers on: parse JSON, extract from YAML, query config, Docker Compose, K8s manifests, GitHub Actions workflows, package.json, filter data."

数据处理 203 7个月前
Project-N-E-K-O

pdf

Project-N-E-K-O

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI 工具 2682 7个月前