抓取

网页抓取与数据提取

显示 121-144 / 共 708 个技能
forcedotcom

playwright-e2e

forcedotcom

writing, running, and debugging Playwright tests. working with their output from github actions

CI/CD 1033 6个月前
aiskillstore

sitemapkit

aiskillstore

Discover and extract sitemaps from any website using SitemapKit. Use this skill whenever the user wants to find pages on a website, get a list of URLs from a domain, audit a site's structure, crawl a sitemap, check what pages exist on a site, or do anything involving sitemaps or site URL discovery — even if they don't explicitly say "sitemap". Requires the sitemapkit MCP server configured with a valid SITEMAPKIT_API_KEY.

抓取 414 5个月前
tavily-ai

tavily-extract

tavily-ai

Extract clean markdown or text content from specific URLs via the Tavily CLI. Use this skill when the user has one or more URLs and wants their content, says "extract", "grab the content from", "pull the text from", "get the page at", "read this webpage", or needs clean text from web pages. Handles JavaScript-rendered pages, returns LLM-optimized markdown, and supports query-focused chunking for targeted extraction. Can process up to 20 URLs in a single call.

CLI 工具 462 5个月前
openakita

browser-get-content

openakita

Extract page content and element text from current webpage. When you need to read page information, get element values, scrape data, or verify page content.

数据处理 1973 7个月前
WenJunDuan

webapp-testing

WenJunDuan

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

自动化 186 7个月前
tradingstrategy-ai

extract-vault-protocol-logo

tradingstrategy-ai

Extract a logo for vault protocol metadata

Docker 825 8个月前
tradingstrategy-ai

extract-project-logo

tradingstrategy-ai

Extract a project's logo from its website, brand kit, or other sources

代码评审 825 7个月前
adobe

scrape-webpage

adobe

Scrape webpage content, extract metadata, download images, and prepare for import/migration to AEM Edge Delivery Services. Returns analysis JSON with paths, metadata, cleaned HTML, and local images.

Docker 170 7个月前
jaechang-hits

histolab-wsi-processing

jaechang-hits

"Whole slide image processing for digital pathology. Tissue detection, tile extraction (random, grid, score-based), filter pipelines for H&E/IHC preprocessing. Use for dataset preparation, tile-based deep learning, and slide quality assessment. For advanced spatial proteomics or multiplexed imaging use pathml."

分析 344 6个月前
aiskillstore

webapp-testing

aiskillstore

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

自动化 414 7个月前
aiskillstore

extract-transcripts

aiskillstore

Extract readable transcripts from Claude Code and Codex CLI session JSONL files

认证鉴权 414 7个月前
phodal

pdf

phodal

Extract and analyze information from PDF documents

代码评审 4534 8个月前
QuixiAI

email-digest

QuixiAI

Digest and ingest emails into memory, surfacing important threads and action items

代码生成 607 6个月前
freekmurze

pdf

freekmurze

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

CLI 工具 1000 7个月前
bear2u

web-to-markdown

bear2u

웹페이지 URL을 입력받아 마크다운 형태로 변환하여 저장합니다. 웹 문서를 로컬 마크다운 파일로 아카이빙하거나 정리할 때 유용합니다.

提示词 916 10个月前
bear2u

web-to-markdown

bear2u

웹페이지 URL을 입력받아 마크다운 형태로 변환하여 저장합니다. 웹 문서를 로컬 마크다운 파일로 아카이빙하거나 정리할 때 유용합니다.

提示词 916 10个月前
colonelpanic8

playwright-cli

colonelpanic8

Automate browser interactions from the shell using Playwright via the playwright-cli command (open/goto/snapshot/click/type/screenshot, tabs/storage/network). Use when you need deterministic browser automation for web testing, form filling, screenshots/PDFs, or data extraction.

认证鉴权 221 6个月前
openakita

webapp-testing

openakita

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

自动化 1973 7个月前
antfu

vue-testing-best-practices

antfu

Use for Vue.js testing. Covers Vitest, Vue Test Utils, component testing, mocking, testing patterns, and Playwright for E2E testing.

抓取 5801 7个月前
yzfly

douyin-video

yzfly

"抖音无水印视频下载和文案提取工具. 从抖音分享链接获取无水印视频下载链接, 下载视频, 提取视频中的语音文案并自动保存到文件. 适用场景包括获取抖音视频信息, 下载无水印视频, 批量提取视频文案. 当用户需要处理抖音视频链接或提取视频内容时触发."

API 开发 1258 7个月前
GPTomics

bio-atac-seq-nucleosome-positioning

GPTomics

Extract nucleosome positions from ATAC-seq data using NucleoATAC, ATACseqQC, and fragment analysis. Use when analyzing chromatin organization, identifying nucleosome-free regions at promoters, or characterizing nucleosome occupancy patterns from ATAC-seq fragment size distributions.

CLI 工具 1194 6个月前
GPTomics

bio-atac-seq-footprinting

GPTomics

Detect transcription factor binding sites through footprinting analysis in ATAC-seq data using TOBIAS. Use when identifying TF occupancy patterns within accessible regions, as TF binding protects DNA from Tn5 cutting.

CLI 工具 1194 6个月前
GPTomics

bio-isoform-switching

GPTomics

Analyzes isoform switching events and functional consequences using IsoformSwitchAnalyzeR. Predicts protein domain changes, NMD sensitivity, ORF alterations, and coding potential shifts between conditions. Use when investigating how splicing changes affect protein function.

代码生成 1194 6个月前
GPTomics

bio-clinical-databases-hla-typing

GPTomics

Call HLA alleles from NGS data using OptiType, HLA-HD, or arcasHLA for immunogenomics applications. Use when determining HLA genotype for transplant matching, neoantigen prediction, or pharmacogenomic screening.

数据处理 1194 6个月前