name: data-spider description: Scrape any webpage and extract structured data as JSON, table, or list. Supports schema-guided extraction. acceptLicenseTerms: true metadata: clawdbot: emoji: "🕷️" homepage: https://aiprox.dev requires: env: - AIPROX_SPEND_TOKEN
Scrape and extract structured data from any webpage. Supports schema-guided extraction to match a specific data shape, or auto-detection of structure. Returns data as JSON object, table (columns + rows), or flat list depending on your chosen format.
urlschema object — data will be extracted to match that exact shapeformat: json (default), table, or list| Permission | Scope | Reason |
|---|---|---|
| Network | aiprox.dev | API calls to orchestration endpoint |
| Env Read | AIPROX_SPEND_TOKEN | Authentication for paid API |
curl -X POST https://aiprox.dev/api/orchestrate \
-H "Content-Type: application/json" \
-H "X-Spend-Token: $AIPROX_SPEND_TOKEN" \
-d '{
"url": "https://example.com/pricing",
"schema": {"free_tier": null, "pro_price": null, "enterprise": null},
"format": "json"
}'
{
"data": {"free_tier": "$0/month, 1000 API calls", "pro_price": "$29/month", "enterprise": "custom pricing"},
"summary": "SaaS pricing page with three tiers.",
"source": "https://example.com/pricing",
"format": "json"
}
curl -X POST https://aiprox.dev/api/orchestrate \
-H "Content-Type: application/json" \
-H "X-Spend-Token: $AIPROX_SPEND_TOKEN" \
-d '{
"task": "extract pricing tiers as a table",
"url": "https://example.com/pricing",
"format": "table"
}'
{
"columns": ["Plan", "Price", "API Calls"],
"rows": [
["Free", "$0/month", "1,000"],
["Pro", "$29/month", "50,000"],
["Enterprise", "Custom", "Unlimited"]
],
"summary": "Three-tier SaaS pricing.",
"source": "https://example.com/pricing",
"format": "table"
}
{
"items": ["$0/month — Free tier, 1000 API calls", "$29/month — Pro, 50,000 calls", "Enterprise — custom pricing"],
"summary": "SaaS pricing tiers extracted as flat list.",
"source": "https://example.com/pricing",
"format": "list"
}
Data Spider fetches and analyzes webpage contents via URL. Content is processed transiently and not stored. Analysis is performed by Claude via LightningProx. Respects robots.txt and rate limits. Your spend token is used for payment only.
这个 Skill 质量中等偏上,文档清晰易懂,示例丰富,配置完整。但文件较少,缺少实际使用案例和故障处理说明,对外部 API 依赖较强。建议补充更多功能示例和常见问题解答。整体上适合有经验的开发者快速集成,对新手不太友好。