Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
-
Updated
Jul 23, 2026 - Rust
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Use LLMs to robustly extract web data
Fully automated and hands-free, accurately extracting and understanding web content — powered by machine learning agents.
Windows desktop app for collecting, reviewing, ranking, and exporting Google Scholar search results.
Low-Cost Cross-Domain Web Structured Information Extraction using specialized LoRA adapters.
Replayable Browser Agent
基于Scala Akka的分布式主题网络爬虫
A swipeable shortlist of live San Francisco apartment listings, powered by Context.dev.
Self-hosted web scraping and Markdown extraction for AI agents
Automatic extraction of the information on local event from a webpage with Machine Learning
Free local web search/extraction router for AI agents. Go CLI + MCP, BYOK/free-first routing, keyless DDGS/Scrapling fallback, setup writers and client guides.
A powerful and lightweight web scraping library with LLM extraction capabilities. This library combines web scraping with AI-powered content extraction using either OpenAI or OpenRouter APIs.
Hermes plugin to improve web search with SearXNG for higher-quality search results and Crawl4AI for LLM-optimized webpage extraction.
Fast local Tavily-compatible web search and extraction backend for Hermes
Predicting product recommendation score using the data available on the website of the client
Programming assignments for Web Information Extraction and Retrieval, FRI UL, 2021. PA1: standalone webcrawler of .gov.si web sites, PA2: approaches of the structured web data extraction, PA3: Data processing and indexing and Data retrieval.
Standalone Crawl4AI web extraction and bounded crawling plugin for Hermes Agent
Glasses Web Reader for Even Realities G2 — three-layer browser (sources → articles → reader) using r.jina.ai for clean URL extraction.
Validate public URLs and turn them into reviewable AI-ready Markdown or JSON.
Add a description, image, and links to the web-extraction topic page so that developers can more easily learn about it.
To associate your repository with the web-extraction topic, visit your repo's landing page and select "manage topics."