<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Crawl4ai - 标签 | 飞污熊小站</title><link>https://xiongneng.me/tags/crawl4ai/</link><description>飞污熊小站</description><generator>Hugo 0.165.0 &amp; FixIt v0.4.6-20260512073637-464c4659</generator><language>zh-CN</language><managingEditor>yidao620@163.com (XiongNeng)</managingEditor><webMaster>yidao620@163.com (XiongNeng)</webMaster><copyright>XiongNeng</copyright><lastBuildDate>Mon, 31 Aug 2026 12:53:38 +0000</lastBuildDate><atom:link href="https://xiongneng.me/tags/crawl4ai/index.xml" rel="self" type="application/rss+xml"/><item><title>crawl4ai专为LLM而生的爬虫</title><link>https://xiongneng.me/posts/ai/crawl4ai/</link><pubDate>Sat, 29 Aug 2026 20:35:10 +0800</pubDate><author>yidao620@163.com (XiongNeng)</author><guid>https://xiongneng.me/posts/ai/crawl4ai/</guid><category domain="https://xiongneng.me/categories/ai/">AI</category><description>&lt;p&gt;做 AI 应用的人，迟早都会撞上同一个问题&amp;ndash;喂给模型的数据从哪来？&#10;模型自己不会上网。你要给它知识，就得先去网上抓内容、清洗、转成它能吃的格式。而网页抓取这件事，从来都不简单：动态加载、反爬、登录态、表格、分页、JS 渲染……每个都是坑。&#10;有人专门为 LLM 场景做了一个爬虫，把抓取到的网页直接转成干净的 Markdown。这个项目叫 &lt;code&gt;crawl4ai&lt;/code&gt;，是 GitHub 上 star 最多的爬虫，作者是 unclecode。&lt;/p&gt;</description></item></channel></rss>