<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>PDF Extraction on Top AI Skills</title><link>https://topaiskills.com/tags/pdf-extraction/</link><description>Recent content in PDF Extraction on Top AI Skills</description><generator>Hugo</generator><language>en</language><lastBuildDate>Thu, 09 Jul 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://topaiskills.com/tags/pdf-extraction/index.xml" rel="self" type="application/rss+xml"/><item><title>Kreuzberg (Xberg): Document Intelligence &amp; PDF Extraction</title><link>https://topaiskills.com/skills/general/kreuzberg/</link><pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate><guid>https://topaiskills.com/skills/general/kreuzberg/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Kreuzberg (now part of the Xberg ecosystem) is a polyglot document intelligence framework built with a Rust core. It extracts text, metadata, images, and structured information from over 97 file formats — PDFs, Office documents, images, EPUB, HTML, and more. If you&amp;rsquo;ve ever fought with a PDF library that couldn&amp;rsquo;t handle a simple table or mangled your Chinese characters, Kreuzberg is the relief you&amp;rsquo;ve been looking for.&lt;/p&gt;
&lt;p&gt;The project started as a standalone skill for AI coding agents and grew into a full framework with bindings for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, and TypeScript (Node/Bun/Wasm/Deno). You can also use it via CLI, REST API, or MCP server. The Rust core means it&amp;rsquo;s fast — we&amp;rsquo;re talking milliseconds per page on modern hardware.&lt;/p&gt;</description></item></channel></rss>