The web-content chunk-quality problem
Across Pinecone, Chroma, Weaviate, Qdrant, and pgvector, the public benchmarks all measure retrieval algorithm quality and largely ignore input quality. In practice, input quality is the single largest source of variance in retrieval accuracy. Embed a corpus of web articles as raw HTML text on the same vector DB as the same articles converted to Markdown, and top-K accuracy differs by 15-30 points.
Chunking strategy for web Markdown
Split first by H2 (article sections), then sub-split by H3 if your H2 sections are long. For prose-heavy articles without strong subheadings, fall back to a recursive character splitter at 600-1000 tokens. Always store the source URL and heading path as chunk metadata: it powers source attribution in synthesis and per-source filtering at query time.