Document Network
40466 documents positioned by embedding similarity — documents close together share similar content
Color by:
40466 documents plotted by semantic similarity. Documents are grouped into 8 clusters.
About this visualization
Each point represents a document from the corpus. Documents are positioned using Principal Component Analysis (PCA) on 768-dimension embeddings generated by Google EmbeddingGemma. Documents with similar semantic content appear closer together.
Clusters are computed via K-Means (k=8) on the full embedding space. Cluster boundaries in 2D are approximate — the true cluster separation exists in the original 768 dimensions.
- Color by Cluster — Groups documents by semantic similarity. Hover over points to see document IDs.
- Color by Source — Shows documents colored by their origin dataset (FOIA, DOJ, etc.).