Semantic similarity
Score similarity based on meaning, not just shared keywords - so rewrites and translations are recognised too.
Find near-duplicates, semantic similarity and topical clusters across archives.
Talk to us about accessThe Similarity API estimates how close two documents are in meaning - not just in shared keywords - so you can detect near-duplicates, semantic overlap and plagiarism across very large archives.
Beyond pairwise similarity, the API can cluster a whole corpus into coherent groups: trending narratives, repeated coverage, related complaints or topical themes. Useful any time the question is 'what's the same here?'
Common use cases:
Built and maintained by the Identrics applied AI team. Trained on human-in-the-loop data, benchmarked continuously, and shipped to you as a single, well-documented API.
Score similarity based on meaning, not just shared keywords - so rewrites and translations are recognised too.
Compare documents written in different languages and still surface the matches that share intent.
Group thousands of documents into coherent topic clusters in a single API call.
Tune similarity thresholds to your workflow - duplicate detection, related content or loose topical grouping.
POST a pair of documents for similarity scoring, or a full batch for clustering.
Documents are embedded into a shared semantic space and compared or clustered.
Get back similarity scores, near-duplicate flags and topical clusters ready for downstream use.
Filter regulatory content and internal documentation by topic, identify near-duplicate filings and keep risk reviews focused on what's actually new.
Avoid duplicate coverage, detect plagiarism and surface related stories across vast archives - so editorial teams can focus on original reporting.
Every Kaspian enrichment ships as its own API, and they all share a common contract - so you can call Similarity on its own today, or chain it with any of the others tomorrow to build a richer pipeline. No extra integration work.
Detect people, organisations, brands and locations across multilingual text.
Explore APIQuantify how audiences feel about brands, products and topics - at scale.
Explore APISort documents into your taxonomy with high-accuracy multilingual classifiers.
Explore APISurface emerging topics from unstructured text without predefined labels.
Explore APIMeasure attention and engagement around the topics that matter to you.
Explore APITrack every mention of your brand and competitors with contextual precision.
Explore APICompress long-form text into fact-checked, readable abstractive summaries.
Explore API