Skip to main content
Identrics
API Framework

Document Similarity & ClusteringSimilarity API by Identrics

Find near-duplicates, semantic similarity and topical clusters across archives.

Talk to us about access
Use it standalone, or compose it with any other Kaspian enrichment - sentiment, NER, topics, summarisation, and more - in a single pipeline.
Kaspian API
What the API Does

A single API endpoint, production-ready.

The Similarity API estimates how close two documents are in meaning - not just in shared keywords - so you can detect near-duplicates, semantic overlap and plagiarism across very large archives.

Beyond pairwise similarity, the API can cluster a whole corpus into coherent groups: trending narratives, repeated coverage, related complaints or topical themes. Useful any time the question is 'what's the same here?'

Common use cases:

  • 01Detect plagiarism and reused content across publishing archives.
  • 02Avoid duplicate coverage by surfacing semantically related stories.
  • 03Cluster legal, regulatory and compliance documents by topic.
  • 04Group customer feedback into recurring themes for faster triage.
API Capabilities

What you get out of the Similarity endpoint.

Built and maintained by the Identrics applied AI team. Trained on human-in-the-loop data, benchmarked continuously, and shipped to you as a single, well-documented API.

01

Semantic similarity

Score similarity based on meaning, not just shared keywords - so rewrites and translations are recognised too.

02

Multilingual matching

Compare documents written in different languages and still surface the matches that share intent.

03

Clustering at scale

Group thousands of documents into coherent topic clusters in a single API call.

04

Custom thresholds

Tune similarity thresholds to your workflow - duplicate detection, related content or loose topical grouping.

How the API Works

From request to enriched response.

  1. 01

    Send your documents

    POST a pair of documents for similarity scoring, or a full batch for clustering.

  2. 02

    Models embed & compare

    Documents are embedded into a shared semantic space and compared or clustered.

  3. 03

    Receive structured output

    Get back similarity scores, near-duplicate flags and topical clusters ready for downstream use.

Where It Fits

Industries already running this API in production.

Risk & Compliance

Filter regulatory content and internal documentation by topic, identify near-duplicate filings and keep risk reviews focused on what's actually new.

Publishers

Avoid duplicate coverage, detect plagiarism and surface related stories across vast archives - so editorial teams can focus on original reporting.

Similarity API

Ready to plug Similarity into your stack? Tell us the job; we'll spin up a standalone or composed deployment.