Synthetic Entity Detection: What Manufactured Authority Signals Look Like From the Inside
AI systems are increasingly targeted by manufactured authority signals: synthetic entities, coordinated sameAs networks, schema over-claiming, and citation farming. SiteNexis detects these patterns — not to accuse, but to help you identify and correct signals that look synthetic even when they are not.
Synthetic entity detection is defensive intelligence. We built it not to flag competitors but to help domain owners understand how their own authority signals appear to AI systems — because inadvertent synthetic-looking patterns are more common than deliberate manipulation. A company that launched multiple regional versions of the same LinkedIn profile, a blogger with identical author bios across 12 syndication partners, a review page with aggregate ratings imported from a third-party platform with no on-page review content: none of these are deliberate fraud, but all produce synthetic detection signals.
The Five Pattern Categories
- Fake Entity Patterns — entities with no verifiable external presence, word-for-word attributes across unrelated domains
- AI Authority Networks — unusual sameAs clustering, multiple entities sharing creation-date and description-length patterns
- Schema Manipulation — aggregate ratings without review content, datePublished pre-dating content metadata, mismatched schema types
- Citation Farming — pages that cite only internal sources for factual claims, sudden high-density fact insertion between audits
- Unnatural Entity Clustering — entity graphs with density far above expected for domain type, single-hub topology
Confidence Scores, Not Binary Flags
Every detected pattern carries a confidence score from 0 to 1. We never report binary "is fake" verdicts. A schema manipulation pattern detected with 0.91 confidence and two specific evidence signals is very different from one detected with 0.34 confidence and one weak indicator. The confidence score represents the strength of the evidence, not a judgment about intent. Users see the exact evidence that triggered each pattern, enabling them to investigate and correct the specific signal.
●Synthetic detection results are shown only to the domain owner — never in competitive analysis views. A medium-confidence pattern on your own domain is an actionable finding. The same finding exposed about a competitor, without their context, creates misleading impressions. This is an architectural constraint, not a UI choice.
Inadvertent Synthetic Signals
The most common inadvertent patterns: aggregate rating schema added by a plugin before any reviews exist, author schema generated for contributors who are not mentioned in the article body, sameAs links pointing to profiles that have since been deactivated, and entity descriptions that are copy-pasted across multiple pages. Each of these creates a detectable pattern that reduces the Entity Authenticity Confidence score. Fixing them is straightforward once they are identified.
The Entity Authenticity Confidence Score
The Synthetic Entity Risk Score (0–100, higher is riskier) is the composite of all five pattern categories weighted: Fake Entity Patterns (25%), Authority Networks (25%), Schema Manipulation (20%), Citation Farming (15%), Unnatural Clustering (15%). Entity Authenticity Confidence is the inverse: 100 minus the risk score. A domain scoring 82 on Authenticity Confidence has very few detectable synthetic signals. The score feeds into Machine Trust Score as a minor contributing factor — entity authenticity is one of several trust signals, not the dominant one.