You've landed on a polished article about mortgage refinancing, but something feels wrong. The opening is unusually clean, the examples are perfectly tidy, and every paragraph uses the same measured tone. You search for an answer to “is this AI generated?” and quickly find a dozen scanners offering different verdicts.

That situation is now routine for editors, SEO teams, teachers, recruiters, and researchers. The mistake is treating detection as a courtroom verdict. A detector score is evidence, not proof. Reliable assessment comes from combining technical fingerprints, writing patterns, metadata, source quality, and tool outputs into a confidence-weighted conclusion.

Table of Contents

Why the AI or Human Question Is Harder Than It Sounds

A human writer can produce formulaic work, especially when following an SEO brief. An AI-assisted writer can also revise a draft until it contains personal anecdotes, uneven sentence lengths, and casual phrasing. The visible result may be neither purely human nor purely machine-generated.

A large 2025 study of 900,000 newly created web pages found that 74.2% contained AI-generated content, but only 2.5% were classified as pure AI. The same research classified 71.7% as a mix of AI and human writing, which means the practical question is often whether a page contains machine assistance, not whether a machine wrote every word. (Ahrefs' analysis of AI-generated content)

That distinction matters. A page may contain a human-written introduction, an AI-generated product section, and a final edit from a subject-matter expert. A binary label hides that composition.

A mind map graphic titled Why the AI or Human Question Is Harder Than It Sounds listing indicators.

Replace the verdict with a confidence note

Start with a better question: What evidence suggests AI involvement, how strong is that evidence, and what could produce the same signals?

Writing style is a weak signal by itself. Repeated transitions, neutral phrasing, and unusually smooth paragraphs may indicate model assistance, but they can also reflect a careful editor or an institutional style guide. Technical fingerprints are often more concrete, while source verification tells you whether the author actually understands the subject.

Detection also has measurable limitations. A 2024 meta-analysis reported overall detection accuracy of 55.54%, with text detection at 52.00%, image detection at 53.16%, and video deepfake detection at 57.31%. (The generative AI survey and detection meta-analysis) Those results are not a reason to abandon detection. They're a reason to stop pretending that one score settles authorship.

Practical rule: Use a detector to decide what deserves inspection, not to decide who is guilty.

The Technical Signals You Can Inspect Yourself

Before pasting text into a scanner, inspect the page as a web document. Technical evidence can reveal how a site was assembled, although it won't always prove who wrote the copy.

Right-click the page and choose View Source. Use the browser's search function for terms such as prompt, generated, GPT, OpenAI, completion, and token. Look through the <head> element for unusual generator metadata, placeholder author fields, or publication dates that appear to match a deployment event rather than an editorial workflow.

Read the build fingerprints

Inspect class names and data attributes. Names such as .ai-content, .generated-text, .auto-blog, or hidden containers labelled with prompt-related terms can provide useful context. None is conclusive, because developers can use descriptive names for ordinary automation, but several related markers deserve attention.

Then review script tags. A page that loads inference-related services, unusual third-party bundles, or generator-specific assets may indicate an AI-assisted publishing workflow. Check the surrounding code rather than counting a single script. Many websites use external APIs for search, analytics, chat, or personalization, and those integrations don't establish that the visible article was generated by AI.

Look at metadata next:

  • Author fields: Placeholder names, missing biographies, or generic “staff” labels weaken provenance.
  • Dates: A publication timestamp that aligns with a batch deployment may warrant review, especially if the site has many similar pages.
  • Canonical tags: Canonicals pointing to feeds, aggregators, or unexpected templates can expose an automated publishing structure.
  • Accessibility attributes: Empty labels and repetitive hidden elements may reveal a rushed or templated build.

For a deeper explanation of how page-level signals can be collected and interpreted, use the AI Website Detector methodology. Treat the output as an evidence layer, not an authorship certificate.

Screenshot from https://example.com/screenshots/view-source-ai-markers.png

Separate builder evidence from writing evidence

A site can be built with an AI-first website builder while its copy was written by a person. The reverse is also possible. A human-built site may publish AI-assisted articles through a normal CMS.

That distinction prevents a common error: concluding that a platform fingerprint proves the content itself was machine-generated. Technical inspection tells you about the production environment. You still need to examine the words, citations, author history, and editorial process before making a claim about authorship.

Writing Style and Metadata Red Flags

Textual clues become useful when you compare the suspected page with material that has a known history. Don't ask whether one paragraph “sounds like AI.” Ask whether the page shows a repeated pattern that differs from the author's established voice.

Look for uniform paragraph rhythm, predictable transitions, exhaustive lists, and conclusions that restate the introduction without adding judgment. Model-assisted writing often explains every side of an issue with similar weight, while a specialist usually prioritizes a position, excludes irrelevant options, or uses domain-specific examples.

Metadata can strengthen or weaken that impression. Check whether the author has a coherent biography, whether older posts use the same vocabulary, and whether the profile has evidence of a real publishing history. A missing byline isn't proof of automation, but it reduces the amount of provenance available for verification.

Compare patterns rather than hunting magic phrases

Common signals include repeated openers such as “Moreover,” “Furthermore,” and “It is worth noting.” They're not forbidden phrases, and human writers use them. The useful question is whether the same transitions appear across many paragraphs, alongside balanced claims, generic examples, and a lack of first-person experience.

Also examine sentence variety. If ten consecutive sentences begin with only a small set of recurring structures, record it as a weak cue. Don't turn it into a mathematical test. A formal report, legal explainer, or carefully edited technical page may naturally have low variation.

Signal Category AI-Generated Tells Human-Written Tells
Paragraph rhythm Similar length and cadence throughout Noticeable variation based on the idea
Transitions Repeated connectors and stock framing Transitions chosen for the specific argument
Examples Broad, tidy, and interchangeable Concrete details tied to experience or context
Tone Consistently neutral and heavily hedged Clear preferences, emphasis, or qualified conviction
Authorship Generic staff identity and thin biography Traceable expertise and consistent publishing history
Metadata Missing or conflicting dates and author fields Coherent dates, bylines, and editorial signals

Build a red-flag tally, but keep it qualitative. Five weak style cues don't outweigh a verified author draft, while one strong provenance contradiction may deserve more attention than a page full of polished prose.

A polished sentence is not suspicious by itself. A page that repeatedly avoids risk, experience, specificity, and accountability is more revealing.

Detection Tools Worth Using in 2026

No single tool sees the whole problem. General text detectors inspect language patterns, originality systems look for reuse, browser tools surface signals during research, and website-level systems inspect the page around the copy.

General AI text detectors are useful for sampling distinctive passages. GPTZero, OriginalityAI, and similar services can identify patterns associated with model-generated text and provide probability-style outputs. They struggle with short samples, translated material, technical writing, and text that a human has substantially revised.

Plagiarism and originality checkers answer a different question. They can find copied passages, close paraphrases, and reused material that a language detector may miss. A clean originality result doesn't prove human authorship, because newly generated text may not match an indexed source.

Browser extensions and inline research tools help when you're reviewing many pages. They reduce the friction of collecting suspect passages and recording page context, but they can encourage snap judgments. Use them to flag items for review, not to label an author while skimming.

The AI Website Detector stack approach combines HTML parsing, technical fingerprinting, metadata review, and pattern matching at page level. It can help identify whether a site appears to use an AI-first builder or an AI-augmented stack, while text detectors focus on selected passages. The two questions overlap, but they aren't identical.

Tool Category Best At Failure Modes
General text detector Sampling prose for machine-like patterns Short, edited, translated, or technical text
Originality checker Finding reuse and close source matches Newly generated material with no indexed match
Browser extension Flagging pages during live research Fast labels without enough context
Website-level analysis Connecting code, metadata, assets, and page patterns Builder evidence doesn't prove copy authorship

For teams designing a broader monitoring process, the practical principles in this anomaly detection systems guide are useful because they emphasize baselines, signal combination, and investigation instead of reacting to one unusual observation. You can also review this free AI detector workflow when you need a quick first pass.

Reading Scores, Confidence Notes, and Conflicting Verdicts

A probability score is a model's estimate that the submitted text resembles material associated with machine generation. It isn't a measurement of the author's identity, and it doesn't tell you how much AI was used.

Read the confidence note before the headline score. The note may indicate whether the sample contains enough distinctive material, whether the result is stable, or whether the tool saw only weak signals. A high score on a short footer is less useful than a moderate score on the page's main argument.

Interpret signal lists in context

Terms such as perplexity, burstiness, and uniform sentence length describe statistical properties of the text. They can help explain a result, but none is decisive. Technical documentation, compliance copy, translated writing, and highly edited prose can share those properties.

A practical weighting system looks like this:

  1. Give the greatest weight to evidence tied directly to the page's main body.
  2. Discount boilerplate such as navigation, disclaimers, cookie notices, and footers.
  3. Give more weight to a detector designed for the relevant language or content type.
  4. Treat agreement among independent tools as stronger than an isolated result.
  5. Record what changed when you re-scan a different passage.

If one tool reports a low likelihood and another reports a high likelihood, don't average the figures and call the result precise. Compare their sampled text, model assumptions, confidence notes, and known blind spots. A conflict often means the page contains mixed authorship or that both tools are reacting to a shared feature, such as repetitive technical language.

A five step infographic explaining how to read and interpret AI detection scores and confidence notes.

A recent large-scale workflow from Pew illustrates why sampling decisions matter. Researchers sampled 490,000 pages, using 10,000 English-language pages from each of 49 Common Crawl snapshots, and treated a detector score of 0.2 or higher as evidence of AI authorship or editing. The study also found that only about 10% to 15% of sampled pages exposed a publication-date field in HTML, which makes date-based attribution difficult. (Pew Research Center's AI-content methodology)

For another page-level perspective, see the website builder checker. Scores inform your conclusion. They don't make it.

Your Follow Up Verification Workflow

A scan should trigger a process, not an accusation. Start by deciding how much verification the situation deserves. A hobby blog, a supplier portfolio, a medical explanation, and a student submission carry different consequences, so the same ambiguous signal shouldn't produce the same response.

Establish the stakes first

For high-impact content, inspect the page manually before communicating a conclusion. Check the author bio, follow important claims to primary sources, and compare the suspected page with older material. Read enough to determine whether the author has a consistent vocabulary, relevant experience, and an identifiable editorial history.

Next, inspect the page source and run a second detector that uses a different approach from the first. Compare the passages tested, the signal descriptions, and the areas where both tools agree. Don't submit only the opening paragraph. Test the main explanatory sections, because introductions and conclusions often contain generic language that can produce noisy results.

Keep an investigation log

Record the URL, date of review, tools used, passages sampled, scores, confidence notes, and manual observations. Add the reasons for your conclusion, including evidence that weakened the AI hypothesis. This log protects against memory bias and lets another reviewer reproduce the investigation.

Use a simple decision rule:

  • Converging evidence: Independent tool outputs and manual inspection point in the same direction. State the conclusion with an appropriate confidence level.
  • Mixed evidence: Technical signals suggest automation, but authorship history and writing samples look consistent. Describe the page as ambiguous or potentially AI-assisted.
  • Weak evidence: The result rests mainly on tone, formatting, or a single detector score. Don't make an authorship claim.
  • High stakes and conflict: Escalate to source verification, author review, or an editorial conversation rather than relying on automation.

A five-step flowchart illustrating a follow-up verification workflow for assessing information credibility and AI-generated content.

The most important safeguard is proportionality. Detection systems can produce false positives on repetitive boilerplate, code, configuration files, and other low-diversity content. Recent research specifically highlights this risk for structured text and warns that watermarking or detector signals can be altered by translation, paraphrasing, light editing, or mixed-tool workflows. (Research on false positives in low-entropy content)

Putting It All Together

A reliable answer to “is this AI generated?” comes from three evidence layers.

First, inspect the technical layer. Review source code, metadata, scripts, assets, headers, and publishing structure. This can reveal an AI-first builder, an automated content pipeline, or a templated deployment pattern.

Second, evaluate the editorial layer. Compare the page's voice with older work, assess the specificity of its examples, verify its claims, and look for meaningful author accountability. Treat repetitive style as a prompt for investigation, not as proof.

Third, interpret the tool layer. Run a website-level scan where the question concerns the site's construction, and run text analysis where the question concerns selected copy. Preserve the sampled passages and confidence notes. A low or middling score should remain inconclusive when the surrounding evidence doesn't support it.

The strongest conclusion is often nuanced: “The site shows technical evidence of AI-assisted construction, while the article's authorship remains unverified.” That statement is more useful than “AI generated” because it identifies what you really know.

A fast verification checklist

  • Inspect the source: Search HTML, metadata, scripts, and classes for relevant markers.
  • Review the site: Identify the builder, stack, publishing structure, and unusual deployment signals.
  • Test the copy: Scan representative body passages with a suitable text detector.
  • Record the reasoning: Save the URL, samples, outputs, and manual observations.

Detection is a verification habit. Build a fingerprint, compare evidence, investigate conflicts, and write the conclusion that the evidence supports, no stronger.


AI Website Detector analyzes a site's technical signals, including its HTML, scripts, metadata, assets, and builder fingerprints, then returns an AI probability score, verdict label, and explanation. Visit AI Website Detector to scan a URL and add page-level evidence to your authorship and technology verification workflow.