VTSS™ · v2.0 · Published

Verified Layer Trust
Standard v2.0

An AI-driven, evidence-based methodology for ranking software vendors. Quality and Presence are scored separately, and never blended into one number.

53
structured prompts
3
AI platforms
159
evaluations / vendor / month
8
scored dimensions
01

Overview and Purpose

The VerifiedLayer Trust Score is a structured, automated, evidence-based assessment of any software vendor operating in the B2B technology market. It is produced by running 53 structured prompts across 3 major AI platforms, combining community sentiment signals, and verifying product evidence, all without requiring vendor participation or submission.

The methodology is designed to be transparent, reproducible, and useful to software buyers making real purchasing decisions. Every component of the score is documented on this page. Vendors who understand the methodology can identify exactly what public evidence would improve their score; that is the intended behavior.

VTSS measures how leading AI models perceive software vendors based on publicly available evidence. It does not certify legal compliance, financial health, or factual truth.
DOCUMENT REFERENCE
VL-VTSS-V2 · Published Methodology
TOTAL EVALUATIONS
159 / vendor / month (53 prompts × 3 platforms)
SCORE RANGE
0–100, each axis
AXES PUBLISHED
Quality & Presence: never blended
UPDATE CYCLE
Monthly automated pipeline
HUMAN INVOLVEMENT
None in scoring
02

Guiding Principles

Evidence precedes confidence, and confidence precedes score: never the reverse, on any single query.
AI does not invent ratings. Every finding a model claims is checked against a predefined Evidence Rating Dictionary wherever a matching entry exists.
Three independent model families evaluate every vendor. Agreement is measured, not assumed.
Quality (how good) and Presence (how big/visible) are never blended into a single number.
Every vendor is evaluated on the same monthly cycle, with the same prompts and formulas.
Every published score is explainable and traceable back to the evidence that produced it.
03

Architecture at a Glance

Every vendor is evaluated across three independent signal-source layers. Weights are never fixed globally: they vary per dimension (Section 8).

S1

AI Platform Signal

Prompts probing what leading AI models know from training and public sources
40 prompts × 3 platforms
S2

Community Voice Signal

Prompts on Reddit, LinkedIn, Glassdoor, G2, TrustRadius sentiment
8 prompts × 3 platforms
S3

Product Evidence Signal

Prompts verifying certifications, analyst recognition, customer references, financial stability
5 prompts × 3 platforms
Total53 prompts × 3 platforms = 159 evaluations / vendor / month
Three independent platforms
Three separate foundation-model companies, each with an independent training pipeline, so agreement between them means something real.
Two axes, never blended
Every piece of evidence is tagged as measuring Quality, Presence, or Both, and the two resulting scores are calculated and published separately.
A third, supplementary axis
Momentum tracks quarter-over-quarter improvement, independent of absolute size.
04

How Evidence Is Collected

The fundamental problem with asking an AI model to "score this vendor 0–100" is that the model tends to assign a score first and generate justification for it afterward, producing overconfident, inconsistent results. VerifiedLayer inverts this order on every single query, without exception:

STEP 1

Collect Evidence

Collect evidence from public sources before considering any score. Sources are named and findings are quoted specifically. If fewer than 2 independent sources are found, the score for that query cannot exceed 45, regardless of what the evidence otherwise suggests.
STEP 2

Declare Confidence

Declare confidence (HIGH, MEDIUM, or LOW) based on what was found, before writing any score. HIGH requires 3+ independent sources, recent (within 24 months), of different types. MEDIUM requires 2 independent sources, or older sources. LOW means only 1 source, or indirect or contradictory evidence.
STEP 3

Assign the Score

Assign the score, constrained by the evidence and confidence already declared. A LOW-confidence query is still scored and published, but flagged.

EVERY QUERY RETURNS A CONSISTENT SET OF FIELDS

EVIDENCE_FOUNDsources named, findings quoted specifically
EVIDENCE_GAPSwhat was searched for but not found
CONFIDENCEHIGH / MEDIUM / LOW + reasoning
SCOREinteger 0–100 + reasoning
AXISQuality / Presence / Both

APPROVED PUBLIC SOURCES

Vendor sites & product documentationCertification bodies & regulatory registriesReview platformsGitHubNews publicationsCompany blogsPartner ecosystems
Vendor marketing copy is never treated as independent evidence.
05

Confidence Calibration & Score Ceilings

Confidence is not a courtesy label; it actively constrains what a score is allowed to be:

ConfidenceRequirement
HIGH3+ independent sources, recent (within 24 months), of different types
MEDIUM2 independent sources, or sources older than 12 months
LOWOnly 1 source, indirect evidence, or contradictions found
A vendor with fewer than 2 independent public sources on a given question cannot score above 45 on that question, no matter how positive the single available source is. This prevents a single self-reported claim from carrying the same weight as corroborated public evidence.
06

Evidence Rating Dictionary (ERD)

Common findings map to a predefined rating so that the same evidence is treated the same way for every vendor, every month:

FindingRating
SOC 2 Type II Certified (verified via audit registry)+95
SOC 2 Type I only, or self-attested+10 (maximum)
ISO 27001 Certified+90
Named Gartner MQ / Forrester Wave position, with year and report cited+85
Generic "recognized by Gartner," no specifics0
GDPR Compliance documented+40
Named enterprise reference with quantified outcome+30
Anonymous case study / generic testimonial0
Multi-Factor Authentication available+20
Public API documentation+15
Pricing fully published with tiers and inclusions+15
Contact-sales-only pricing0 (neutral)
"AI-powered" claimed in marketing with no technical substancecapped at 40
Active Critical Security Incident (unresolved)−50
Active Regulatory Action−70
No Glassdoor presence45 (unknown, not treated as bad)
A model does not decide what a certification or an incident is "worth" in the moment; it recognizes the finding, and this table supplies the anchor. This is also what makes a vendor's score improvable in a predictable way: an organization that genuinely earns SOC 2 Type II, publishes transparent pricing, or documents a named enterprise outcome knows in advance what that will do to its score.
07

Cross-Model Aggregation

Each of the 3 platforms answers every query independently. Their scores are combined using a reliability-weighted method, so agreement is rewarded with a simple average and disagreement is handled conservatively rather than blindly averaged away.

Platform AgreementAggregation Method Used
High agreementSimple average across all 3 platforms
Good agreementMedian (the middle score of the 3)
Moderate agreementTrimmed mean (the one outlier score is dropped, remaining 2 averaged)
Low agreementConservative floor (lowest of the 3 scores used), and the query is flagged for review
This produces one combined score per query, which feeds into its dimension along the axis (Quality, Presence, or Both) that evidence was tagged with.
08

The 8 Scored Dimensions & Signal-Source Weights

Every vendor is scored across 8 fixed dimensions. Weights are derived from research on what actually drives enterprise purchase decisions, and sum to exactly 100%.

D1

Product Quality

18%

Feature completeness, stability, real-world effectiveness.

ANTI-GAMINGRequires actively investigating weaknesses; scores above 70 require documented negative-source review.
D2

Market Reputation

16%

Analyst, media, and peer regard.

ANTI-GAMINGRequires a specific report name, year, and position; generic mentions are rejected.
D3

Community Sentiment

15%

What real practitioners say on Reddit and forums.

ANTI-GAMINGAbsence of signal scores 40, not 100; silence cannot be gamed into a good score.
D4

Customer Experience

14%

Onboarding, support, post-sale reality vs. promise.

ANTI-GAMINGOnly user-reported timelines count, not vendor-stated ones.
D5

Technical Credibility

13%

Certifications, API quality, compliance.

ANTI-GAMINGSOC 2 Type II earns full credit; Type I is capped; self-attestation earns nothing.
D6

Value for Money

11%

Pricing transparency, documented ROI, true total cost of ownership.

ANTI-GAMINGMarketing ROI claims with no client attribution score 0.
D7

Innovation & Vision

8%

Genuine innovation vs. marketing claims.

ANTI-GAMING"AI-powered" with no technical substance caps the score at 40.
D8

Buying Safety

5%

Financial stability, lock-in risk, legal history.

ANTI-GAMINGLock-in difficulty must be documented by users who tried to exit, not by the vendor.

SIGNAL-SOURCE WEIGHTS PER DIMENSION

DimensionAI SignalCommunity SignalEvidence SignalWhy
Product Quality40%40%20%Equally about perception and community voice
Market Reputation35%35%30%Blends AI knowledge, community, and verifiable analyst position
Community Sentiment20%70%10%Community is the primary signal
Customer Experience35%45%20%Post-sale reality is best captured in community voice
Technical Credibility30%30%40%Certifications are verifiable facts: evidence dominates
Value for Money35%35%30%Blends perception, community feedback, and pricing evidence
Innovation & Vision40%30%30%Best assessed through AI awareness and analyst citation
Buying Safety30%30%40%Depends on verifiable facts: funding, legal, certifications
09

Quality Score & Presence Score

Every finding is tagged as measuring Quality, Presence, or Both. A dimension therefore produces two sub-scores (a Quality-axis score and a Presence-axis score) built only from the evidence tagged to that axis. Evidence tagged Both contributes at half weight to each axis, so it is never counted at full strength twice.

FORMULAS
Quality Score = Σ (Dimension Weight × Dimension Quality-axis Score) across all 8 dimensions
Presence Score = Σ (Dimension Weight × Dimension Presence-axis Score) across all 8 dimensions
Both scores range from 0 to 100 and are published side by side, never combined into a single blended number. This is what prevents large, high-visibility vendors from automatically outranking better, smaller products: a vendor with excellent Quality and modest Presence is shown exactly as such, not averaged down.
10

Confidence Score

This measures the reliability of the evidence base behind a vendor's score this month, not the vendor's quality.

FactorDescriptionWeight
AI AgreementAverage cross-model agreement across all queries40%
Evidence CoverageVerified items collected vs. expected baseline for the category20%
Source DiversityIndependent source categories represented20%
Evidence FreshnessAverage recency across all findings used20%
11

Momentum Score

Momentum rewards vendors improving quickly, regardless of where they started:

Momentum = 70% × (change in Quality Score over the last 2 quarters) + 30% × (growth in reviews, mentions, and references over the same period)
A vendor whose Momentum reaches +15 or more is shown with a "Rising" flag, regardless of how small it still is.
12

Automated Quality Gates

Four checks run automatically before any score is published. No human is involved in scoring (the pipeline always publishes, it never holds a score indefinitely) but low-confidence results are flagged directly on the public profile.

1

Cross-Model Reliability

Determines which aggregation method is used for each query (High / Good / Moderate / Low agreement).
2

Evidence Spot-Check

A sample of cited evidence is validated against a live web search. A high fabrication rate floors the affected score and flags it as low confidence; a moderate rate triggers a single re-run with a stricter check; a low rate passes cleanly.
3

Month-over-Month Change Gate

If a vendor's score changes sharply from the prior month, the published score is blended with the prior month's result rather than jumping immediately, and a "Significant movement detected" flag is shown. Prevents a single event (a viral post, a review bomb) from swinging a score overnight.
4

Confidence Gate

If a large share of a vendor's monthly queries return low confidence, the published score is capped and flagged "Limited Evidence" or "Partial Evidence." If the evidence base is too thin to trust at all, the leaderboard shows "Unranked: Insufficient Evidence" instead of a numeric score.
13

Leaderboard, Badges & Category Rank

Badge: a grid of Quality × Presence, never a single blended score.

Quality ScorePresence ScoreBadge
≥ 75≥ 65👑 LEADER
≥ 65< 65🚀 HIGH PERFORMER
50–74≥ 65🏢 ESTABLISHED
50–64< 65🛡 VERIFIED
< 50any⚠️ EMERGING / CAUTION
(any of the above)Momentum ≥ +15+ "RISING" flag added
A vendor with high Quality but modest Presence earns High Performer, not a diluted mid-tier badge. A vendor with high Presence but mediocre Quality earns Established, not an inflated Leader badge.

CATEGORY RANK FORMULA

Rank Score = 70% × Confidence Score + 20% × Review Volume Score + 10% × Category Relevance
Category Rank orders vendors within a category page only; it does not affect a vendor's Quality, Presence, or badge.

CATEGORY ASSIGNMENT PRIORITY

PrioritySourceWhy
1: HighestG2 Category PageResearch-backed, reflects real buyer evaluation behavior.
2: HighGartner / Forrester ReportIndependent analyst classification, names category directly.
3: MediumTechnology PressJournalists generally research before publishing.
4: LowestVendor's Own WebsiteVendors self-describe aspirationally; used only if 1–3 yield nothing.
Vendors matching no approved category are placed in "Emerging AI Software" and may correct their category assignment by claiming their profile.
14

Monthly Publication Cycle

1
Trigger the monthly evaluation for every vendor
2
Execute all 53 standardized prompts across all 3 platforms
3
Apply the evidence-first protocol to every query
4
Aggregate cross-model scores; compute dimension, Quality, and Presence scores
5
Compute Confidence and Momentum scores
6
Run all automated quality gates and apply any caps
7
Assign the badge and Category Rank
8
Publish the leaderboard, each month archived as an immutable historical snapshot
15

Limitations & Disclaimer

VTSS measures AI-perceived, evidence-based reputation, not absolute truth. Rankings:

Reflect only publicly available information as of the evaluation date.
Depend on the current Evidence Rating Dictionary; ratings can be revised going forward but are never applied retroactively to past monthly snapshots.
Are not legal, financial, security, or compliance certifications, and should never be treated as such.
Can be temporarily capped or withheld when evidence is thin; this reflects the confidence of the evaluation, not a judgment on the vendor.
Should be read as a structured, monthly AI-driven assessment, not a guarantee of vendor performance or a substitute for a buyer's own due diligence.
16

Glossary

Query

One of the 53 standardized monthly prompts, asked of all 3 platforms.

Signal Source

The AI Platform, Community Voice, or Product Evidence category a query belongs to.

Evidence Rating Dictionary (ERD)

The fixed, versioned table mapping findings to predefined ratings.

Confidence (query-level)

The HIGH/MEDIUM/LOW label a model assigns a single query, based on evidence found.

Confidence Score

The separate, aggregate measure of how reliable the whole month's evidence base was.

Axis (Quality / Presence / Both)

Which of the two never-blended scores a piece of evidence contributes to.

Quality Score / Presence Score

The two core published scores; never combined into one number.

Momentum Score

Rate of quarter-over-quarter improvement.

Category Rank Score

The separate formula used only to order vendors within a category page.

Badge

The label (Leader, High Performer, Established, Verified, Emerging/Caution, + Rising) from the Quality × Presence grid.
17

Worked Example

VENDOR
"Nimbus Data"

QUALITY-AXIS DIMENSION SCORES

DimensionWeightQuality Score
Product Quality18%88.00
Market Reputation16%82.00
Community Sentiment15%79.00
Customer Experience14%84.00
Technical Credibility13%87.00
Value for Money11%75.00
Innovation & Vision8%80.00
Buying Safety5%90.00

PRESENCE-AXIS DIMENSION SCORES

DimensionWeightPresence Score
Product Quality18%70.00
Market Reputation16%85.00
Community Sentiment15%90.00
Customer Experience14%60.00
Technical Credibility13%65.00
Value for Money11%55.00
Innovation & Vision8%72.00
Buying Safety5%68.00
Quality Score = 0.18(88)+0.16(82)+0.15(79)+0.14(84)+0.13(87)+0.11(75)+0.08(80)+0.05(90) = 83.03
Presence Score = 0.18(70)+0.16(85)+0.15(90)+0.14(60)+0.13(65)+0.11(55)+0.08(72)+0.05(68) = 71.76
CONFIDENCE SCORE
88.60
well above the evidence floor, published without a cap
MOMENTUM
11.72
below the +15 threshold, no "Rising" flag this month
BADGE: QUALITY 83.03 (≥75) AND PRESENCE 71.76 (≥65) →
👑 LEADER
Published profile: Quality 83.03 · Presence 71.76 · Confidence 88.60 · Momentum 11.72 · Badge: LEADER

See your score in action

Browse the leaderboard to see how vendors perform under VTSS v2.0, or run your own vendor audit.