The AI Cowboys

Third-Party Validation: Proving a Product Does What It Claims

Vendors measure their own products, and buyers cannot tell a real result from a well-lit one. Here is what independent third-party validation is, what it measures for AI hardware and software, and the structural conditions that make a validator credible.

A vendor arrives with a datasheet. The model is 97% accurate. The device runs the workload at 12 watts. The agent resolves 80% of tickets without a human. Every number is plausible, every number was measured by the company that needs it to be true, and no buyer on the other side of the table can tell the difference between a real result and a well-lit one.

That gap is what third-party validation closes. Increasingly it is what companies come to us for: not to build the system, but to independently test the one they already built and put our name on what it actually does.

What third-party validation means

Third-party validation is an independent technical assessment of a product's performance claims, run by an organization that did not build the product and does not benefit from the result. The validator writes a test protocol, measures the product against it under stated conditions, and publishes what was found, including the conditions under which the claim fails.

The word "third" is doing real work. There are three parties to any claim:

  • First-party testing is the vendor testing its own product. It is necessary, it is usually competent, and it is structurally unable to settle an argument, because the tester chose the test.
  • Second-party testing is the buyer testing before purchase. Honest, but expensive, slow, and rarely repeatable: most buying teams do not own the instrumentation or the adversarial expertise, and every buyer repeats the same work.
  • Third-party validation is an independent lab with no stake in the outcome, testing to a protocol written before the results are known.

One clarification we make before every engagement, because the words get used interchangeably and they are not the same thing. Certification and accreditation are formal statuses conferred by a designated body under a published scheme. Validation is the technical work underneath: the protocol, the measurement, the evidence, the report. We do the second, we do it to the discipline of the first, and we say so plainly rather than implying a stamp nobody issued.

Why AI claims need validation more than most product claims

A pump either moves 400 gallons a minute or it does not. AI systems fail in ways that ordinary product testing was not designed to catch.

  • The number has no conditions attached. "97% accurate" is not a claim until you know on what distribution, against what labels, at what threshold, and how it degrades when the input drifts away from the test set.
  • Benchmark contamination. If the evaluation data was anywhere in the training corpus, the score measures memory, not capability. Only an outside party holding a held-out set can rule it out.
  • Demo conditions. Latency measured on an unloaded machine, accuracy measured on clean audio, autonomy measured with an engineer sitting beside the robot. All three are real measurements of an unreal situation.
  • Nondeterminism. The same prompt, the same model, and a different answer. A single passing run is not evidence; a distribution over many runs is.
  • Adversarial surface. Prompt injection, tool misuse, data exfiltration through an agent's own permissions. These are not found by the tests a builder writes for the behavior they intended.
  • Energy and thermals. On hardware, the claimed power figure is often steady-state on a bench at room temperature. What matters is the sustained draw at load, in the enclosure, at the temperature the thing will actually live at.

None of this makes vendors dishonest. It makes their measurements unfalsifiable, which for a buyer with a budget and a mission is the same problem.

Case study: how the hardest industry in the world validates a claim

Consider extreme ultraviolet lithography, the process that prints the smallest features on modern chips. ASML is the only company in the world that builds EUV lithography systems. Its High-NA generation raises the numerical aperture of the optics to 0.55 to resolve finer features than the previous 0.33 systems, and a single machine represents one of the largest capital purchases any manufacturer makes.

Now consider the buyer's position. The performance claims are at the edge of physics. Almost nobody outside the company has the instrumentation to check them. The purchase decision commits a fabrication plant to a process node for years. This is the worst possible case of the problem above: an unverifiable claim, an enormous consequence, and a single source.

The semiconductor industry did not solve this by trusting harder. It solved it structurally, through imec, the independent nanoelectronics research institute in Leuven, Belgium. ASML and imec operate a joint High-NA EUV lab, and imec's role in the industry is precisely the third-party one: chipmakers, materials suppliers and equipment vendors evaluate on neutral ground, to shared protocols, in a facility whose institutional interest is the result rather than the sale. When a resist supplier and a tool vendor disagree about whose component limits the process, the arbitration happens somewhere neither of them owns.

Three features of that arrangement are the whole lesson, and they transfer directly to AI:

  • The validator is separate from both seller and buyer. Not neutral in intention. Neutral in structure, with no equity, no resale margin, and no fee contingent on the answer.
  • The protocol is fixed before the measurement. What counts as passing is written down while it can still be argued about honestly.
  • The result is reproducible by someone else. A validation that cannot be repeated is an opinion with a logo on it.

To be explicit, because it matters in a field full of borrowed credibility: ASML and imec are cited here as the public, well-documented example of how a serious industry structures independent validation. They are not clients of The AI Cowboys, and nothing here should be read as a claim that they are.

What a validation engagement actually looks like

Six phases. Every one of them produces an artifact the client keeps.

  1. Claim decomposition. We take the datasheet, the pitch deck and the marketing page and turn each assertion into something testable, with its conditions written out. "Real-time" becomes a p99 latency budget at a stated concurrency. "Secure" becomes a named threat model. Claims that cannot be made testable are reported as such, which is itself a finding.
  2. Protocol design. The test plan, the data, the instrumentation, the acceptance criteria and the sample sizes, agreed and frozen before anything is measured. Where a public standard applies, we test to it rather than inventing our own.
  3. Instrumented measurement. Accuracy against held-out data the vendor has never seen. Latency at the percentiles that matter rather than the mean. Power at the wall under sustained load. Throughput under the concurrency the deployment will really see.
  4. Adversarial testing. Red-team work against the model and the system around it: prompt injection, jailbreaks, tool abuse, data extraction, and for hardware, out-of-spec inputs, thermal stress and fault injection. This is where the interesting failures live.
  5. Reproducibility package. Scripts, seeds, dataset manifests, hardware configuration and environment, delivered so a competent engineer at another organization can rerun the whole thing and get the same numbers.
  6. The report. What the product does, under exactly which conditions, where it degrades and where it fails, written for a technical buyer and for the contracting officer who will read it later. Both the passing and the failing results appear, because a report that only contains good news is marketing.

What we measure

Depending on what is being validated, an engagement covers some or all of:

  • Task performance with its conditions: accuracy, precision and recall, calibration, and behavior under distribution shift.
  • Latency and throughput at p50, p95 and p99 under realistic load, not best case on an idle host.
  • Energy per inference and sustained watts under load, measured at the wall. We do a great deal of this work on neuromorphic and edge hardware, where the power budget is the constraint that decides what can be fielded at all.
  • Determinism and repeatability across runs, seeds and model versions.
  • Security against a stated threat model: adversarial inputs, prompt injection, model extraction, agent permission boundaries, supply chain and SBOM.
  • Data provenance: where the training data came from, what license governs it, and whether the evaluation set is genuinely held out.
  • Governance alignment to the frameworks federal and regulated buyers are held to, including the NIST AI Risk Management Framework, the NIST adversarial machine learning taxonomy, secure software development practice under Executive Order 14028, and the evidence CMMC and FedRAMP reviewers ask for.

Hardware and software are different jobs

For hardware, the questions are physical: sustained performance in the enclosure rather than on the bench, power and thermals at the duty cycle it will really run, behavior at the edges of the operating envelope, and what happens when a sensor lies or a supply sags. We run this in our own lab in San Antonio, on instrumentation we control.

For software and models, the questions are statistical and adversarial: held-out evaluation, contamination checks, calibration, drift, and a red team that treats the system as a target rather than a demo. Where the product is an agent, the boundary between what it may do alone and what needs a human is itself something to be tested, not something to be read in a policy document.

For systems that are both, and most interesting products now are, the validation has to cross the boundary. A model that hits its accuracy target at a power draw the platform cannot supply has not met the claim.

What independence requires

A validator with a side deal is not a third party. Ours are the conditions we hold ourselves to, and the ones any buyer should demand of anyone offering this service:

  • No equity in the vendor, no reseller margin, no referral fee.
  • No fee contingent on the outcome. The price is the same whether the product passes or fails.
  • The protocol agreed and frozen before measurement begins.
  • Failures reported. A client may keep a report private, but a client cannot have the failures removed from it.
  • Conflicts disclosed in the report itself, including any prior engineering relationship with the vendor.

Common questions

Is a validation report the same as a certification?

No. A certification is a status granted under a published scheme by the body that owns that scheme. A validation report is independent technical evidence: what was tested, under what conditions, and what happened. In practice the report is what a technical buyer, a program office or an investor's diligence team actually reads, and it is frequently what a certification body will ask you to produce anyway.

How long does an engagement take?

A focused validation of one or two specific claims is measured in weeks. A full assessment across performance, energy, security and governance alignment is measured in months, and the schedule is dominated by test design and by access to representative data, not by the measurement itself.

Can the results stay confidential?

Yes. Most engagements are covered by an NDA and the report belongs to the client. What is not negotiable is the content of the report: we do not remove findings, and we do not issue a summary that contradicts the evidence behind it.

Can you work on sensitive or air-gapped systems?

Yes. We are a Service-Disabled Veteran-Owned Small Business registered in SAM.gov, CAGE 9V9EO, working with federal agencies and the defense community from San Antonio. Testing can run on premises, air-gapped, or in the client's own cloud, and nothing has to leave the perimeter for us to measure it.

Do you validate products you helped build?

Not as a third party, no. If we engineered it, we are the first party, and we will say so. That is the whole point of the arrangement.

Get the claim tested

If you are bringing hardware or software to a federal program, an enterprise procurement or a funding round, the fastest way past "prove it" is evidence somebody else produced. Tell us what your product claims, and we will tell you what it would take to demonstrate it, along with the conditions under which we expect it to hold.

Talk to The AI Cowboys about a validation engagement, or email contact_us@theaicowboys.com. We are at the UT San Antonio College of AI, Cybersecurity and Computing in downtown San Antonio, and the web address is www.theaicowboys.com.

Further reading