
Traditional software warranty approaches don’t fit AI products. Yesterday's webinar with Laura Belmont and Matt Kohel on this topic made that clear. Both speakers kept returning to the same starting point. Before anyone drafts the warranty, someone has to figure out what the AI product is supposed to do.
This disconnect happens in part because the products are so different. Software is deterministic. The same input produces the same output, making it easier to tell if it matches the documentation or not. But generative AI output is probabilistic. The same inputs can produce different outputs, so there is no “right” output to measure the product against. Matt Kohel highlighted how AI products keep evolving. Vendors update and switch models or make other significant updates on a regular basis, so the warranty has to work as a living commitment and not a snapshot of the product at signing.
The disconnect isn’t just between software and AI products. The other reason is that not all AI products are the same. They range from a chatbot that answers any question, to a tool that extracts terms from invoices, to a system that answers only from a defined set of documents. A warranty written for deterministic software may not fit any of them, and the same general AI product warranty will not work for all of them.
I really loved Laura Belmont’s framework for evaluating AI product warranties. She starts with the job the tool is supposed to do. Then she asks whether that job has a right answer, and if it does, how many right answers are possible. If the parties agree that the AI product produces a correct answer, then they can create a warranty that reflects the specific way that it works.
We should be building the warranty tailored to the job it does and make sure it includes six essential elements. These elements are:
The metric itself. That states what the output has to get right, such as the payment terms pulled from each invoice.
The dataset. The dataset is the data the test runs against, and the parties should say whether it is the vendor's test data or the customer's own data.
The threshold. The threshold is the number the output has to reach to pass, and the number below which the vendor has a breach to cure.
The testing method. The testing method states who runs the test, over what period, and on what sample, so the parties do not run different tests and reach different results.
The time of measurement. A product that passes at signing may not pass a year later, because datasets change and models drift, so the standard has to run for the subscription and allow retesting when the vendor changes the model.
The consequence. A warranty exists to name what happens when the product fails, whether that is a cure within a set number of days, a credit, or a termination right.
The framework is straightforward to describe and harder to draft. It requires that we figure out what the tool does, whether the job has a right answer, and how to measure it. But doing so allows us to create warranty terms that work for the specific product instead of arguing over adjectives.





