Preprint
Article

This version is not peer-reviewed.

AI as a Credence Good: Quality Competition, Limited Verification, and the Industrial Organization of Disclosure

Submitted:

10 March 2026

Posted:

11 March 2026

You are already at the latest version

Abstract
In this paper we study competition between AI providers when users cannot fully verify model quality. Many AI services are not well described as standard search goods, and they are not pure experience goods either. Price, interface quality, and latency are usually observable, but reliability, hallucination risk, and the downstream cost of error are often only imperfectly observable even after use. Building on the emerging view that algorithmic advice can exhibit credence-good features, we embed that insight in a static industrial-organization model of vertical quality differentiation, costly certification, and limited user comprehension. Two firms choose whether to offer a low-quality or high-quality model. High-quality AI reduces error risk, but it is slower and costlier. A high-quality firm can purchase credible certification or disclosure, yet only a subset of users can interpret it. We characterize the pure-strategy equilibrium set, show how low-quality pooling can arise even when superior technology exists, and identify a quality-trap region in which the unique market equilibrium is low-quality pooling although an allocation with one high-quality provider is welfare superior. We then analyze policy. Standardized certification works through the demand side by increasing the fraction of users who can reward quality; minimum quality standards work directly but bluntly; liability shifts firms' cost incentives and weakly shrinks the region in which a low-quality industry outcome can be sustained. Contrary to a common rhetorical move in AI policy debates, these instruments are not interchangeable. The model also clarifies that our framework is a certification model rather than a full Grossman-Milgrom unraveling game: the key distortion comes from limited user comprehension of costly, truthful quality communication. The results offer a tractable industrial-organization foundation for current debates over hallucinations, model evaluation, AI documentation, and governance.
Keywords: 
;  ;  ;  ;  

1. Introduction

Generative AI services are increasingly sold and adopted in market environments in which quality is hard to verify. A user can usually observe subscription price, rough speed, interface design, integration with other software, and perhaps a few headline benchmark claims. But the user often cannot observe the true reliability of the model before purchase, and may not be able to verify correctness after use either. A fluent answer may contain fabricated citations; a programming solution may fail only in edge cases; a legal or medical recommendation may look persuasive to a nonexpert despite being wrong. For a large class of economically important uses, AI quality therefore has a substantial credence-good component.
That conceptual placement is not itself novel. In particular, Biermann et al. (2022) explicitly argue, and experimentally document, that algorithmic advice can be perceived as a credence good even after repeated use. What remains underdeveloped is the industrial-organization theory that follows once one takes that observation seriously. If AI quality is difficult to verify, what kind of market equilibrium should we expect? Under what conditions do firms invest in lower-hallucination systems, more robust retrieval, or stronger safeguards? How much can credible certification or documentation accomplish when only some users can decode it? And when do liability or minimum quality standards outperform disclosure-based governance?
These questions are central because contemporary AI markets display a salient vertical trade-off. More reliable systems are often more expensive to build and, at least in some applications, slower to run or more likely to abstain. Faster and cheaper models can look attractive on dimensions that users immediately notice, while their reliability deficit may be much harder to evaluate. That combination is precisely what makes industrial-organization analysis valuable. Competition alone does not guarantee that firms will optimize socially important dimensions of quality when those dimensions are weakly rewarded by demand.
Our starting point is classic. Akerlof (1970) established that quality uncertainty can degrade market outcomes, while Darby and Karni (1973) and the subsequent credence-good literature emphasized that some qualities remain difficult to evaluate even after consumption. AI fits that structure in an especially forceful way. Many answers are not self-verifying. Even when a user eventually discovers a mistake, the discovery may come too late, or only after the user has already made an important decision. Moreover, a large recent literature in machine learning and human-computer interaction emphasizes that model documentation, benchmark design, and evaluation accessibility matter because ordinary users and even expert organizations often struggle to diagnose AI failure modes (Bommasani et al. 2021; Jakesch et al. 2023; Ji et al. 2023; Liang et al. 2023; Mitchell et al. 2019).
We develop a static duopoly model in which firms first choose whether to offer a low-quality or high-quality model, then choose certification intensity if they adopt the high-quality technology, and finally compete in prices. The high-quality model has a higher marginal cost and an additional fixed cost; it also incurs a latency penalty relative to the low-quality model. Certification is truthful and credible but costly. Crucially, only a fraction of users can understand or act on certification. Those users correctly recognize the reliability gain from the high-quality product; the rest mostly respond to price and speed. We interpret certification broadly: standardized evaluation reports, third-party audits, model cards, or any institution that truthfully communicates reliability in a way that is at least partially legible to the market.
The model yields four main results.
First, we characterize the pure-strategy equilibrium set in threshold form. There is always a region in which a symmetric low-quality industry outcome is sustainable, a region in which a symmetric high-quality industry outcome is sustainable, and, when the fixed cost of quality lies between those regions, a region in which one firm differentiates upward and the other stays low quality. The existence of a premium reliable tier is therefore endogenous rather than imposed.
Second, certification need not be privately sufficient to support separation. The reason is not that certification is false or noisy; it is that only some users can convert truthful certification into willingness to pay. A high-quality firm buys certification only if the resulting demand response is strong enough to offset the cost and latency disadvantages of higher quality. This is the central market failure: quality may be real and socially valuable without being adequately rewarded in demand.
Third, there exists a quality-trap region in which the unique pure-strategy equilibrium is low-quality pooling even though an allocation with one high-quality provider and one low-quality provider is welfare superior. This region emerges when the fixed cost of building a reliable system is too high for private incentives but not too high for social efficiency. The wedge is driven by limited user comprehension of certification. In equilibrium, market demand underweights reliability relative to its realized social value.
Fourth, policy tools work through distinct margins. Standardized certification and better model evaluation raise the fraction of users who can reward quality and can lower the effective cost of quality communication. Minimum quality standards remove low-quality actions entirely, which can be desirable in uniformly high-stakes applications but is blunt in heterogeneous markets. Liability operates differently again: by raising the expected private cost of lower-quality AI, it shifts firms’ incentives directly rather than relying on user inference. We show that liability weakly shrinks the region in which low-quality pooling is sustainable and weakly expands the region in which symmetric high-quality supply is sustainable. The effect on the separating region is more subtle and generally ambiguous.
Two clarifications are important. First, our paper does not claim originality for the proposition that AI can have credence-good properties. Rather, our contribution is to embed that premise in a tractable oligopoly model of quality choice, certification, and policy. Second, our framework is not a full unraveling model in the Grossman–Milgrom sense. We do not study a revelation game in which every type chooses whether and how much to disclose and users draw Bayesian inferences from silence. Instead, we study costly certification by higher-quality providers when users have limited ability to process that certification. Grossman–Milgrom style insights remain useful for comparison, but the core distortion here is different: certification is truthful yet insufficiently effective because comprehension is limited.
The paper contributes to several literatures. Relative to classic vertical-differentiation models, we endogenize certification and impose a speed–quality trade-off. Relative to credence-good models, we analyze platform competition rather than expert fraud or overtreatment. Relative to the disclosure literature, we study a market in which quality communication is costly and only partially understood. Relative to AI-governance work, we supply a micro-founded industrial-organization explanation for why markets may reward fluency, speed, and price more strongly than reliability.
The rest of the paper proceeds as follows. Section 2 positions the paper in the relevant literatures. Section 3 presents the model. Section 4 characterizes equilibrium, the quality trap, and liability. Section 5 develops the welfare and policy analysis. Section 6 concludes. Proofs are collected in the Appendix.

3. Model

3.1. Firms, Quality, and Certification

There are two firms, indexed by i ∈ { 1 , 2 } , located at the endpoints of a Hotelling line [ 0 , 1 ] . Users are uniformly distributed along the line. Horizontal differentiation captures interface familiarity, integration with existing software, prompt libraries, switching frictions, and other nonquality reasons why users may prefer one provider over another. Let t > 0 denote the Hotelling transportation parameter.
Each firm chooses one of two technologies:
q i ∈ { L , H } .
The low-quality technology L is normalized as the baseline. The high-quality technology H reduces expected error and hallucination risk. We denote the realized per-user value of that reliability gain by v > 0 . The high-quality technology also imposes a latency or convenience penalty ℓ > 0 and a marginal production cost premium c > 0 . In addition, any firm that adopts H pays a fixed cost F > 0 .
If a firm adopts H, it may purchase certification intensity x ≥ 0 . Certification is truthful and verifiable, and we interpret it broadly: audited evaluation reports, standardized benchmark disclosure, model cards with externally verifiable content, or third-party certification. Certification is costly:
K ( x ) = k 2 x 2 , k > 0 .
Only a fraction of users can understand and exploit certification. Let that fraction be
λ ( x ) = λ 0 + ρ x ,
where λ 0 ∈ [ 0 , 1 ) is baseline market comprehension and ρ > 0 is the effectiveness of certification. We assume parameters such that λ ( x ) ∈ [ 0 , 1 ] on the equilibrium path.
This reduced form captures two realistic features of AI markets. First, even truthful quality communication may be difficult for ordinary users to process. Second, standardized, comparable certification can make high-quality systems more legible.

3.2. Users and Demand

Users differ horizontally by location z ∈ [ 0 , 1 ] . Consider the asymmetric case in which firm 1 offers H with certification intensity x and firm 2 offers L. A user who can understand certification correctly perceives the reliability benefit v from the high-quality product. A user who cannot understand certification does not attach that value ex ante. All users perceive the latency penalty ℓ.
Accordingly, the average perceived vertical advantage of the high-quality product is
Δ ( x ) = λ ( x ) v − ℓ .
A user at location z has utilities
U 1 = u ¯ + Δ ( x ) − p 1 − t z , U 2 = u ¯ − p 2 − t ( 1 − z ) ,
where u ¯ is large enough to guarantee full market coverage.
The indifferent user satisfies
u ¯ + Δ ( x ) − p 1 − t z = u ¯ − p 2 − t ( 1 − z ) ,
so
z * = 1 2 + Δ ( x ) − p 1 + p 2 2 t .
Hence demand for firm 1 is
D 1 = 1 2 + Δ ( x ) − p 1 + p 2 2 t , D 2 = 1 − D 1 .
Two points are worth stressing. First, the market response depends on perceived quality, not actual quality. Users who do not understand certification underweight reliability in their purchase decision even though they will later enjoy the realized reliability gain if they consume the high-quality product. Second, the model intentionally compresses user heterogeneity into a tractable reduced form. One can reinterpret v as the expected value of error reduction across users, while λ ( x ) captures the mass of users with enough verification ability, sophistication, or organizational support to act on certification. The same logic survives if one lets users differ in risk sensitivity, the cost of being wrong, or the value of speed; what matters is that only a subset of the market internalizes reliability at the point of purchase.

3.3. Timing

The game has three stages.
  • Firms simultaneously choose quality q i ∈ { L , H } .
  • Any firm choosing H selects certification intensity x i ≥ 0 .
  • Firms compete in prices.
We solve for subgame-perfect equilibrium.

3.4. Regularity Conditions

Let
a ≡ ρ v , D ≡ 9 t k − a 2 .
We assume
Assumption 1.
D = 9 t k − a 2 > 0 .
Assumption 1 guarantees strict concavity of the certification problem.
Let
B ≡ 3 t − c − ℓ + λ 0 v .
We focus on the economically relevant interior region in which the premium provider has positive demand in the asymmetric subgame and the low-quality rival does not disappear. This is ensured by:
Assumption 2.
0 < B + a x * < 6 t ,
where x * denotes the equilibrium certification level derived below.
Assumption 2 rules out corner solutions in the pricing stage.

4. Equilibrium Analysis

4.1. Pricing and Certification in the Asymmetric Subgame

Suppose firm 1 chooses H and firm 2 chooses L. Firm profits are
Π H = ( p 1 − c ) D 1 − F − k 2 x 2 , Π L = p 2 D 2 .
Using (2), we obtain the pricing equilibrium.
Lemma 1.
Given an asymmetric quality profile ( H , L ) and certification intensity x, the unique Nash equilibrium in prices is
p H * ( x ) = t + Δ ( x ) + 2 c 3 ,
p L * ( x ) = t + c − Δ ( x ) 3 .
The corresponding market shares are
s H ( x ) = 3 t + Δ ( x ) − c 6 t ,
s L ( x ) = 3 t − Δ ( x ) + c 6 t .
Profits are
Π H ( x ) = 3 t + Δ ( x ) − c 2 18 t − F − k 2 x 2 ,
Π L ( x ) = 3 t − Δ ( x ) + c 2 18 t .
Lemma 1 already shows the central trade-off. Higher quality is rewarded only to the extent that certification-induced comprehension makes reliability salient enough to overcome the latency penalty and the marginal cost premium. A firm can be objectively better and yet only weakly differentiated in demand.
Using (1), equation (7) becomes
Π H ( x ) = ( B + a x ) 2 18 t − F − k 2 x 2 .
The high-quality firm’s certification problem is therefore quadratic.
Proposition 1.
Under Assumption 1 and B > 0 , the high-quality firm’s profit in the ( H , L ) subgame is strictly concave in x. The unique optimal certification level is
x * = a B D .
The resulting gross operating profit of the high-quality firm, net of certification cost but before subtracting the fixed quality cost F, is
Π ¯ H ≡ max x ( B + a x ) 2 18 t − k 2 x 2 = k B 2 2 D .
Proposition 1 yields immediate comparative statics. Certification is more valuable when baseline user comprehension λ 0 is higher, when certification is more legible ( ρ is higher), and when the underlying reliability benefit v is larger. It is less valuable when quality is slower (ℓ is larger), when producing quality is more expensive (c is larger), and when certification itself is expensive (k is larger). These are not just engineering facts; they are determinants of equilibrium industrial structure.

4.2. Quality-Stage Equilibrium

When both firms choose the same technology, there is no incentive to certify in equilibrium because certification no longer changes relative demand. Hence
Π L L = t 2 , Π H H = t 2 − F .
Define two threshold values:
F H ≡ Π ¯ H − t 2 ,
F L ≡ t 2 − Π L ( x * ) .
The threshold F H is the largest fixed cost consistent with a profitable deviation from ( L , L ) to high quality. The threshold F L is the largest fixed cost consistent with high quality being a best response to a high-quality rival.
Proposition 2.
Under Assumptions 1 and 2, the pure-strategy quality-stage equilibria are characterized as follows.
(i) 
( L , L ) is a pure-strategy equilibrium if and only if
F ≥ F H .
(ii) 
( H , H ) is a pure-strategy equilibrium if and only if
F ≤ F L .
(iii) 
An asymmetric pure-strategy equilibrium, ( H , L ) or ( L , H ) , exists if and only if
F L ≤ F ≤ F H .
If F L > F H , the asymmetric interval is empty and the model exhibits a multiplicity region
F ∈ [ F H , F L ]
in which both symmetric profiles, ( L , L ) and ( H , H ) , are pure-strategy equilibria.
Proposition 2 is the core equilibrium classification. It clarifies that three qualitatively distinct industry structures may arise: low-quality pooling, high-quality pooling, or vertical differentiation with a premium reliable provider and a low-cost rival. The model also permits multiplicity among symmetric equilibria when the incentive to deviate upward from ( L , L ) is weaker than the incentive to stay at H once both firms are already there. That possibility matters for welfare because it shows that equilibrium selection cannot be assumed away.
Several implications follow.
First, limited user comprehension expands the low-quality region. Since Π ¯ H depends positively on the effectiveness of certification, lower λ 0 or lower ρ reduce F H and thereby make ( L , L ) easier to sustain.
Second, the market can fail to create a premium reliable tier even when quality is technically feasible. If the fixed cost F exceeds F H , no firm wants to be the first mover into a high-quality niche.
Third, stronger competition in the sense of easier substitution is not guaranteed to improve reliability incentives. What matters is not the number of firms per se, but whether the market reward to reliability is strong enough to support costly differentiation.
Corollary 1.
In the interior solution, x * and F H are increasing in λ 0 , ρ, and v, and decreasing in c, ℓ, and k.
Corollary 1 gives a compact comparative-static rationale for standardized evaluation, interoperable certification, and more legible model documentation. Those policies do not merely “improve transparency” in a loose sense; they expand the parameter region in which market incentives support reliable AI.

4.3. A Quality Trap

We now turn to welfare. The crucial distinction is between perceived quality, which governs demand, and realized quality, which governs surplus. When users who do not understand certification consume the high-quality product, they still enjoy the lower error rate ex post even though they did not fully value it ex ante.
Consider again the asymmetric allocation with one high-quality firm and one low-quality firm. Relative to ( L , L ) , welfare equals the incremental benefit from users served by the high-quality firm, minus certification cost, minus the fixed cost of developing high quality, minus the extra Hotelling distortion from asymmetric market shares. Using (5), the welfare gain from an asymmetric allocation is
W H L ( x ) − W L L = s H ( x ) ( v − ℓ − c ) − k 2 x 2 − F − t s H ( x ) − 1 2 2 .
Define the welfare threshold for an asymmetric allocation:
F W ≡ sup x ≥ 0 s H ( x ) ( v − ℓ − c ) − k 2 x 2 − t s H ( x ) − 1 2 2 .
with x ≥ 0 : 0 ≤ s H ( x ) ≤ 1 . Thus F W is the largest fixed cost for which an allocation with one high-quality firm is socially preferable to ( L , L ) .
Now define
F ¯ ≡ max { F H , F L } .
Proposition 3.
If
F W > F ¯ ,
then for every fixed cost
F ∈ ( F ¯ , F W ) ,
the unique pure-strategy equilibrium is ( L , L ) , yet an asymmetric allocation with one high-quality firm and one low-quality firm is welfare superior to ( L , L ) .
Proposition 3 formalizes the quality trap. The key wedge is simple. In demand, the reliability gain is weighted by λ ( x ) because only some users understand certification. In welfare, the realized reliability gain accrues to every user who consumes the high-quality product. When the market underperceives quality, private incentives can be too weak to sustain it even though the social return is positive.
This is not merely a repackaged lemons result. There is no exogenous population of hidden product types. Quality is chosen endogenously. The failure is therefore one of investment incentives: the market supplies too little reliability because users do not sufficiently reward it. In AI applications where hallucinations are hard to detect and costly when they occur, that is precisely the distortion policymakers worry about.

4.4. Liability

We now add a liability rule. Let d > 0 denote the reduction in expected harm per query when a firm moves from L to H. Suppose the firm faces expected liability m d per query avoided by using high quality, where m ≥ 0 is a policy parameter. Then the effective marginal cost premium of high quality becomes
c ( m ) = c − m d .
Liability does not directly change user comprehension; it changes private incentives by making unreliable AI more expensive to supply.
All the formulas above remain valid after replacing c with c ( m ) , provided the interior assumptions continue to hold. Define
B ( m ) = 3 t − c ( m ) − ℓ + λ 0 v .
Then
x * ( m ) = a B ( m ) D , Π ¯ H ( m ) = k B ( m ) 2 2 D ,
and the thresholds become
F H ( m ) = Π ¯ H ( m ) − t 2 , F L ( m ) = t 2 − Π L x * ( m ) ; c ( m ) .
Proposition 4.
For every admissible m such that the interior conditions hold, the following statements are true.
(i) 
x * ( m ) is weakly increasing in m.
(ii) 
F H ( m ) is strictly increasing in m.
(iii) 
F L ( m ) is strictly increasing in m.
Therefore liability weakly shrinks the set of fixed costs for which ( L , L ) is a pure-strategy equilibrium and weakly expands the set of fixed costs for which ( H , H ) is a pure-strategy equilibrium.
Proposition 4 is the clean liability result. Once unreliable AI generates expected legal or regulatory costs, the private profitability of quality rises. The low-quality industry outcome becomes harder to sustain, and the high-quality industry outcome becomes easier to sustain. This is strongest precisely when demand-side governance is weakest, because liability does not rely on user sophistication.
However, liability does not generally have a monotone effect on the width of the separating region [ F L ( m ) , F H ( m ) ] .
Corollary 2.
Suppose the interior solution holds and define
C ( m ) ≡ λ 0 v − ℓ − c ( m ) .
Then
d d m F H ( m ) − F L ( m ) = d k C ( m ) 18 k t − a 2 + 3 a 2 t D 2 .
Hence liability expands the separating region if and only if
C ( m ) > − 3 a 2 t 18 k t − a 2 .
Otherwise the separating region contracts.
Corollary 2 is important for policy interpretation. Liability unambiguously pushes the industry away from low-quality pooling and toward high-quality outcomes, but it need not enlarge the set of parameters that generate vertical differentiation. In some environments, liability makes symmetric high-quality supply more attractive and thereby compresses the region in which one firm remains low quality. The policy lesson is that liability is primarily a tool for moving the market toward reliability, not necessarily for preserving a two-tier structure.

5. Welfare, Market Failure, and Policy

5.1. Why the Market Underprovides Reliability

The model produces underprovision of reliability for a structurally clear reason. The willingness to pay that firms face is governed by perceived quality, not actual quality. Users who cannot decode certification still enjoy the realized benefits of more reliable AI if they end up using it, but that surplus is not fully capitalized into demand. As a result, the private return to quality can be below the social return.
This wedge is especially likely to be large in AI markets for three reasons.
First, verification is costly. Checking citations, testing edge cases, or validating factual claims often requires outside expertise or time. Second, mistakes are unevenly distributed. Many low-quality outputs look fine until they fail in a high-consequence context. Third, AI systems are often adopted by organizations in which the buyer, the user, and the party bearing downstream harm are not the same. Even without adding explicit externalities to the model, these institutional features make it realistic that market demand underweights reliability.
The model also illustrates why price competition alone is not a sufficient remedy. If firms mostly compete on dimensions users can easily observe, then more competition may intensify the race along those dimensions rather than along reliability. In AI, those dimensions are often speed, convenience, and price. The outcome can be a competitive market that is nevertheless skewed toward low-quality provision.

5.2. Certification Policy

A first policy family works by improving the information environment. Standardized benchmarking, common taxonomies of failure modes, third-party audits, certification labels, and clearer model cards map naturally into the model in three ways:
(a)
they raise baseline user comprehension λ 0 ;
(b)
they increase the effectiveness of certification ρ ;
(c)
they lower the private cost of certification k.
Corollary 1 implies that all three changes raise the private return to quality. In equilibrium language, they increase x * and F H , thereby shrinking the region in which low-quality pooling can survive.
This provides a precise economic rationale for current efforts to standardize AI evaluation and documentation (Liang et al. 2023; Mitchell et al. 2019). The point is not simply that more information is always better. Rather, standardized and legible certification changes market incentives by increasing the share of users who can reward reliability. In a market where quality is otherwise hidden behind fluency and speed, that can be the difference between a viable premium segment and a market that pools on low reliability.
At the same time, the model disciplines optimism about disclosure-based governance. Certification works only through the users who can understand it. If λ 0 remains small, if certification is costly, or if the quality improvement comes with a large latency penalty, then transparency alone may not be enough. In such cases, policymakers should not expect a disclosure mandate to replicate the effects of liability or minimum standards.
This point also clarifies how our analysis relates to unraveling. In a full Grossman–Milgrom environment, nondisclosure can itself reveal bad news if consumers reason sharply enough. Our model does not rely on that mechanism. The problem is not merely that some firms choose silence. The deeper problem is that even truthful certification has limited reach because users do not uniformly understand it. For AI governance, that distinction is practical. A model card that exists but cannot be interpreted by most buyers will not reliably discipline the market.

5.3. Minimum Quality Standards

A second policy family imposes minimum quality standards. In the model, a standard that bans low-quality AI forces the industry into ( H , H ) as long as firms continue to operate. Relative to ( L , L ) , the welfare change is
W H H − W L L = v − ℓ − c − 2 F .
Hence a minimum standard is socially desirable if and only if
v − ℓ − c > 2 F .
This criterion yields a clear policy distinction. Minimum standards are especially attractive when applications are uniformly high stakes and verification is hard. In medicine, legal compliance, safety-critical coding, or other domains where undetected hallucinations can generate substantial harm, forcing low-quality systems out of the market may be efficient. In such environments, the bluntness of a standard is not a major drawback because the social value of reliability is large for almost every user.
The picture changes when applications are heterogeneous. Many AI uses are low stakes, easily checked, or complementary to strong human oversight. In those domains, a blanket standard can overcorrect by suppressing fast and cheap tools that are socially useful despite being less reliable. Our model therefore suggests that minimum standards should be use-case specific when possible. A standard that is optimal for professional legal research need not be optimal for casual drafting or brainstorming.
In industrial-organization terms, minimum standards dominate when the regulator wants to compress the market onto a single high-quality technology. Certification policies dominate when the regulator wants to preserve segmentation while making reliable quality commercially sustainable.

5.4. Liability, Safe Harbors, and the Allocation of Responsibility

A third policy family uses liability. Proposition 4 shows that liability works through a fundamentally different channel from certification. Certification raises demand-side rewards to quality. Liability changes the firm’s effective cost comparison between H and L.
This difference matters when users are unsophisticated. If firms know much more about residual hallucination risk than buyers do, then demand-side tools can be weak because users cannot accurately map disclosure into expected harm. Liability is attractive precisely because it does not depend on that mapping. It directly internalizes part of the harm generated by lower-quality AI.
This does not mean maximal liability is always optimal. Excessive liability can chill beneficial deployment, raise entry barriers, or induce defensive overrefusal. Moreover, some AI harms are hard to attribute to a particular provider or to distinguish from user misuse. These are real concerns. But the model suggests a clear principle: liability is most valuable where the downstream cost of undetected error is large and traceable, and where user-side verification is especially poor.
A useful institutional implication follows. Policymakers can combine liability with safe harbors. For example, a provider that satisfies certain certification or audit requirements could face reduced liability exposure. In the model, such a regime would jointly improve comprehension and alter cost incentives. That combination is likely to outperform either instrument in isolation when both user understanding and firm incentives are distorted.

5.5. Heterogeneous Applications and Regulatory Targeting

The paper’s baseline model suppresses some heterogeneity for tractability, but its policy logic points toward targeted regulation. In practice, users differ in at least three ways: their ability to verify outputs, the cost they incur when the AI is wrong, and the value they place on speed. These differences can be accommodated by interpreting v and ℓ as averages over application-specific segments.
Doing so yields several useful lessons.
First, markets with a large mass of low-stakes users may rationally sustain a low-quality tier. That is not by itself a market failure. The failure occurs when the same tier spills into high-stakes applications whose users cannot verify quality. This is one reason why enterprise settings often demand procurement rules, vendor audits, or ex ante certification.
Second, broad platform-level regulation can be too crude. The optimal policy mix depends on whether the relevant market segment is dominated by uninformed users, by high-stakes use cases, or by strong downstream monitoring. The model therefore supports application-contingent regulation more strongly than one-size-fits-all rules.
Third, public or industry-funded evaluation infrastructure can have large leverage. When third parties generate credible, legible comparisons across models, the effective values of λ 0 and ρ rise for the whole market. That can tilt competition toward reliable systems without requiring pervasive command-and-control regulation.

5.6. Dynamic Interpretation

Although our model is static, its logic has a natural dynamic reading. Repeated adoption and learning do not automatically eliminate the information problem if users cannot accurately diagnose AI errors. In that sense, a static certification model is not merely a shortcut for a reputation model; it captures an environment in which reputation itself may be slow or noisy to form.
This perspective helps explain why AI policy discussions have focused so heavily on evaluation and documentation. In standard experience-good markets, firms can often build reputation through repeated consumer feedback. In credence-good environments, reputation is weaker because users may not know whether they were served well. The same is true for AI when answers sound persuasive regardless of correctness. Certification and liability therefore become substitutes for missing reputational discipline.

6. Conclusion

We have developed a theory of AI competition when model quality is difficult for users to verify. The paper takes seriously the increasingly common observation that algorithmic advice can have credence-good features and derives the corresponding industrial-organization consequences.
Three broad conclusions follow.
First, markets can rationally pool on low-quality AI even when better technology exists. The core reason is that reliability is not fully priced when only some users understand certification. Second, there is a quality-trap region in which low-quality pooling is the unique pure-strategy equilibrium even though an allocation with one reliable provider is welfare superior. Third, policy instruments should be distinguished by the margins through which they operate. Certification improves the market reward to quality; minimum standards remove low-quality actions; liability alters firms’ private cost comparison directly.
The model is deliberately parsimonious, but it yields an implication of wider significance. In AI markets, competition can reward what is legible before purchase rather than what matters after use. Speed, convenience, and low price are legible. Reliability often is not. Where that gap is large, the case for policy is not a generic distrust of markets. It is a standard industrial-organization response to a market in which socially valuable quality is hard to verify.

Appendix A. Proofs

Proof of Lemma 1.
Using (2), firm profits in the asymmetric subgame are
Π H = ( p 1 − c ) 1 2 + Δ ( x ) − p 1 + p 2 2 t − F − k 2 x 2 ,
Π L = p 2 1 2 + − Δ ( x ) + p 1 − p 2 2 t .
Differentiating with respect to own price yields the first-order conditions
∂ Π H ∂ p 1 = 1 2 + Δ ( x ) − p 1 + p 2 2 t − p 1 − c 2 t = 0 ,
∂ Π L ∂ p 2 = 1 2 + − Δ ( x ) + p 1 − p 2 2 t − p 2 2 t = 0 .
Multiplying by 2 t gives
t + Δ ( x ) − 2 p 1 + p 2 + c = 0 ,
t − Δ ( x ) + p 1 − 2 p 2 = 0 .
Solving the linear system yields (3) and (). Substituting those prices into demand gives (5) and (). Substituting prices and shares into profits gives (7) and (). Uniqueness follows from strict concavity of each firm’s profit in its own price. □
Proof of Proposition 1.
From Lemma 1,
Π H ( x ) = ( B + a x ) 2 18 t − F − k 2 x 2 .
Differentiating,
d Π H d x = a ( B + a x ) 9 t − k x , d 2 Π H d x 2 = a 2 9 t − k = − D 9 t < 0
by Assumption 1. Hence the objective is strictly concave. Setting the first derivative equal to zero gives
a ( B + a x ) = 9 t k x ,
so
a B = ( 9 t k − a 2 ) x = D x ,
and therefore
x * = a B D .
Substituting into the objective,
Π ¯ H = ( B + a x * ) 2 18 t − k 2 ( x * ) 2 = k B 2 2 D .
□
Proof of Proposition 2.
In the symmetric low-quality profile, each firm earns
Π L L = t 2 .
A unilateral deviation to high quality yields profit
Π ¯ H − F .
Hence ( L , L ) is an equilibrium if and only if
t 2 ≥ Π ¯ H − F ⇔ F ≥ Π ¯ H − t 2 = F H .
In the symmetric high-quality profile, each firm earns
Π H H = t 2 − F .
A unilateral deviation from H to L against a rival that remains high quality and chooses x * yields profit
Π L ( x * ) .
Hence ( H , H ) is an equilibrium if and only if
t 2 − F ≥ Π L ( x * ) ⇔ F ≤ t 2 − Π L ( x * ) = F L .
Finally, an asymmetric profile is an equilibrium if and only if high quality is a best response to low quality and low quality is a best response to high quality. These conditions are
Π ¯ H − F ≥ t 2 ⇔ F ≤ F H ,
and
Π L ( x * ) ≥ t 2 − F ⇔ F ≥ F L .
Combining them gives
F L ≤ F ≤ F H .
If F L > F H , the interval is empty and there is no asymmetric pure-strategy equilibrium. In that case ( L , L ) exists for all F ≥ F H and ( H , H ) exists for all F ≤ F L , implying coexistence of the symmetric equilibria on [ F H , F L ] . □
Proof of Corollary 1.
From (9),
x * = ρ v 3 t − c − ℓ + λ 0 v 9 t k − ( ρ v ) 2 .
The denominator is positive by Assumption 1. The numerator is increasing in λ 0 , ρ , and v, and decreasing in c and ℓ. The denominator is increasing in k, so x * is decreasing in k. Since
F H = k B 2 2 D − t 2 ,
the same directional results hold for F H . □
Proof of Proposition 3.
Let
F ¯ = max { F H , F L } .
If F > F ¯ , then F > F H and Proposition 2 implies that no asymmetric pure-strategy equilibrium exists. Also, F > F L , so ( H , H ) is not a pure-strategy equilibrium. Since F > F H , ( L , L ) is a pure-strategy equilibrium. Hence ( L , L ) is the unique pure-strategy equilibrium for every F > F ¯ .
Now suppose additionally that F < F W . By the definition of F W in (14), there exists some certification level x such that
s H ( x ) ( v − ℓ − c ) − k 2 x 2 − t s H ( x ) − 1 2 2 − F > 0 .
By (13), this is equivalent to
W H L ( x ) − W L L > 0 .
Therefore an asymmetric allocation is welfare superior to low-quality pooling. Combining the two steps proves that for every
F ∈ ( F ¯ , F W ) ,
the unique pure-strategy equilibrium is ( L , L ) even though an asymmetric allocation is welfare superior. □
Proof of Proposition 4.
Under liability m, the effective marginal cost premium becomes
c ( m ) = c − m d .
Thus
B ( m ) = 3 t − c ( m ) − ℓ + λ 0 v = 3 t − c + λ 0 v − ℓ + m d ,
so
d B ( m ) d m = d > 0 .
From (9),
x * ( m ) = a B ( m ) D ,
hence
d x * ( m ) d m = a d D > 0 .
This proves part (i).
Next,
F H ( m ) = k B ( m ) 2 2 D − t 2 ,
so
d F H ( m ) d m = k d B ( m ) D > 0 .
This proves part (ii).
For part (iii), let
Y ( m ) ≡ B ( m ) + a x * ( m ) .
Using (9),
Y ( m ) = B ( m ) 1 + a 2 D = 9 t k D B ( m ) ,
so
d Y ( m ) d m = 9 t k d D > 0 .
From Lemma 1,
Π L ( x * ( m ) ; c ( m ) ) = 6 t − Y ( m ) 2 18 t .
Under Assumption 2, Y ( m ) < 6 t , so
d Π L d m = − ( 6 t − Y ( m ) ) Y ′ ( m ) 9 t < 0 .
Therefore
d F L ( m ) d m = − d Π L d m > 0 .
Hence both thresholds move upward with liability. Because ( L , L ) requires F ≥ F H ( m ) , the low-quality region shrinks. Because ( H , H ) requires F ≤ F L ( m ) , the high-quality region expands. □
Proof of Corollary 2.
Let
S ( m ) ≡ F H ( m ) − F L ( m ) .
From the proof of Proposition 4,
F H ′ ( m ) = k d B ( m ) D , F L ′ ( m ) = k d ( 6 t − Y ( m ) ) D .
Therefore
S ′ ( m ) = k d D B ( m ) − 6 t + Y ( m ) .
Using
Y ( m ) = 9 t k D B ( m ) , B ( m ) = 3 t + C ( m ) ,
we obtain
S ′ ( m ) = k d D 3 t + C ( m ) − 6 t + 9 t k ( 3 t + C ( m ) ) D .
After simplification,
S ′ ( m ) = d k C ( m ) 18 k t − a 2 + 3 a 2 t D 2 ,
which is (16). The sign condition follows immediately. □

References

  1. Akerlof, George A. 1970. The market for “lemons”: Quality uncertainty and the market mechanism. The Quarterly Journal of Economics 84(3), 488–500. [CrossRef]
  2. Biermann, Jan, John J. Horton, and Johannes Walter. 2022. Algorithmic advice as a credence good. ZEW Discussion Paper 22-071, ZEW – Leibniz Centre for European Economic Research. [CrossRef]
  3. Board, Oliver. 2009. Competition and disclosure. The Journal of Industrial Economics 57(1), 197–213. [CrossRef]
  4. Bommasani, Rishi, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, et al. 2021. On the opportunities and risks of foundation models. [CrossRef]
  5. Darby, Michael R. and Edi Karni. 1973. Free competition and the optimal amount of fraud. The Journal of Law and Economics 16(1), 67–88. [CrossRef]
  6. Daughety, Andrew F. and Jennifer F. Reinganum. 1995. Product safety: Liability, R&D, and signaling. The American Economic Review 85(5), 1187–1206.
  7. Daughety, Andrew F. and Jennifer F. Reinganum. 2008a. Communicating quality: A unified model of disclosure and signalling. The RAND Journal of Economics 39(4), 973–989. [CrossRef]
  8. Daughety, Andrew F. and Jennifer F. Reinganum. 2008b. Imperfect competition and quality signalling. The RAND Journal of Economics 39(1), 163–183. [CrossRef]
  9. Dranove, David and Ginger Zhe Jin. 2010. Quality disclosure and certification: Theory and practice. Journal of Economic Literature 48(4), 935–963. [CrossRef]
  10. Dulleck, Uwe and Rudolf Kerschbamer. 2006. On doctors, mechanics, and computer specialists: The economics of credence goods. Journal of Economic Literature 44(1), 5–42. [CrossRef]
  11. Emons, Winand. 1997. Credence goods and fraudulent experts. The RAND Journal of Economics 28(1), 107–119. [CrossRef]
  12. Fishman, Michael J. and Kathleen M. Hagerty. 2003. Mandatory versus voluntary disclosure in markets with informed and uninformed customers. Journal of Law, Economics, and Organization 19(1), 45–63. [CrossRef]
  13. Grossman, Sanford J. 1981. The informational role of warranties and private disclosure about product quality. The Journal of Law and Economics 24(3), 461–483. [CrossRef]
  14. Jakesch, Maurice, Jeffrey T. Hancock, and Mor Naaman. 2023. Human heuristics for AI-generated language are flawed. Proceedings of the National Academy of Sciences 120(11), e2208839120. [CrossRef]
  15. Ji, Ziwei, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Delong Chen, Wenliang Dai, Ho Shu Chan, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys 55(12), 248:1–248:38. [CrossRef]
  16. Liang, Percy, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, et al. 2023. Holistic evaluation of language models. Transactions on Machine Learning Research.
  17. Lizzeri, Alessandro. 1999. Information revelation and certification intermediaries. The RAND Journal of Economics 30(2), 214–231. [CrossRef]
  18. Milgrom, Paul R. 1981. Good news and bad news: Representation theorems and applications. The Bell Journal of Economics 12(2), 380–391. [CrossRef]
  19. Mitchell, Margaret, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pp. 220–229. [CrossRef]
  20. Mussa, Michael and Sherwin Rosen. 1978. Monopoly and product quality. Journal of Economic Theory 18(2), 301–317. [CrossRef]
  21. Shapiro, Carl. 1983. Premiums for high quality products as returns to reputations. The Quarterly Journal of Economics 98(4), 659–679. [CrossRef]
  22. Wolinsky, Asher. 1993. Competition in a market for informed experts’ services. The RAND Journal of Economics 24(3), 380–398. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.