{"id":1081,"date":"2026-07-09T04:43:29","date_gmt":"2026-07-09T04:43:29","guid":{"rendered":"https:\/\/gk.palem.in\/articles\/?p=1081"},"modified":"2026-07-09T04:55:13","modified_gmt":"2026-07-09T04:55:13","slug":"online-dispute-resolution-why-ai-should-not-settle-disputes-based-on-truth-alone","status":"publish","type":"post","link":"https:\/\/gk.palem.in\/articles\/online-dispute-resolution-why-ai-should-not-settle-disputes-based-on-truth-alone\/","title":{"rendered":"Online Dispute Resolution: Why AI Should Not Settle Disputes Based on Truth Alone"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Can an AI be unsure about the truth and still be confident about a settlement?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">At first, that sounds contradictory.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If the system does not know exactly what happened, how can it recommend what should happen next?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But this is precisely how many real-world decisions work. A doctor may not know the exact biological pathway that caused a symptom, but still know the safest clinical intervention. A compliance officer may not prove money laundering conclusively, but still know that a transaction should be blocked, investigated or escalated. A business leader may not know every internal failure that caused a project delay, but still know how to renegotiate delivery, protect the relationship and reduce further loss.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This distinction is critical for <strong>AI-powered dispute resolution<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Most people assume that a dispute-resolution AI must first determine the truth, then assign responsibility, then propose an outcome. That sounds logical, but it is not how many commercial disputes actually behave in the real world.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In low-value, high-volume disputes, the cost of proving the truth may exceed the value of the dispute itself. The missing evidence may no longer exist. The warehouse CCTV may have been overwritten. The courier may not have item-level weight logs. The signed delivery note may prove that a pallet arrived but not that the contents were individually verified. The claimant may have a photo showing eight units, but the photo may not prove that only eight units crossed the delivery threshold.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So the real question is not always:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><strong>Can we prove exactly what happened?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The better product question is:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><strong>Can we propose a fair, explainable and commercially acceptable resolution under uncertainty?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That is where the concept of <strong>Resolution Confidence<\/strong> becomes important.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Truth Confidence Is Not Resolution Confidence<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A serious AI dispute platform should not operate with a single confidence score.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">One score is too crude. It collapses different kinds of uncertainty into one number and creates the illusion of precision.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Instead, the system must separate at least three layers.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Truth Confidence<\/strong> asks: what most likely happened?<\/li>\n\n\n\n<li><strong>Responsibility Confidence<\/strong> asks: who should bear the loss, risk or burden?<\/li>\n\n\n\n<li><strong>Resolution Confidence<\/strong> asks: can the AI safely propose a fair settlement without human review?<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">These are related, but they are not the same.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A system may have only moderate confidence about factual truth but high confidence that a compromise settlement is fair. Conversely, it may know the facts very well but still be uncertain about the correct legal or commercial allocation of loss.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This distinction matters deeply for CEOs and founders building AI-native platforms. If you confuse these three layers, your product will either over-automate risky decisions or over-escalate commercially resolvable cases to expensive human review.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Neither scales.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The Warehouse Example<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Consider a simple business dispute.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A claimant ordered and paid for ten units of equipment. The respondent says all ten were dispatched. The claimant says only eight arrived. The contract only says the respondent must deliver the quantity paid for. There is no special proof-of-delivery clause.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The respondent has a courier dispatch manifest showing ten units handed over to the courier. The claimant has an internal goods-received note and photos showing eight units after unboxing. The signed delivery confirmation says the claimant received one pallet, but contents were not individually verified.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A weak AI system will ask: <strong>Who is more credible?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A better system will ask: <strong>What does each evidence artifact actually prove?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The dispatch manifest supports the proposition that ten units entered the dispatch process. It does not prove that ten units were individually delivered to the claimant\u2019s premises.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The delivery confirmation supports the proposition that one pallet was delivered. It does not prove that ten units were inside the pallet at delivery.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The photos support the proposition that eight units were visible when the photos were taken. They do not prove that only eight units arrived unless metadata, timing or first-opening evidence supports that interpretation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The claimant\u2019s goods-received note supports the proposition that the claimant internally recorded eight units. It does not automatically prove that the respondent breached the delivery obligation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So the system should not jump from documents to judgment. It should pass through a structured evidentiary model.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In simplified form:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><strong><math data-latex=\"[ \\text{Evidence} \\rightarrow \\text{Propositions} \\rightarrow \\text{Custody Hypotheses} \\rightarrow \\text{Confidence Scores} \\rightarrow \\text{Settlement Policy} ]\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">[<\/mo><mtext>Evidence<\/mtext><mo stretchy=\"false\">\u2192<\/mo><mtext>Propositions<\/mtext><mo stretchy=\"false\">\u2192<\/mo><mtext>Custody&nbsp;Hypotheses<\/mtext><mo stretchy=\"false\">\u2192<\/mo><mtext>Confidence&nbsp;Scores<\/mtext><mo stretchy=\"false\">\u2192<\/mo><mtext>Settlement&nbsp;Policy<\/mtext><mo form=\"postfix\" stretchy=\"false\">]<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">[ \\text{Evidence} \\rightarrow \\text{Propositions} \\rightarrow \\text{Custody Hypotheses} \\rightarrow \\text{Confidence Scores} \\rightarrow \\text{Settlement Policy} ]<\/annotation><\/semantics><\/math><\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That middle layer is where most AI products fail.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Evidence Does Not Speak for Itself<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In regulated domains, documents are not merely \u201cdata.\u201d They are <strong>evidence<\/strong>, and evidence must always be interpreted against a proposition.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The same document may be strong evidence for one proposition and weak evidence for another.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A signed delivery note may be strong evidence that a pallet arrived. It may be weak evidence that the correct item count arrived. A dispatch manifest may be strong evidence that a warehouse process recorded ten units. It may be weak evidence that the claimant received ten units. A photo may be strong evidence that eight items were visible. It may be weak evidence that the shipment originally contained only eight.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is why treating dispute resolution as a document summarization problem is a category error.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An LLM can summarize the documents beautifully and still misunderstand what they prove.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For any evidence item <math data-latex=\"e_i\"><semantics><msub><mi>e<\/mi><mi>i<\/mi><\/msub><annotation encoding=\"application\/x-tex\">e_i<\/annotation><\/semantics><\/math>, the system must evaluate its relevance to a proposition <math data-latex=\"\\phi_m\"><semantics><msub><mi>\u03d5<\/mi><mi>m<\/mi><\/msub><annotation encoding=\"application\/x-tex\">\\phi_m<\/annotation><\/semantics><\/math>, not to the dispute in general. A simplified support score may look like this:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\"[ S_m = \\sum_i A_i \\gamma_{im} q_{im} \\lambda_{im} ]\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">[<\/mo><msub><mi>S<\/mi><mi>m<\/mi><\/msub><mo>=<\/mo><msub><mo movablelimits=\"false\">\u2211<\/mo><mi>i<\/mi><\/msub><msub><mi>A<\/mi><mi>i<\/mi><\/msub><msub><mi>\u03b3<\/mi><mrow><mi>i<\/mi><mi>m<\/mi><\/mrow><\/msub><msub><mi>q<\/mi><mrow><mi>i<\/mi><mi>m<\/mi><\/mrow><\/msub><msub><mi>\u03bb<\/mi><mrow><mi>i<\/mi><mi>m<\/mi><\/mrow><\/msub><mo form=\"postfix\" stretchy=\"false\">]<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">[ S_m = \\sum_i A_i \\gamma_{im} q_{im} \\lambda_{im} ]<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here, <math data-latex=\"A_i\"><semantics><msub><mi>A<\/mi><mi>i<\/mi><\/msub><annotation encoding=\"application\/x-tex\">A_i<\/annotation><\/semantics><\/math> indicates whether the evidence is present, <math data-latex=\"(\\gamma_{im})\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msub><mi>\u03b3<\/mi><mrow><mi>i<\/mi><mi>m<\/mi><\/mrow><\/msub><mo form=\"postfix\" stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">(\\gamma_{im})<\/annotation><\/semantics><\/math> measures relevance to the proposition, <math data-latex=\"(q_{im})\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msub><mi>q<\/mi><mrow><mi>i<\/mi><mi>m<\/mi><\/mrow><\/msub><mo form=\"postfix\" stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">(q_{im})<\/annotation><\/semantics><\/math> measures evidence quality, and <math data-latex=\"(\\lambda_{im})\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msub><mi>\u03bb<\/mi><mrow><mi>i<\/mi><mi>m<\/mi><\/mrow><\/msub><mo form=\"postfix\" stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">(\\lambda_{im})<\/annotation><\/semantics><\/math> measures the evidentiary direction and strength.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is the principle: A document should not receive one global credibility score. Its strength depends on what proposition it is being used to prove.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The Missing Evidence Problem<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The most important evidence in a dispute is often the evidence that is absent.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In the warehouse case, the crucial missing element is item-level delivery verification at the claimant\u2019s premises. The respondent proves dispatch better than delivery. The claimant proves receipt-shortage observation better than delivery-stage causation. Neither proves item-level delivery or item-level shortage at the exact custody handoff.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This creates a critical evidence gap.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Let:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\"[ S_D = \\text{support for dispatch of ten units} ]\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">[<\/mo><msub><mi>S<\/mi><mi>D<\/mi><\/msub><mo>=<\/mo><mtext>support&nbsp;for&nbsp;dispatch&nbsp;of&nbsp;ten&nbsp;units<\/mtext><mo form=\"postfix\" stretchy=\"false\">]<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">[ S_D = \\text{support for dispatch of ten units} ]<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\"[ S_R = \\text{support for observed receipt shortage} ]\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">[<\/mo><msub><mi>S<\/mi><mi>R<\/mi><\/msub><mo>=<\/mo><mtext>support&nbsp;for&nbsp;observed&nbsp;receipt&nbsp;shortage<\/mtext><mo form=\"postfix\" stretchy=\"false\">]<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">[ S_R = \\text{support for observed receipt shortage} ]<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\"[ \\operatorname{Cov}_L = \\text{coverage of item-level delivery evidence} ]\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">[<\/mo><msub><mi>Cov<\/mi><mi>L<\/mi><\/msub><mo>\u2061<\/mo><mspace width=\"0.1667em\"><\/mspace><mo>=<\/mo><mtext>coverage&nbsp;of&nbsp;item-level&nbsp;delivery&nbsp;evidence<\/mtext><mo form=\"postfix\" stretchy=\"false\">]<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">[ \\operatorname{Cov}_L = \\text{coverage of item-level delivery evidence} ]<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Then, in this case, we may have:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\"[ S_D \\uparrow,\\quad S_R \\text{ moderate},\\quad \\operatorname{Cov}_L \\downarrow ]\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">[<\/mo><msub><mi>S<\/mi><mi>D<\/mi><\/msub><mo stretchy=\"false\">\u2191<\/mo><mo separator=\"true\">,<\/mo><mspace width=\"1em\"><\/mspace><msub><mi>S<\/mi><mi>R<\/mi><\/msub><mtext>&nbsp;moderate<\/mtext><mo separator=\"true\">,<\/mo><mspace width=\"1em\"><\/mspace><msub><mi>Cov<\/mi><mi>L<\/mi><\/msub><mo>\u2061<\/mo><mo stretchy=\"false\">\u2193<\/mo><mo form=\"postfix\" stretchy=\"false\">]<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">[ S_D \\uparrow,\\quad S_R \\text{ moderate},\\quad \\operatorname{Cov}_L \\downarrow ]<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That low item-level delivery coverage creates what I call a <strong>delivery evidence gap<\/strong>:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\"[ G_L = \\rho_L(1-\\operatorname{Cov}_L) ]\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">[<\/mo><msub><mi>G<\/mi><mi>L<\/mi><\/msub><mo>=<\/mo><msub><mi>\u03c1<\/mi><mi>L<\/mi><\/msub><mo form=\"prefix\" stretchy=\"false\">(<\/mo><mn>1<\/mn><mo>\u2212<\/mo><msub><mi>Cov<\/mi><mi>L<\/mi><\/msub><mo>\u2061<\/mo><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo form=\"postfix\" stretchy=\"false\">]<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">[ G_L = \\rho_L(1-\\operatorname{Cov}_L) ]<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">where <math data-latex=\"(\\rho_L)\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msub><mi>\u03c1<\/mi><mi>L<\/mi><\/msub><mo form=\"postfix\" stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">(\\rho_L)<\/annotation><\/semantics><\/math> represents the importance of item-level delivery proof in that dispute type.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This gap should not be treated as automatic proof against either party. It is not a magical shortcut to liability. It is an uncertainty amplifier.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">And that uncertainty affects the system differently depending on which confidence score is being computed.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It lowers <strong>Truth Confidence<\/strong> because the exact failing link in the chain-of-custody is unclear.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It lowers <strong>Responsibility Confidence<\/strong> because the system cannot confidently assign the loss to one party.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But it may increase the defensibility of a <strong>risk-sharing settlement<\/strong>, because a one-sided outcome becomes harder to justify when the critical evidence gap is large.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is the design insight.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Missing evidence does not always mean \u201cdo nothing.\u201d Sometimes it means \u201cdo not adjudicate, but propose a proportionate settlement.\u201d<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The Settlement Layer<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A settlement proposal should be evaluated differently from a factual finding.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A factual finding asks whether the system knows what happened. A settlement proposal asks whether a proposed outcome is fair, explainable, commercially rational and likely to resolve the dispute without escalation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For a settlement proposal (s), a practical resolution-confidence model may be expressed abstractly as:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\" C_S(s)=\\operatorname{Cal}_S \\left[ D_s^{\\alpha} A_s^{\\beta} F_s^{\\delta} P_s^{\\eta} Q_s^{\\theta} V_s^{\\kappa} \\right] \"><semantics><mrow><msub><mi>C<\/mi><mi>S<\/mi><\/msub><mo form=\"prefix\" stretchy=\"false\">(<\/mo><mi>s<\/mi><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo>=<\/mo><msub><mi>Cal<\/mi><mi>S<\/mi><\/msub><mo>\u2061<\/mo><mrow><mo fence=\"true\" form=\"prefix\">[<\/mo><msubsup><mi>D<\/mi><mi>s<\/mi><mi>\u03b1<\/mi><\/msubsup><msubsup><mi>A<\/mi><mi>s<\/mi><mi>\u03b2<\/mi><\/msubsup><msubsup><mi>F<\/mi><mi>s<\/mi><mi>\u03b4<\/mi><\/msubsup><msubsup><mi>P<\/mi><mi>s<\/mi><mi>\u03b7<\/mi><\/msubsup><msubsup><mi>Q<\/mi><mi>s<\/mi><mi>\u03b8<\/mi><\/msubsup><msubsup><mi>V<\/mi><mi>s<\/mi><mi>\u03ba<\/mi><\/msubsup><mo fence=\"true\" form=\"postfix\">]<\/mo><\/mrow><\/mrow><annotation encoding=\"application\/x-tex\"> C_S(s)=\\operatorname{Cal}_S \\left[ D_s^{\\alpha} A_s^{\\beta} F_s^{\\delta} P_s^{\\eta} Q_s^{\\theta} V_s^{\\kappa} \\right] <\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">where the terms represent <strong>evidentiary defensibility<\/strong>, <strong>acceptance probability<\/strong>, <strong>fairness<\/strong>, <strong>proportionality<\/strong>, <strong>explanation quality<\/strong> and <strong>external survivability<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The exact weights should not be guessed in a conference room. They should be calibrated through product data, expert validation, settlement outcomes and downstream challenge behavior.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The function <math data-latex=\"(\\operatorname{Cal}_S) \"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msub><mi>Cal<\/mi><mi>S<\/mi><\/msub><mo>\u2061<\/mo><mo form=\"postfix\" stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">(\\operatorname{Cal}_S) <\/annotation><\/semantics><\/math> is the calibration layer. It converts a raw model score into an operational probability. If <math data-latex=\"(C_S(s)=0.72)\"><semantics><mrow><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msub><mi>C<\/mi><mi>S<\/mi><\/msub><mo form=\"prefix\" stretchy=\"false\">(<\/mo><mi>s<\/mi><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo>=<\/mo><mn>0.72<\/mn><mo form=\"postfix\" stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">(C_S(s)=0.72)<\/annotation><\/semantics><\/math>, that should mean proposals of comparable type, risk and evidence profile have historically resolved successfully around seventy-two percent of the time.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Without calibration, the number is just random pick &#8211; Nothing more, nothing less.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is where many AI products become dangerous. They display confidence scores that look quantitative but are not tied to observed outcomes. A founder should be allergic to such systems. A CTO should reject them before they enter production.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Why 72 Percent Is Not a Magic Number<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">When people ask at what confidence level an AI should propose settlement without human review, they usually expect a fixed number.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That is the wrong framing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The threshold should be dynamic.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For a low-value, non-binding Stage 1 settlement proposal, the threshold may initially sit around 0.72. But this number is not a universal legal constant. It is a product-risk threshold.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A simplified decision threshold can be written as:<br><math data-latex=\"\\tau(D)=\\inf_{\\tau} \\left\\{ \\tau: P(Y_{\\text{fail}}=1 \\mid C_S \\geq \\tau,D) \\leq \\epsilon(D) \\right\\}\"><semantics><mrow><mi>\u03c4<\/mi><mo form=\"prefix\" stretchy=\"false\">(<\/mo><mi>D<\/mi><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo>=<\/mo><msub><mi>inf<\/mi><mi>\u03c4<\/mi><\/msub><mo>\u2061<\/mo><mrow><mo fence=\"true\" form=\"prefix\">{<\/mo><mi>\u03c4<\/mi><mo lspace=\"0.2222em\" rspace=\"0.2222em\">:<\/mo><mi>P<\/mi><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msub><mi>Y<\/mi><mtext>fail<\/mtext><\/msub><mo>=<\/mo><mn>1<\/mn><mo lspace=\"0.22em\" rspace=\"0.22em\" stretchy=\"false\">|<\/mo><msub><mi>C<\/mi><mi>S<\/mi><\/msub><mo>\u2265<\/mo><mi>\u03c4<\/mi><mo separator=\"true\">,<\/mo><mi>D<\/mi><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo>\u2264<\/mo><mi>\u03f5<\/mi><mo form=\"prefix\" stretchy=\"false\">(<\/mo><mi>D<\/mi><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo fence=\"true\" form=\"postfix\">}<\/mo><\/mrow><\/mrow><annotation encoding=\"application\/x-tex\">\\tau(D)=\\inf_{\\tau} \\left\\{ \\tau: P(Y_{\\text{fail}}=1 \\mid C_S \\geq \\tau,D) \\leq \\epsilon(D) \\right\\}<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This says: choose the lowest threshold at which the failure probability stays within the acceptable risk tolerance for that dispute class.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The case context D matters. A \u00a3300 dispute over missing inventory should not have the same automation threshold as a \u00a3300,000 dispute involving regulatory exposure, reputational consequences or complex contractual dependencies.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The tolerated failure rate <math data-latex=\"\\epsilon(D)\"><semantics><mrow><mi>\u03f5<\/mi><mo form=\"prefix\" stretchy=\"false\">(<\/mo><mi>D<\/mi><mo form=\"postfix\" stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">\\epsilon(D)<\/annotation><\/semantics><\/math> should shrink as claim value, legal risk, jurisdictional uncertainty or reputational exposure increase.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So the real answer is not:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>I would use 72 percent.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The real answer is:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>For low-value, non-binding Stage 1 settlement proposals, I would allow autonomous settlement recommendation when Resolution Confidence clears the dynamically calibrated threshold, which may initially be around 72 percent for the lowest-risk dispute classes.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That is a very different product philosophy.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It says the system is not hardcoded. It is governed.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What Not To Do<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The <strong>first mistake<\/strong>: is to use <strong>Truth Confidence<\/strong> as the settlement trigger.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>That sounds reasonable, but it fails in practice. Many disputes cannot cheaply reach high factual certainty. If the system waits for near-perfect truth, it will escalate too many cases and destroy the economic value of automation. If it proposes outcomes using low truth confidence, it will look arbitrary.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The <strong>second mistake<\/strong>: is to let the LLM invent the confidence score.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Language models are good at summarization, classification, extraction and explanation. They are not, by themselves, calibrated decision engines. They can support the reasoning workflow, but the confidence layer must be structured, measured and validated.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The <strong>third mistake<\/strong>: is to treat missing evidence as simple evidence against one party.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Absence can mean many things. The evidence may not have been generated. It may have been generated but not retained. It may be under the control of a non-party. It may be commercially disproportionate to obtain. It may have been withheld. Each of these has different meaning.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The <strong>fourth mistake<\/strong>: is manual threshold tuning.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Manual tuning may work in a pilot. It does not scale in production. If every dispute category requires bespoke human adjustment, the product is not an AI platform. It is a consulting workflow with a chatbot interface.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The <strong>fifth mistake<\/strong>: is to optimize only for acceptance.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>A bad settlement can be accepted if one party is fatigued, confused or commercially weaker. That is not success. A serious platform must also track execution, reopening, fairness perception, complaint rate and external survivability.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">The Business Implication<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For CEOs and founders, <strong>Resolution Confidence<\/strong> is not just a technical idea. It is a business architecture.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It determines when the product can automate, when it should ask for more evidence, when it should escalate to human review and when it should refuse to pretend certainty.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The operating model becomes:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\"\\text{If } C_S(s^\\star) \\geq \\tau(D), \\text{ propose settlement}\"><semantics><mrow><mtext>If&nbsp;<\/mtext><msub><mi>C<\/mi><mi>S<\/mi><\/msub><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msup><mi>s<\/mi><mo form=\"prefix\" stretchy=\"false\">\u22c6<\/mo><\/msup><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo>\u2265<\/mo><mi>\u03c4<\/mi><mo form=\"prefix\" stretchy=\"false\">(<\/mo><mi>D<\/mi><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo separator=\"true\">,<\/mo><mtext>&nbsp;propose&nbsp;settlement<\/mtext><\/mrow><annotation encoding=\"application\/x-tex\">\\text{If } C_S(s^\\star) \\geq \\tau(D), \\text{ propose settlement}<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\"\\text{If } C_S(s^\\star) < \\tau(D), \\text{ acquire the next most informative evidence}\"><semantics><mrow><mtext>If&nbsp;<\/mtext><msub><mi>C<\/mi><mi>S<\/mi><\/msub><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msup><mi>s<\/mi><mo form=\"prefix\" stretchy=\"false\">\u22c6<\/mo><\/msup><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo>&lt;<\/mo><mi>\u03c4<\/mi><mo form=\"prefix\" stretchy=\"false\">(<\/mo><mi>D<\/mi><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo separator=\"true\">,<\/mo><mtext>&nbsp;acquire&nbsp;the&nbsp;next&nbsp;most&nbsp;informative&nbsp;evidence<\/mtext><\/mrow><annotation encoding=\"application\/x-tex\">\\text{If } C_S(s^\\star) &lt; \\tau(D), \\text{ acquire the next most informative evidence}<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><math data-latex=\"\\text{If evidence acquisition has poor expected value, escalate}\"><semantics><mtext>If&nbsp;evidence&nbsp;acquisition&nbsp;has&nbsp;poor&nbsp;expected&nbsp;value,&nbsp;escalate<\/mtext><annotation encoding=\"application\/x-tex\">\\text{If evidence acquisition has poor expected value, escalate}<\/annotation><\/semantics><\/math><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is how a product remains low-cost without becoming low-trust.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The most valuable evidence request is not \u201csend more documents.\u201d It is targeted. Ask for courier pickup weight. Ask for delivery weight. Ask for seal records. Ask for serial-number scans. Ask for timestamped first-opening footage. Ask for ERP goods-received audit logs.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The system should estimate the expected value of each request before asking for it.<\/p>\n\n\n\n<div class=\"wp-block-math\"><math display=\"block\"><semantics><mrow><msup><mi>a<\/mi><mo>\u22c6<\/mo><\/msup><mo>=<\/mo><mrow><mi>arg<\/mi><mo>\u2061<\/mo><mspace width=\"0.1667em\"><\/mspace><\/mrow><munder><mi>max<\/mi><mi>a<\/mi><\/munder><mo>\u2061<\/mo><mrow><mo fence=\"true\" form=\"prefix\">[<\/mo><mi>\ud835\udd3c<\/mi><mrow><msub><mi>C<\/mi><mi>S<\/mi><\/msub><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msup><mi>s<\/mi><mo form=\"prefix\" stretchy=\"false\">\u22c6<\/mo><\/msup><mo lspace=\"0.22em\" rspace=\"0.22em\" stretchy=\"false\">|<\/mo><mi>a<\/mi><mo form=\"postfix\" stretchy=\"false\" lspace=\"0em\" rspace=\"0em\">)<\/mo><\/mrow><msub><mi>C<\/mi><mi>S<\/mi><\/msub><mo form=\"prefix\" stretchy=\"false\">(<\/mo><msup><mi>s<\/mi><mo form=\"prefix\" stretchy=\"false\">\u22c6<\/mo><\/msup><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mspace width=\"0.1667em\"><\/mspace><mi>Cost<\/mi><mo>\u2061<\/mo><mo form=\"prefix\" stretchy=\"false\">(<\/mo><mi>a<\/mi><mo form=\"postfix\" stretchy=\"false\">)<\/mo><mo fence=\"true\" form=\"postfix\">]<\/mo><\/mrow><\/mrow><annotation encoding=\"application\/x-tex\">a^\\star = \\arg\\max_a\n\\left[\n\\mathbb{E}{C_S(s^\\star \\mid a)}\n\nC_S(s^\\star)\n\n\\operatorname{Cost}(a)\n\\right]<\/annotation><\/semantics><\/math><\/div>\n\n\n\n<p class=\"wp-block-paragraph\">Again, the principle matters more than the formula. Evidence acquisition should be economically rational. If the evidence costs more to obtain than the dispute is worth, the platform should recognize that and shift toward proportionate resolution.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Why This Matters Beyond Legaltech<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The same idea applies across regulated AI.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In healthcare, a clinical decision-support system may not know the exact diagnosis but can still recommend the safest next diagnostic step. In anti-money laundering, a system may not prove criminal intent but can still assign a transaction to enhanced due diligence. In insurance, the system may not know precisely when damage occurred but can still propose a settlement based on evidence strength and policy risk. In smart contracts, the code may execute deterministically, but real-world disputes still require evidence-aware exception handling.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The common pattern is this:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>Truth is expensive. Resolution is economic. Trust requires knowing the difference.<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">That is where AI needs human expertise.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Not every uncertainty should be automated away. Some uncertainty must be represented, priced and resolved.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The Real Moat<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Anyone can connect an LLM to a document store and call it legal AI.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That is not the moat.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The moat is knowing that a dispatch manifest proves dispatch, not delivery. A delivery signature proves pallet transfer, not item-level verification. A medical note proves a recorded observation, not a ruled-out diagnosis. An invoice proves declared value, not economic legitimacy. A smart-contract event proves execution, not necessarily fairness.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The moat is in the evidence ontology, the confidence architecture, the calibration loop and the domain judgment embedded into the system.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">AI is a tool. In the hands of people who do not understand the domain, it creates fluent uncertainty. In the hands of experts, it becomes leverage.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For startups and AI businesses, the lesson is simple: do not build products that merely answer questions. Build systems that know when they have enough confidence to act, when they need better evidence and when the economically correct answer is not more truth, but better resolution.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That is the difference between an AI demo and an AI business.<\/p>\n\n\n\n<pre class=\"wp-block-verse\"><a href=\"https:\/\/gk.palem.in\/Contact.html?swcfpc=1\">Contact Me<\/a> if you are building solutions in the AI space and need domain expertise (Healthcare, FinTech, Legal Tech, Retail ...) to ensure correctness and product scalability.<\/pre>\n","protected":false},"excerpt":{"rendered":"<p>Can an AI be unsure about the truth and still be confident about a settlement? At first, that sounds contradictory. If the system does not know exactly what happened, how can it recommend what should happen next? But this is precisely how many real-world decisions work. A doctor may not know the exact biological pathway [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1085,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"advanced_seo_description":"","jetpack_seo_html_title":"","jetpack_seo_noindex":false,"jetpack_post_was_ever_published":false,"_cloudinary_featured_overwrite":false,"fifu_image_url":"https:\/\/live.staticflickr.com\/65535\/55382828039_40e8d0b2a5.jpg","fifu_image_alt":"","footnotes":""},"categories":[28,69,88],"tags":[26,58,86,84],"class_list":["post-1081","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai","category-blog","category-legal-tech","tag-artificial-intelligence","tag-legal","tag-legaltech","tag-startups"],"jetpack_featured_media_url":"https:\/\/live.staticflickr.com\/65535\/55382828039_40e8d0b2a5.jpg","jetpack-related-posts":[],"jetpack_shortlink":"https:\/\/wp.me\/pfLaRd-hr","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/posts\/1081","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/comments?post=1081"}],"version-history":[{"count":3,"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/posts\/1081\/revisions"}],"predecessor-version":[{"id":1084,"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/posts\/1081\/revisions\/1084"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/media\/1085"}],"wp:attachment":[{"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/media?parent=1081"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/categories?post=1081"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/gk.palem.in\/articles\/wp-json\/wp\/v2\/tags?post=1081"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}