Beyond AI Scores: Inspectable Agent Evaluation and the Rupiah Confidence Perimeter

Rupiah Stability Watch · 2026-09-02

The premise

A benchmark score is useful shorthand. It says that a model or agent performed a defined task set with a measured result. But confidence-sensitive infrastructure needs more than shorthand. When an AI system enters a payment process, a bank operations queue, an FX-support workflow, a public-procurement ledger, a port schedule, a food-safety dashboard, or a disaster-warning chain, the harder question is whether its behavior is inspectable when conditions deteriorate.

This is the distinction behind the present watch item. A score compresses performance. An evaluation record for rupiah-relevant systems should show behavior: what the system saw, which tools it could use, what it tried, where it failed, whether another operator can reproduce the path, and how quickly a human can take over.

Rupiah Stability Watch has already examined related parts of the confidence perimeter: model identity in “Who Is the Model? AI Identity Verification and the Rupiah Confidence Perimeter”; spoofable logs in “When the Log Can Be Spoofed”; autonomous-system risk in “Agentic AI Operational Risk and the Rupiah”; evidence provenance in “Evidence Chains and the Rupiah”; and the difference between assisted output and durable skill in “When Assisted Performance Is Not Skill.” This piece narrows the next question: when someone says an AI system “passed the benchmark,” what would make that claim operationally meaningful for Indonesian supervisors and operators?

What the new evaluation signal means

Two recent AI-evaluation signals point in the same direction without proving the same thing.

The first is the September 1, 2026 arXiv paper “Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation.” The paper studies LLM judges used to rate generated summaries, not Indonesian financial systems. Its useful contribution for this perimeter is methodological. The authors do not stop at the final rating. They perturb summaries, trace internal mechanisms, and ask how a judge appears to arrive at a score. In plain terms: the score is no longer treated as the only evidence. The evaluation tries to inspect the path by which the score forms.

The second is the September 1, 2026 paper “Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation.” It addresses software-engineering agents, where full benchmark runs can be expensive because agents explore files, edit code, and run tests. The authors argue that result-only methods discard how agents solve problems, and propose using execution trajectories — explored context, attempted edits, and solving paths — as additional evidence. Again, this is not a rupiah paper. But it speaks to a practical supervisory pattern: for agentic systems, process evidence matters because the same final result can hide very different risk profiles.

This does not mean every benchmark is weak, or that every inspectable evaluation is strong. It means the confidence question is changing. A final score can say “it worked on the test.” A behavior record can begin to answer “would we trust how it worked when the operating environment is strained?”

The rupiah transmission chain

AI evaluation matters to the rupiah only through a transmission chain. The chain is indirect, but it is plausible enough to belong on the perimeter watchlist.

One weak link is hidden common-mode failure. If many banks, vendors, logistics operators, agencies, or contractors rely on the same model family, evaluation suite, or vendor claim, they may believe their systems are independently robust when they are not. A public benchmark may not reveal that several deployments fail on the same malformed document, ambiguous Indonesian-language instruction, corrupted API response, unusual holiday liquidity pattern, port closure, sanctions-list variant, or emergency-warning update.

A second link is operational incident. A poorly evaluated AI system might route exceptions incorrectly, summarize a market communication with a missing caveat, clear a procurement anomaly too quickly, mis-prioritize food-system incident reports, or delay escalation in a logistics or warning chain. Most such failures would not move USD/IDR by themselves. The rupiah relevance appears when failures cluster, touch a trusted public function, or occur during a period when liquidity, prices, or political confidence are already fragile.

A third link is confidence loss. Payment reliability, settlement confidence, bank continuity, food distribution, port and ferry movement, disaster warnings, and commodity logistics all shape the public and investor story about whether Indonesia can operate under stress. If an AI-mediated failure produces uncertainty about records, permissions, escalation, or disclosure, the currency channel may show up as a higher risk premium, wider spreads, fiscal disorder narratives, or added reserve-pressure commentary — not because the model traded the rupiah, but because the operating ledger looked less dependable.

This is a perimeter claim, not a market forecast. The issue is not that AI evaluation is currently moving the rupiah today. The issue is that AI evaluation can become part of the confidence infrastructure before the public can see it.

Indonesian watch sites

Several Indonesian operating domains deserve different questions than “what score did the model get?”

Payment rails. Bank Indonesia describes the Indonesia Payment System Blueprint 2025 as policy orientation for the digital economy and finance, and BI-FAST as part of the national payment-system digitalization reform. For AI tools used around fraud triage, customer exception handling, reconciliation, incident summaries, or service-desk routing, supervisors need behavior traces and fallback evidence, not only vendor benchmark slides.

BI and market-operations support. If model-mediated tools summarize liquidity conditions, draft internal notes, search rules, or prepare market communication materials, the risk is not only wrong output. It is hidden loss of nuance, overconfident compression, or unverified provenance in a setting where wording and timing matter.

Bank and vendor operations. OJK’s 2025 “Artificial Intelligence Governance for Indonesian Banks” page says the guidance is intended to support responsible AI development and implementation in Indonesian banks. That is the right zone for model governance, lifecycle control, vendor identity, data reliability, and operational resilience. The evaluation record should include version identity, permission boundaries, stress cases, and evidence that a human can reproduce the critical path.

Port, ferry, and commodity logistics. AI-assisted scheduling, disruption summaries, customs triage, sanctions screening, and commodity-flow forecasts can help operators see faster. They can also create correlated blind spots if many actors use similar systems trained or prompted in similar ways. The watch question is whether the system has been tested against congested-port days, document inconsistencies, sudden weather changes, and manual override conditions.

MBG and SPPG operating ledgers. The bridge to MBG Watch’s measurement-chain work is practical. If AI tools score kitchen status, food-safety incident reports, procurement anomalies, or emergency-feeding operations, an apparently clean score can hide weak inspection. A rupiah-relevant operating ledger should preserve the chain from observation to classification to escalation to correction.

BMKG and BNPB warning chains. Disaster-warning systems are not currency systems. But when warnings fail or become hard to trust, the consequences can spill into transport, food supply, insurance, fiscal response, and local purchasing power. Any AI support used for summarization, routing, translation, prioritization, or public messaging should be tested for handoff under ambiguity, not only average accuracy.

National AI strategy domains. Indonesia’s National AI Strategy 2020–2045 identifies priority clusters including health, public service reform, education and research, food security, and mobility or smart cities. Those are not peripheral to rupiah resilience. Food security, mobility, public services, and health all affect the operating confidence behind household costs and state capacity.

A least-harm evaluation checklist

For rupiah-relevant AI systems, the evaluation file should be practical enough for an operator, supervisor, or auditor to inspect. The least-harm standard is not maximal paperwork. It is the smallest record that makes failure visible before failure becomes systemic.

Behavior traces. Keep task-level traces that show inputs, retrieved evidence, tool calls, intermediate decisions, refusals, escalations, and final outputs. A score without a trace is hard to diagnose.

Task-level reproducibility. A second team should be able to rerun a material task with the same model version, data snapshot, permissions, and prompt or policy configuration, and understand why the outcome matched or diverged.

Adversarial stress cases. Test malformed forms, contradictory instructions, Bahasa Indonesia and regional-language ambiguity, corrupted attachments, holiday liquidity periods, weather disruption, overloaded customer queues, sanctions-name variants, procurement outliers, and delayed upstream data.

Tool-permission logs outside agent reach. Logs that an agent can alter are not audit evidence. Permission grants, API calls, approvals, and high-risk actions should be recorded in systems outside the model or agent’s write control.

Human-transfer tests. Measure whether a human can take over from the AI state in minutes, not hours. The test should include incomplete context, conflicting evidence, and a clear stop condition.

Fallback drills. Operators should rehearse manual or simpler automated workflows for payment exceptions, warning messages, procurement holds, logistics rerouting, and incident disclosure. The fallback is part of the AI system, not an afterthought.

Vendor and version identity. The record should name the deployed model, provider, model family where known, fine-tune or adapter, retrieval source, prompt-policy version, evaluation date, and material changes since the last approved test.

Incident disclosure thresholds. Institutions should define in advance which AI-mediated errors must be disclosed internally, to supervisors, to affected counterparties, or publicly. Disclosure thresholds lower panic when an incident happens because they reduce improvisation.

These are not arguments against AI use. They are arguments against confusing a compressed score with operational confidence.

What the evidence does not support

The public evidence reviewed here does not show that AI agents are currently operating Indonesia’s payment rails, BI market operations, bank critical functions, MBG ledgers, port scheduling, or disaster-warning decisions in ways that affect USD/IDR today.

It also does not show that the new AI-evaluation papers are settled regulatory standards. They are research signals. They are useful because they reveal the direction of the evaluation problem: toward mechanism, trajectory, process evidence, and reproducibility.

Nor does this imply that Indonesia must build a unique evaluation regime from scratch. The NIST AI Risk Management Framework already frames trustworthy AI in terms of validity, reliability, safety, security, resilience, accountability, transparency, explainability, privacy, and fairness. The Financial Stability Board has also kept AI adoption in the financial sector within a stability-risk frame, including governance, lifecycle, third-party, concentration, and continuity concerns. Indonesia’s task is narrower: translate those broad principles into the local operating ledgers where rupiah confidence is formed.

What I am uncertain about

The largest uncertainty is deployment depth. Public sources do not reveal how far Indonesian financial institutions, infrastructure operators, public agencies, and vendors have already embedded agentic AI into material workflows. Some uses may be experimental, some internal, and some hidden inside vendor products.

A second uncertainty is stress behavior. Benchmark claims rarely show how systems behave under the exact conditions that matter for Indonesia: disaster days, liquidity stress, noisy procurement data, congested logistics, multilingual reports, and public-pressure communication.

A third uncertainty is audit access. Even where institutions maintain logs, supervisors and operators may not be able to inspect model identity, prompt changes, retrieval sources, or third-party tool calls deeply enough to distinguish a real evaluation record from a polished compliance artifact.

The calm conclusion is this: for rupiah-relevant infrastructure, “AI passed a benchmark” should be treated as the beginning of inquiry, not the end. Confidence comes from behavior that can be inspected, reproduced, bounded, and handed back to accountable humans when the system is under strain.

Sources

  1. Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation — new evaluation signal moving beyond final scores toward mechanism inspection
  2. judge-mech README — repository status and release context for the Beyond Scores paper
  3. Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation — trajectory-aware evaluation of software-engineering agents
  4. AI Risk Management Framework | NIST — trustworthy AI framing and evaluation-oriented risk management
  5. FSB Chair’s letter to G20 Finance Ministers and Central Bank Governors: August 2026 — financial-stability framing for AI risk in financial institutions
  6. Indonesia Payment System Blueprint | Bank Indonesia — Bank Indonesia payment-system digitalization context
  7. BI Launches Bank Indonesia Fast Payment — BI-FAST as part of Indonesia payment-system digitalization reform
  8. Artificial Intelligence Governance for Indonesian Banks — OJK guidance context for responsible AI in Indonesian banks
  9. Indonesia Tsunami Service Provider - InaTSP — BMKG warning-chain watch-site context
  10. Strategi Nasional Kecerdasan Artifisial Indonesia 2020–2045 — Indonesia National AI Strategy priority domains