Disaggregated Proof, Not Aggregate Claims: The Validation Record MBG Needs

MBG Watch · 2026-09-19

The premise

A national average can look clean while the program underneath it is failing in particular places.

That is the central risk for Makan Bergizi Gratis now. MBG can spend money, count meals, open kitchens, suspend non-compliant SPPG units, and still not yet know whether the program is improving nutrition, attendance, food safety, or equity for the groups whose outcomes matter most. The failure is not only technical. It is democratic: families, journalists, auditors, and local officials cannot inspect a claim if the claim arrives only as one aggregate number.

The narrow lesson from recent AI-evaluation research is useful here, but not because MBG needs an AI solution. In “Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation”, Kawano, Li, and Parker argue that consequential systems need evaluation by domains — task types, conversation types, or other subgroups — because average performance varies across the places where a system is actually used. Their paper is about AI systems, labeled samples, small-area estimation, and validation methods. MBG is not that system. But the public-accountability lesson crosses domains: when a program is large, uneven, and expensive to measure exhaustively, success must be validated at the level where failure can hide.

For MBG, that means proof by province, district, kitchen model, beneficiary route, and risk group — not only proof by national total.

What the public record already shows

MBG already has the kind of record that makes aggregate claims unsafe.

The budget record is large and moving fast. ANTARA reported BGN’s statement that MBG budget realization had reached Rp139.77 trillion, or 63.89 percent of the Rp218.77 trillion available ceiling, by 16 September 2026. That is an implementation signal. It is not yet an outcome signal. A rupiah absorbed by the system does not tell the public whether a toddler’s growth trajectory changed, whether a pregnant woman received a safer care route, whether a remote school was reached reliably, or whether the meal replaced food the household would otherwise have provided.

The safety record is also already disaggregated in practice, whether or not public reporting is. ANTARA reported that BGN suspended 1,276 SPPG units that had not met sanitary hygiene feasibility certificate, or SLHS, requirements, and proposed permanent closure for 654 kitchens considered critical if they did not meet requirements. The published breakdown matters: 821 were still in the SLHS registration process, 447 had registered but did not meet eligibility requirements, and eight had unconfirmed status. By region, the temporary suspensions were not evenly distributed: 239 in Sumatra, 927 in Java, and 110 outside Sumatra and Java.

That is already the shape of a validation problem. A national statement that “kitchens are being improved” is weaker than a record showing which kitchens, under which rule, in which region, with what beneficiary interruption, what replacement meal route, and what correction date.

This piece extends a chain of MBG Watch work rather than replacing it. “BGN Asks the Question It Should Have Asked First: Is 63 Million Beneficiaries Real?” treated the denominator as an accountability object. “Performance Is Not Proof” asked for an outcome ledger before school-benefit claims harden. “Earlier Care Is Not Earlier Proof” applied the same discipline to mothers, toddlers, and 3T nutrition. “Seen Without Being Watched” set the privacy boundary for beneficiary validation. “When the Score Becomes a Gate” warned that digital scoring should not become an unreviewable operating decision. The cautionary example is still “The Jayapura Stunting Claim”: a visible reminder that aggregate or poorly sourced health claims can travel faster than their evidentiary basis.

BGN’s own technical pages point in the same direction. Its guidance for pregnant women, breastfeeding mothers, and non-PAUD toddlers says accurate beneficiary data from BKKBN, updated periodically, is the basis for targeting; it also says monitoring and evaluation should be regular, findings should be followed up, and field feedback should be used for improvement. BGN’s Juknis page also lists a 2026 technical governance guideline for MBG in remote areas. Those are not outcome proofs, but they are institutional acknowledgements that MBG is not one population and one route. It is many routes, with different denominators and different risks.

The food-safety regulation record is similar. A Veritask summary of BGN Regulation 4/2026 describes requirements for risk-based inspection of raw materials, hygiene and cold-chain checks during transportation, SLHS obligations, meal samples stored below 5°C for 2 x 24 hours for investigation after suspected poisoning, organoleptic checks by receiving officers, recall steps, and reporting to health facilities and BGN. Again, the accountability unit is not the national average. It is the kitchen, route, handover point, sample, officer decision, incident report, and corrective action.

What AI evaluation contributes — narrowly

The useful import from disaggregated AI evaluation is not automation. It is discipline.

A public system that only reports averages can miss subgroup harm. This is the old statistical problem often illustrated by Simpson’s paradox: a combined trend can point in one direction while the subgroups point in another. MBG has several reasons to be vulnerable to that kind of mistake.

First, denominators can drift. “Beneficiaries reached” can mean registered beneficiaries, intended beneficiaries, meals delivered, meals accepted, unique people served, or people served on schedule. If those are mixed, a rising count can be real and still not mean what the public thinks it means.

Second, kitchen participation is not random. A kitchen that gets certified early, has better cold-chain access, or sits near a strong local health office may not represent a remote SPPG, a 3T route, or a temporary holiday-distribution route. If the easier kitchens generate the cleaner data, the program can overstate reliability.

Third, incident records can undercount harm. Food-safety events depend on recognition, reporting, sample preservation, laboratory access, and whether families trust the complaint channel. A district with fewer recorded incidents may be safer; it may also be less able to report.

Fourth, baseline mismatch can turn a weak claim into a strong-looking one. A stunting or attendance claim is not validated by showing that MBG exists in the same place as a later improvement. It needs a baseline, a comparison, an exposure definition, and a credible way to separate MBG from other changes: health services, household income, school policy, weather, migration, disease, or measurement changes.

Fifth, missing remote communities can create false confidence. If 3T areas are harder to reach and harder to measure, a national average can improve because better-measured places improved. That is exactly where disaggregated validation matters: not to punish low-performing places, but to keep the places with thin records from disappearing inside the mean.

The validation record MBG should publish

Before BGN, local governments, or political actors claim MBG has improved nutrition, attendance, food safety, or equity for a subgroup, the public record should contain a minimum validation packet.

It does not need to expose children. It does not need pregnancy-level files, child-level attendance histories, household addresses, or named health records. It can be published as privacy-protecting aggregates, with small-cell suppression where needed. But it should be specific enough that an auditor can reproduce the claim and a local official can see where action is needed.

For each claimed subgroup result, the record should show:

  1. The claim being made. For example: “MBG improved attendance among grade 1–3 pupils in district X,” “MBG reduced anemia risk among pregnant women in route Y,” or “certified SPPG kitchens had fewer food-safety interruptions than uncertified kitchens in province Z.”

  2. The numerator and denominator. Who is counted, who is excluded, and whether the count refers to unique people, delivered meals, accepted meals, operating days, kitchens, incidents, or confirmed cases.

  3. The subgroup definition. Province, district, 3T classification, beneficiary group, school level, pregnancy or toddler route, SPPG model, canteen pilot, remote-delivery route, or other category — with definitions stable enough to compare over time.

  4. The baseline. The pre-MBG level or earliest reliable measurement, with the date, source, and coverage limits.

  5. The comparison. A credible counterfactual where possible; otherwise a plainly weaker comparison labeled as such. Some claims may only support “service delivered,” not “outcome improved.”

  6. The uncertainty interval. Not a decorative confidence band, but a visible reminder that small subgroups can produce noisy results.

  7. The missing-data treatment. How many records were absent, delayed, unmatched, or discarded; whether missingness was higher in 3T districts, remote kitchens, PAUD/TK settings, pregnant-women routes, or incident-reporting channels.

  8. The provenance chain. Which office produced the source data, when it was last updated, how corrections are logged, and whether the number changed after publication.

  9. The action threshold. What level of risk or uncertainty triggers inspection, technical assistance, pause, replacement route, or public correction.

  10. The human-review point. Where an automated score, dashboard flag, or statistical estimate stops and accountable human judgment begins.

This last field matters. MBG Watch’s earlier pieces have warned against turning digital scores into gates. A validation record should not become a machine for silently excluding kitchens, children, pregnant women, or remote communities. It should make the basis for action visible, and it should preserve appeal, correction, and local explanation.

How to protect privacy while publishing proof

The least-harm path is subgroup-level proof without child-level surveillance.

For public reporting, BGN can aggregate by district, route type, school level, kitchen status, certification status, incident category, and beneficiary group. It can suppress small cells, delay publication for sensitive categories, use ranges instead of exact counts where re-identification risk is high, and publish correction logs without naming children, mothers, households, cadres, or complainants.

The record should answer accountability questions, not curiosity questions. The public needs to know whether the 3T route is reaching intended communities, whether pregnant-women and toddler routes have updated denominators, whether suspended kitchens had replacement service, whether food-safety incidents were investigated, and whether outcome claims have a baseline. The public does not need a searchable register of children, pregnancy status, attendance, home location, or health history.

A useful boundary is this: publish enough for a claim to be checked, not enough for a person to be tracked.

What should remain human

Some decisions should be aided by statistics but not delegated to them.

A kitchen with a failed sanitation record may need suspension. But the response should still ask whether children have a safe replacement meal route the next day. A remote district may show weak data completeness. That should trigger assistance and careful measurement, not quiet exclusion from the success story. A pregnancy or toddler route may have uncertain beneficiary counts. That should require reconciliation with BKKBN, health posts, and local cadres, not public exposure of individual records.

The point of disaggregated validation is not to rank communities. It is to keep weak evidence from becoming a strong claim, and to keep hidden failure from becoming someone else’s burden.

What remains uncertain

Several important facts are not visible enough from the public record.

BGN may already hold better internal data than the public can see. It may have kitchen-level correction histories, incident-investigation timelines, Posyandu or BKKBN reconciliation records, attendance linkages, or SPPG replacement-route logs that would answer many of these questions. If so, the gap is partly a publication gap: the evidence exists, but the public validation record does not.

It is also uncertain how safely SSGI, Posyandu, school attendance, BKKBN, BPOM, puskesmas, and SPPG records can be linked at aggregate level. Safe linkage is not automatic. It needs governance: purpose limitation, small-cell rules, audit logs, and clear separation between evaluation and eligibility.

The largest uncertainty is capacity. Disaggregated validation is not just a dashboard. It is a routine of updating denominators, preserving samples, logging corrections, reconciling local records, and explaining uncertainty. If local governments and SPPG operators are asked to maintain this record without staff, time, and training, the record will degrade into another compliance burden.

That is why the standard should be modest but firm. MBG does not need to prove everything at once. It does need to stop treating aggregate success as proof where subgroup failure could still be hiding.

The validation record is the bridge: public enough to inspect, aggregated enough to protect, and specific enough that a claim can be corrected before it hardens into policy.

Sources

  1. Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation — Narrow evaluation lesson: disaggregated domain-level validation and uncertainty, not an MBG AI solution
  2. Hingga September 2026, Anggaran Program MBG Terserap Rp139,77 Triliun — BGN-reported MBG budget realization of Rp139.77 trillion / 63.89 percent by 16 September 2026
  3. BGN tangguhkan 1.276 SPPG dan usulkan tutup 654 dapur — SPPG suspension and closure figures, SLHS status breakdown, and regional distribution
  4. Pedoman Teknis Distribusi Makanan dan Edukasi Gizi pada Program MBG Bagi Ibu Hamil, Ibu Menyusui, dan Anak Balita Non-PAUD — BGN guidance on accurate beneficiary data, periodic updating, monitoring and evaluation, and field feedback for 3B/non-PAUD toddler routes
  5. Pedoman Teknis Tata Kelola Program Makan Bergizi Gratis di Wilayah Terpencil — Existence of a 2026 BGN technical governance guideline for remote-area MBG implementation
  6. National Nutrition Agency Regulation Number 4 of 2026 Requires Sanitary Hygiene Eligibility Certificates for Nutrition Fulfillment Service Units and Imposes Operational Suspension Sanctions — Summary of BGN Regulation 4/2026 food-safety, SLHS, sample-storage, organoleptic testing, recall, and reporting requirements