Making Government AI Work: Monitoring the System, the Decision and the Outcome
Artificial intelligence is beginning to move from pilot projects and chatbots into the everyday machinery of urban governance. Chandigarh Municipal Corporation’s 2026–27 budget reflects this shift. Of its total ₹1,712 crore budget, ₹12 crore has for the first time been allocated for a package of information-technology upgrades, including AI-based systems, digital file tracking and online billing and payment systems. The city has also previously introduced BIRBAL, an AI-enabled chatbot that helps residents navigate services provided by the Chandigarh Administration and Municipal Corporation.
These tools can help municipal governments respond to complaints, identify service gaps and manage resources more effectively. But the shift is not simply from manual to digital administration. As technology begins to prioritise complaints, flag irregularities or influence administrative decisions, the performance of the system itself becomes part of service delivery.
As an earlier ResGov blog argued, governments already report scheme outputs while providing far less information on whether the digital platforms through which services are delivered are reliable. AI adds another layer to this question. Even when a system is technically functioning, governments still need to know whether its outputs are accurate, how officials are acting on them and what happens when they are wrong.
This blog argues that monitoring must follow this entire chain—from the functioning of the technology, to its use in administrative decisions, to its effect on services and citizens. Drawing on lessons from earlier welfare technologies and the Government of India’s AI Governance Guidelines, it sets out the questions that departments should build into procurement, implementation and review.
The decision matters more than the label
“AI” is increasingly used as a broad label for technologies that work in very different ways. Rule-based software, database matching, biometric authentication, predictive models and generative-AI chatbots work differently and thus do not require identical safeguards.
What matters for oversight is thus what administrative function it performs. A chatbot that provides information carries relatively limited consequences when it produces an incorrect answer, particularly where citizens can easily reach a human official. An error in a sanitation-routing system could affect service delivery in a neighbourhood. But where a system flags households for investigation or influences access to an entitlement, the consequences are far more serious.
The Government of India’s AI Governance Guidelines recognise this distinction. They call for risk assessment suited to the Indian context, human oversight at critical decision points and accountability proportionate to the function performed and the potential harm involved.
Monitoring should therefore be proportionate to the nature of the decision and the harm an error could cause. The greater the potential impact on rights, entitlements or access to essential services, the stronger the requirements for testing, documentation, human review and appeal.
What earlier welfare systems teach us
One of the most common uses of data-driven systems in welfare administration is to identify duplicate or ineligible beneficiaries and reduce leakage. These are legitimate objectives, particularly where governments are administering large programmes with limited staff and fragmented records.
Telangana’s Samagra Vedika platform was introduced partly for this purpose, bringing together information from more than 30 government databases to support eligibility checks across welfare schemes. A state government presentation on a 2016 ration-card pilot in Hyderabad reported that around 100,000 cards had initially been removed. Of the approximately 19,000 people who subsequently reapplied, around 14,000 had their cards restored following verification. The presentation reported a net removal of approximately 86,000 cards and estimated savings of ₹4.6 crore per month.
A joint investigation published in 2024, however, documented cases in which incorrect data matching and inferences may have contributed to eligible households being flagged as ineligible. Between 2014 and 2019, more than 1.86 million food-security cards were cancelled and 142,086 fresh applications rejected without individual notice, although these decisions cannot all be attributed to Samagra Vedika. Following a court-directed re-verification process, 15,471 of the 205,734 applications processed by July 2022 were approved—suggesting that at least some of the earlier exclusions had been incorrect.
Both sets of numbers matter. The cards identified and estimated savings tell us whether the system may be reducing errors of inclusion. Restored cards and successful appeals tell us whether it is also creating errors of exclusion. A monitoring system that reports only the first is incomplete.
Experience with Aadhaar-enabled welfare delivery points to a similar distinction between a safeguard on paper and its operation in practice. Central government guidance provides for alternative methods of identification and advises that genuine beneficiaries should not be denied foodgrains because of biometric-authentication failure, missing Aadhaar seeding, network problems or other technical issues.
But the existence of a fallback provision does not tell a department how often authentication fails, whether the alternative is readily available or how long a beneficiary must wait for the error to be resolved. These are questions of implementation—and therefore of monitoring.
What should governments monitor?
The India AI Governance Guidelines already provide broad direction. They recommend risk classification, human oversight, audit trails, incident reporting, grievance redress and proportionate responsibility. The next step is translating these principles into a small set of management questions that departments, municipalities and frontline officials can answer routinely.
1. What problem is the system meant to solve?
Monitoring has to begin before procurement. Departments should identify the administrative problem, establish a baseline and specify what improvement the technology is expected to produce. This means recording existing processing times, service coverage, unresolved complaints, error rates and administrative costs before a new system is introduced.
A sanitation-routing tool, for example, should not be judged primarily by the number of routes it generates. The relevant questions are whether collection becomes more regular, fewer locations are missed, complaints fall and differences across wards narrow.
The procurement document should therefore contain outcome indicators alongside technical specifications and software deliverables. Without a baseline, governments may be able to confirm that a system has been installed, but not whether it has improved the service it was purchased to support.
2. Is the system working—and for whom?
The first layer of monitoring concerns the technology itself. The appropriate indicators will depend on the system. For a chatbot, they might include response accuracy, unanswered questions, language performance and successful transfer to a human official. For a system used to flag possible fraud, they should include how many flags are subsequently confirmed, how many are overturned and what types of data errors recur.
Performance should also be examined across the locations and groups most relevant to the service. An average accuracy rate may conceal weaker performance in areas with poor connectivity, across particular languages, or for people whose names and addresses are recorded differently across databases. This does not require every system to apply a standard demographic checklist. It requires departments to identify the likely pathways through which the particular technology could work unevenly, subject to lawful and responsible use of data.
Monitoring must also continue after deployment. Data change, administrative rules are revised and software is updated. A system that performed adequately during a pilot may not necessarily do so when expanded to a larger population or used in a different context.
3. How does an output become an administrative decision?
Technical performance is only one part of the story. Governments must also examine how the system is used inside the administration.
Is its output merely one piece of information available to an official? Does it determine which cases receive greater scrutiny? Or does it automatically trigger a rejection, suspension or closure? The same technology can carry very different consequences depending on how it is inserted into an administrative process.
Where human review is required, departments should assess whether it is meaningful. The relevant question is not simply whether an official clicked an approval button. It is whether the official could understand the basis of the recommendation, had access to other relevant information, and possessed both the authority and the time to disagree with it.
Departments should consequently maintain reviewable records of the relevant data, system recommendation, final administrative decision and any override. Changes to the model or decision rules should also be recorded. This information need not all be publicly released, particularly where it contains personal or security-sensitive data. But an authorised reviewer should be able to reconstruct how a consequential decision was reached.
The AI Governance Guidelines similarly emphasise human oversight, audit trails and the ability to review or override AI outputs before they cause harm.
4. What happens when the system fails?
Government systems will make errors. The important question is whether those errors are visible and capable of being corrected.
Citizens should know when an automated system has materially influenced a consequential decision, where they can seek assistance and what information is required to challenge it. Essential services should have workable alternatives where a technological channel is unavailable or produces an evidently incorrect result.
Departments should monitor more than the number of grievances closed. They should track the nature of errors, time taken to restore a service or entitlement, frequency with which decisions are overturned and whether the same failure is recurring.
Incident logs can serve an important management function here. They allow officials and technology teams to distinguish isolated mistakes from systematic problems with data, software or administrative processes. Where appropriate, public-facing status information can also help citizens and frontline staff understand whether the problem lies with their individual case or with the platform itself.
The Government of India’s guidelines recommend both an AI-incident mechanism and accessible grievance channels, with the feedback generated through complaints feeding back into system improvements.
5. Who is responsible—and has monitoring been funded?
Government departments often procure technology from private vendors, but they cannot outsource responsibility for a public decision.
Contracts should clarify who will maintain system records, investigate incidents and correct errors. They should also give the government access to the documentation and logs required to evaluate performance; specify when significant changes to a system must be disclosed; and establish rights relating to testing, data portability and exit from the contract.
These are not solely legal or technical questions. Officials must be able to manage the procurement and understand the system well enough to question vendor claims. The AI Governance Guidelines themselves identify limited technical capacity among public officials as a constraint and recommend training for officials responsible for procurement, risk management and oversight.
There is also a public-finance implication. An allocation for government AI should not cover only software, hardware and vendor fees. It must provide for data cleaning, staff training, cybersecurity, maintenance, grievance redress, periodic testing and evaluation.
Not every application will require an expensive independent audit. But systems that materially affect entitlements, enforcement or essential services should be periodically assessed by reviewers who are not responsible for building or operating them. The depth and frequency of review can be proportionate to the consequences of an error.
Monitoring as part of implementation
This is not an argument for creating a compliance-heavy regime around every government chatbot or digital tool. It is an argument for treating monitoring as part of implementation rather than as an additional ethics exercise undertaken after deployment.
For a low-risk chatbot, monitoring may focus on response accuracy, language coverage, uptime and successful handover to a human official. For a higher-risk eligibility system, it must also capture false positives, human overrides, notices, appeals, fallback use and restoration of benefits.
As Chandigarh and other cities invest in AI-enabled administration, they have an opportunity to define these measures alongside the technology itself. The measure of success should not be how many AI systems have been launched or how many decisions have been automated. It should be whether services improve, errors become easier to identify, and citizens and officials are able to correct them when they occur.

