Government employees are already using generative AI, but an OECD review of 14 countries found that many governments lack clear instructions for its use and consistent ways to determine whether pilot projects work, create risks, or should expand.
Key Takeaways
- Governments’ use of generative AI is moving faster than practical oversight.
- Some countries provide detailed checklists, training and testing tools, while others mainly set broad ethical principles.
- Few governments systematically measure whether AI pilots improve performance, create new risks, or are ready for wider use.
- The OECD proposed five areas for judging pilots: performance, public value, cost and feasibility, usability, and risk and compliance.
The Organisation for Economic Co-operation and Development (OECD) reviewed official AI guidance from 14 countries and found that generative AI in government is spreading faster than the practical rules needed to guide employees and test whether public-sector pilots work.
In a working paper published Monday, the OECD found that public employees are using generative AI, often without formal approval or clear oversight, faster than agencies are developing practical rules and consistent methods for evaluating pilot projects. The finding is based on official guidance from 14 countries, along with research and case studies on government AI pilots.
Guidance varies widely
The OECD found that government guidance ranges from broad principles to detailed instructions for running and evaluating AI projects. Some examples include:
- Australia sets general responsibilities. Employees remain accountable for AI-assisted decisions and should assume that information entered into public AI tools could become public.
- Canada gives employees practical dos and don’ts, including how to handle data, review AI output and begin with low-risk uses.
- Singapore provides project-level instructions for selecting uses, choosing data and measuring success.
The OECD said employees in agencies with only high-level principles may still lack clear answers about what data they can enter, who approves an experiment and who is responsible when a system produces a bad result. This can lead some agencies to avoid useful tests while others adopt inconsistent practices.
Governments rarely conduct meaningful evaluations
Very few governments meaningfully evaluate whether AI systems work or whether pilot projects are ready for wider use. Most rely on broad measures such as website traffic or user satisfaction, which do not show whether the AI:
- Improves the speed or quality of government work.
- Produces accurate and dependable results.
- Discriminates or violates existing rules.
- Delivers benefits that justify its costs and risks.
A few governments take a more structured approach:
- Singapore uses a central platform to track digital-service performance and analyze public feedback.
- The Netherlands created a government team to develop standard tests and tools for measuring the risks and benefits of generative AI.
- Italy’s draft guidance recommends specific measures covering accuracy, reliability, cost, usability, discrimination and regulatory compliance.
The OECD said these approaches remain exceptions. Most governments still lack the evidence needed to compare pilots, learn from failures and decide which projects to expand, change or stop.
OECD calls for shared rules and earlier evaluation
The OECD proposed evaluating government AI pilots in five areas:
- Performance: The quality of the data used and the results produced.
- Public value: Does the pilot deliver the intended benefit?
- Cost and feasibility: Is it affordable and does it work with existing operations?
- Usability: Are people able and willing to use it?
- Risk and compliance: Does the AI pilot manage risks and follow government rules?
The paper also called for:
- Shared guidance across agencies.
- Practical training for employees.
- Common data and testing tools.
- Clear measures, review schedules and criteria for changing, expanding or stopping a pilot, set before testing begins.
An earlier analysis cited in the paper found that only 6.3% of 1,050 government AI policy initiatives explicitly addressed testing or pilot projects.
The OECD systematically reviewed official guidance across 14 countries whose main purpose was to set rules, principles or instructions for generative AI in government. The review covered Australia, Canada, Finland, Ireland, Italy, Japan, New Zealand, Norway, Sweden, Switzerland, the Netherlands, the United Kingdom, the United States and Singapore. Also, it drew on academic research and documented case studies.

