Meta’s Oversight Board found that leading AI models were substantially less willing to criticize governments in countries where laws restrict criticism of political leaders than in countries with stronger protections for political speech.
Key Takeaways
- AL Models rejected 34% of prompts that criticize governments and political leaders in countries where such criticism can carry legal penalties, compared with 14% in countries with stronger protections for political speech.
- China and Thailand had the highest refusal rates, at 45% and 43%, while the United Kingdom and the United States had the lowest, at 8% and 9%, respectively. Taiwan was an outlier, with a refusal rate of 24%.
- The gap emerged when models were asked to create political content, such as protest flyers and poems. By contrast, models rejected opinion-based political questions at the same rate, 41%, regardless of the country involved.
- Model behavior varied widely. Gemini 3 Flash and Grok 4 Fast did not refuse any requests for protest flyers or poems, while other systems frequently declined to generate political content.
- The Meta Oversight Board called on AI companies to incorporate human rights reviews into model testing and to provide clearer explanations when systems reject politically sensitive requests.
The Meta Oversight Board, an independent body that makes binding content decisions for Facebook, Instagram, and Threads, published its first study examining how major AI models respond to political speech.
The report tested 10 models from Anthropic, DeepSeek, Google, Meta, OpenAI, and xAI by asking them to generate protest flyers, political poems, and opinions about governments and political leaders. Researchers found that the models were much more likely to reject criticism involving countries where speech against political leaders can be punished by law.
Models refused criticism of some governments more often than others
Across all models, requests involving China, Saudi Arabia, Thailand, Cambodia, and Turkey were rejected 34% of the time, compared with 14% for prompts involving the United States, Taiwan, Japan, Chile, and the United Kingdom.
China produced the highest refusal rate, at 45%, followed by Thailand at 43%. At the other end of the spectrum, the United Kingdom and the United States recorded refusal rates of 8% and 9%, respectively. Taiwan stood apart from the other countries in its group, with a refusal rate of 24%.
The gap appeared in content creation, not opinion questions
The differences were concentrated in requests to create political materials. Models were more likely to refuse requests for protest flyers and poems criticizing governments and political leaders, particularly in countries with restrictions on political speech.
However, when researchers asked the models for political opinions, such as whether a government should be supported or protested, refusal rates were identical across both groups of countries at 41%.
Some models rejected far more requests than others
Performance varied significantly across the 10 systems tested. Gemini 3 Flash and Grok 4 Fast did not refuse any protest-flyer or poem prompts. Other systems were far more restrictive.
Meta’s Llama 4 Maverick, for example, refused every protest-flyer request involving Chinese President Xi Jinping, Thailand’s King Vajiralongkorn, and Cambodia’s King Sihamoni, often citing legal, safety, or political concerns.
Researchers also found that models frequently gave inconsistent explanations for their refusals. Some cited local laws or safety concerns, while others invoked policies that appeared to be applied unevenly. The report cautioned that these explanations should not be treated as reliable accounts of how the systems actually work.
Recommendations
The Meta Oversight Board said AI companies should incorporate human-rights reviews into model development, explain more clearly why systems reject politically sensitive requests, and disclose whether restrictions stem from company policies, legal obligations, or government demands.
Methodology
The study examined 10 models from Anthropic, DeepSeek, Google, Meta, OpenAI, and xAI. The tests ran in March 2026 using seven prompts that asked models to generate protest flyers and satirical poems, express opinions about political leaders and institutions, and respond to questions involving political violence.
Each prompt was tested across 10 jurisdictions and four leaders or institutions in each country. The researchers compared China, Saudi Arabia, Thailand, Cambodia, and Turkey, where some forms of political criticism can carry legal penalties, with the United States, Taiwan, Japan, Chile, and the United Kingdom, which have stronger protections for political speech.

