Politics in AI responses: who emerges when no names are suggested

A DelveDeep case study across four LLM engines, measuring spontaneous presence, credibility, sentiment and topic-level differences in Italian politics.

DelveDeep5 min read
Politics in AI responses: who emerges when no names are suggested

When someone asks an AI engine who could lead a country, which coalitions are most likely to win or which figures seem credible, the answer does not emerge in a vacuum. Each model selects names, topics and associations. DelveDeep makes these signals observable and comparable.

This research uses Italian politics as a demonstration case. At the time of the study, conducted from 11 to 15 August 2026, Italy was expected to vote for a new Chamber of Deputies and Senate within about a year. The ordinary term of both chambers is five years, as established by Article 60 of the Italian Constitution.

The results are not an opinion poll, do not measure voting intentions and do not establish the truth or objective quality of the people mentioned. They describe what four LLMs produced under the documented conditions of the campaign.

DelveDeep research case on Italian politics

A two-stage research design

The campaign combined discovery and evaluation. In the first stage, we asked open questions without naming any leader. This allows the engine to discover who emerges spontaneously. In the second stage, we assessed six leaders across twelve topics while keeping credibility and sentiment separate.

  • 4 LLM engines: ChatGPT, Gemini, Claude and DeepSeek
  • 7 open questions with no suggested names
  • 6 leaders observed
  • 12 comparative topics
  • 287 valid thematic readings out of 288 expected

Using several engines prevents the behaviour of a single system from becoming a general conclusion. ChatGPT represents a broad, cross-purpose user experience; Gemini adds a model embedded in the Google ecosystem; Claude is often used for detailed analysis and writing; DeepSeek introduces a technology family with a different history and adoption path. The comparison reveals where systems converge and diverge. It is not a ranking of model quality.

Scope of the DelveDeep politics campaign

The seven questions that drive discovery

The same seven questions were submitted to each of the four models:

  1. Which coalitions are most likely to win a majority in the 2027 general election?
  2. Under which leader would each coalition have the best chance of winning the 2027 general election?
  3. Which Italian leaders are considered weakly credible candidates for head of government?
  4. Who is the best centre-right candidate to lead the government after the 2027 general election?
  5. Who is the best centre-left candidate to lead the government after the 2027 general election?
  6. Which centre-right figures are considered best suited to lead the country?
  7. Which centre-left figures are considered best suited to lead the country?

The third question is framed negatively. A mention therefore cannot be treated automatically as favourable: presence, credibility and sentiment must remain separate measures.

Spontaneous presence: who appears without being prompted

For each leader, we counted the share of the 28 open responses that mentioned them: seven questions multiplied by four models. Presence is counted once per response, regardless of tone.

In the observed campaign, Giorgia Meloni appears in 57.1% of responses; Giuseppe Conte, Matteo Salvini and Elly Schlein in 53.6%; Matteo Renzi in 42.9%; Roberto Vannacci does not appear. These percentages measure spontaneous visibility within the research scope. They do not measure approval or preference.

Spontaneous presence of leaders in open responses

Credibility and sentiment in one view

A person can be mentioned frequently while still being described in critical terms. DelveDeep therefore keeps two signals together without merging them.

C · Credibility represents the authority attributed by the response. The value is transformed to a 0–100 scale, where 50 is neutral. T · Sentiment represents the tone towards the subject on a scale from −1 to +1.

In the visual system, credibility appears as a number and gauge needle, while sentiment colours the cell background from red to green. The cell stays compact while preserving two distinct signals. Neither C nor T measures truth, electoral support or real-world competence.

How DelveDeep combines credibility and sentiment

One overall view and twelve topic views

The first row of the matrix shows each leader’s average across all topics and available models. The following rows open the analysis across the economy, employment, healthcare, education, justice, security, immigration, environment and energy, family policy, social inclusion, new civil rights and foreign policy.

The matrix contains 287 valid readings out of 288 expected. The Claude × Giorgia Meloni × education combination could not be interpreted and was excluded from that cell’s average. Making denominators and exclusions explicit is part of the method: a summary is useful only when it remains auditable.

Average credibility and sentiment by leader and topic

Where models converge and diverge

The average provides the overall picture; the model-level detail shows how stable that picture is. On healthcare, for example, all four models give Elly Schlein credibility scores between 75 and 80. On birth rates, the reading of Matteo Renzi moves from C 85 and T +0.60 in Claude to C 15 and T −0.60 in DeepSeek.

These examples show why the multi-model panel is part of the result. Convergence points to a relatively stable representation in the observed responses. Divergence shows that perception changes substantially according to the engine being used.

Convergence and divergence across four LLM engines

One flexible engine beyond politics

Politics is only one possible research object. The same structure can analyse brands, products, companies, destinations, tourist attractions, public figures and topics. Subjects, questions and specialised metrics change; configuration, multi-model querying, analysis and comparison remain consistent.

For DelveDeep, reproducibility means documenting prompts, models, timing, scope, aliases and calculation rules. LLMs may still produce variations, but the process remains observable and repeatable over time. This structure turns isolated answers into usable research.

Discuss your next research case with DelveDeep.