PoliLoom

Where loose threads become linked data

  • Who are we?
  • What will we talk about?
  • Why does it matter?

Who are we?

We're representing OpenSanctions

Brenna Maeve

Mastermind

Johan Schuijt

AI/Data Engineer

The problem

Wikidata is the backbone of civic knowledge

But politician data is scattered and lagging

Manual collection is tedious

We want fast, verifiable edits

The data is out there

But scattered across unstructured documents

Wikipedia has some

And there are government publications

Lots of sources

All structured differently

Our solution

Simple infrastructure that proposes edits

Principles

  1. Humans decide, models propose
  2. Citations or it does not land
  3. Transparent pipeline and public artifacts

PoliLoom

Automate the boring bits

Let the model drown in information

So we can focus on verifying facts

Show proposed and current side by side

Automatically cite archived source

Make contribution enjoyable

It can feel like a game

Why OpenSanctions cares

OpenSanctions ❤️ Open Data

We want to see the commons grow

Better data helps everyone who uses Wikidata

Clearer context downstream

Better positions, terms, and affiliations reduce false matches

A healthier commons benefits everyone who uses open data, including us

Why the wiki community wins

Less grunt work

Suggestions come pre-sourced and pre-parsed

Verifiability by default

Every statement carries archived evidence

Coverage where it is thin

Local councils, committees, short stints

What good looks like

Politicians with office terms,

start and end dates,

and sources

Faster time to update after elections or reshuffles

Fewer ambiguous matches in downstream reuse

Community safeguards

Archived sources shown to check claims

Full change logs so edits are traceable

We want to contribute back

So how does it work?

Architecture

Ingest API + Evaluation GUI

Ingest API

  • Ingests wikidata dumps
  • Archives web sources
  • Extracts statements

Evaluation GUI

  • User-friendly evaluation
  • Metawiki OAuth system
  • Highlighting of archived sources

Extracts statements?

Well... extract, then reconcile

  1. Extract strings from documents
  2. Search for matching entities
  3. LLM reconciliation

1. Extract strings

LLM identifies statements:

"Member of Saarbrücken City Council"

2. Search for matches

Semantic search identifies matches:

"Mitglied der Regionalversammlung Saarbrücken"

3. LLM reconciliation

Source strings + Matching entity descriptions

Structured data API: "Q99775411"

Great! But statements from where?

Currently were using existing wikipedia links

But adding support for government portals

Generated webscrapers

For a curated list of pages

What we're figuring out currently

Basically filtering and sorting

Who do you look at first?

Newest? Least known statements?

How to make bite-sized chunks?

"campaigns" / tiers of government

Why does this matter?

Concrete wins for editors and reusers

Better data for research and journalism

A healthier commons that we all rely on

Better integration in civic tech and compliance

Complete picture of global governance

Your role

Evaluate suggestions in your language and jurisdiction

Teach the system with better examples and edge cases

Tell us what to prioritize next

Get involved!

Join the test cohort

Try Poliloom today at loom.everypolitician.org

Thank you!

opensanctions.github.io/wikidatacon2025