Dataset

Aisberg Telegram News UA

A de-identified dataset of Ukrainian Telegram news posts and their comment threads, downloaded 8,165 times to date and around 3,000 times a month. Free, no registration, released under CC BY 4.0, including for commercial use, as long as you credit us.

8,165downloads
2,461cluster files
11monthly snapshots
CC BY 4.0licence

What's inside

The dataset ships in two shapes, because researchers ask for two different things. One is the raw stream; the other is our analysis of it.

Monthly snapshots: data/MM_YYYY.jsonl

The raw stream, one file per calendar month, JSONL/UTF-8.

  • Posts: channel name, anonymized text, publication date, aggregated reactions (emoji to count)
  • Comments: user pseudonym, comment text, timestamp in epoch milliseconds

Use this for corpus work: language modelling, sentiment baselines, engagement statistics.

Per-cluster analysis: clusters/cl_*.json

One file per published report. It is the same JSON that powers the report dashboard.

  • Grouped posts and the generated summary
  • Detected manipulation techniques and signals
  • View / reaction time series and comment sentiment
  • The full fact-checking evidence chain, with sources and verdict

Use this for narrative-level work: spread dynamics, cross-channel timing, verification research.

Coverage

Monthly snapshots currently published. The most recent snapshot is July 2026; the August file is published at the start of September, once the month is complete. October 2025 is the one gap: collection was down that month and no data exists to backfill.

Monthly snapshot availability, as of 26 August 2026
PeriodFileStatus
August 2025 · collection startdata/08_2025.jsonlpublished
November 2025data/11_2025.jsonlpublished
December 2025data/12_2025.jsonlpublished
January 2026data/01_2026.jsonlpublished
February 2026data/02_2026.jsonlpublished
March 2026data/03_2026.jsonlpublished
April 2026data/04_2026.jsonlpublished
May 2026data/05_2026.jsonlpublished
June 2026data/06_2026.jsonlpublished
July 2026data/07_2026.jsonlpublished
August 2026data/08_2026.jsonlearly September

The table shows the first collection month and the most recent ones; the full file list is on Hugging Face. Months before March 2026 contain some collection-downtime days; from March 2026 onward coverage is complete and continuous.

Per-cluster files are uploaded continuously, so a new one appears with every published report. Languages present: Ukrainian, Russian, English.

How the data is de-identified

We publish discussion data from a country at war, so anonymization is not a formality. What we do:

  • User IDs and usernames are replaced with salted pseudonyms (user_<hash>), stable within the dataset so conversation structure survives
  • Emails, phone numbers, mentions and URLs are replaced with the tokens [email], [phone], [mention], [url]
  • Dates in the date field are normalized to day-level granularity
  • Internal Telegram IDs are removed; only channel names are preserved

Channel names are deliberately kept: without them, cross-channel narrative research becomes impossible, because the whole question is who published first and who amplified whom. Channels are public broadcasters, not private individuals.

A limitation worth stating plainly: pseudonymization is not anonymity. A determined party who already holds the original public messages can re-link some comments by matching text. Treat this as de-identified research data, not as a privacy guarantee for the people in it.

How to cite

The licence is CC BY 4.0. Use it freely, including commercially, with attribution.

Plain

Aisberg Public Organization. Aisberg Telegram News UA (anonymized) [Data set].
Hugging Face. https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset

BibTeX

@misc{aisberg_telegram_news_ua,
  title        = {Aisberg Telegram News UA (anonymized)},
  author       = {{Aisberg Public Organization}},
  howpublished = {\url{https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset}},
  note         = {Available at \url{https://aisberg.live/data/}},
  license      = {CC-BY-4.0}
}

If you publish something built on this data, we'd like to know: info@aisberg.live. We keep a list.

Need more than the public snapshots?

Our working corpus is larger and finer-grained than what we publish: full snapshot series, per-post view and reaction curves, comment threading, channel-level aggregates, and analysis outputs that don't fit the monthly format. The public release is deliberately conservative next to it, with day-level dates and a fixed set of fields.

We are a registered Ukrainian non-profit (ЄДРПОУ 46034329). If you're a researcher, journalist, fact-checking organization or NGO and the public files don't cover your question, write to us and describe what you need. We can usually prepare a custom extract for a specific date range, a set of channels, a narrative, or a particular field set. For non-commercial research we do it for free.

Email info@aisberg.live

Please include: what you're researching, the period and channels you need, the fields that matter, and how the result will be published. It helps us say yes faster.

What next

News gets cleaner when somebody checks it

News clusters, fact-check verdicts and a link to the dashboard go out on the Telegram channel every hour. Reading it is free and needs no account. And if you want more of this work to happen, you can join it.

Subscribe to the channel Become a volunteer