Aisberg Telegram News UA
A de-identified dataset of Ukrainian Telegram news posts and their comment threads, downloaded 8,165 times to date and around 3,000 times a month. Free, no registration, released under CC BY 4.0, including for commercial use, as long as you credit us.
What's inside
The dataset ships in two shapes, because researchers ask for two different things. One is the raw stream; the other is our analysis of it.
Monthly snapshots: data/MM_YYYY.jsonl
The raw stream, one file per calendar month, JSONL/UTF-8.
- Posts: channel name, anonymized text, publication date, aggregated reactions (emoji to count)
- Comments: user pseudonym, comment text, timestamp in epoch milliseconds
Use this for corpus work: language modelling, sentiment baselines, engagement statistics.
Per-cluster analysis: clusters/cl_*.json
One file per published report. It is the same JSON that powers the report dashboard.
- Grouped posts and the generated summary
- Detected manipulation techniques and signals
- View / reaction time series and comment sentiment
- The full fact-checking evidence chain, with sources and verdict
Use this for narrative-level work: spread dynamics, cross-channel timing, verification research.
Coverage
Monthly snapshots currently published. The most recent snapshot is July 2026; the August file is published at the start of September, once the month is complete. October 2025 is the one gap: collection was down that month and no data exists to backfill.
| Period | File | Status |
|---|---|---|
| August 2025 · collection start | data/08_2025.jsonl | published |
| ⋯ | ||
| November 2025 | data/11_2025.jsonl | published |
| December 2025 | data/12_2025.jsonl | published |
| January 2026 | data/01_2026.jsonl | published |
| February 2026 | data/02_2026.jsonl | published |
| March 2026 | data/03_2026.jsonl | published |
| April 2026 | data/04_2026.jsonl | published |
| May 2026 | data/05_2026.jsonl | published |
| June 2026 | data/06_2026.jsonl | published |
| July 2026 | data/07_2026.jsonl | published |
| August 2026 | data/08_2026.jsonl | early September |
The table shows the first collection month and the most recent ones; the full file list is on Hugging Face. Months before March 2026 contain some collection-downtime days; from March 2026 onward coverage is complete and continuous.
Per-cluster files are uploaded continuously, so a new one appears with every published report. Languages present: Ukrainian, Russian, English.
How the data is de-identified
We publish discussion data from a country at war, so anonymization is not a formality. What we do:
- User IDs and usernames are replaced with salted pseudonyms (
user_<hash>), stable within the dataset so conversation structure survives - Emails, phone numbers, mentions and URLs are replaced with the tokens
[email],[phone],[mention],[url] - Dates in the
datefield are normalized to day-level granularity - Internal Telegram IDs are removed; only channel names are preserved
Channel names are deliberately kept: without them, cross-channel narrative research becomes impossible, because the whole question is who published first and who amplified whom. Channels are public broadcasters, not private individuals.
A limitation worth stating plainly: pseudonymization is not anonymity. A determined party who already holds the original public messages can re-link some comments by matching text. Treat this as de-identified research data, not as a privacy guarantee for the people in it.
How to cite
The licence is CC BY 4.0. Use it freely, including commercially, with attribution.
Plain
Aisberg Public Organization. Aisberg Telegram News UA (anonymized) [Data set].
Hugging Face. https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset
BibTeX
@misc{aisberg_telegram_news_ua,
title = {Aisberg Telegram News UA (anonymized)},
author = {{Aisberg Public Organization}},
howpublished = {\url{https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset}},
note = {Available at \url{https://aisberg.live/data/}},
license = {CC-BY-4.0}
}
If you publish something built on this data, we'd like to know: info@aisberg.live. We keep a list.
Need more than the public snapshots?
Our working corpus is larger and finer-grained than what we publish: full snapshot series, per-post view and reaction curves, comment threading, channel-level aggregates, and analysis outputs that don't fit the monthly format. The public release is deliberately conservative next to it, with day-level dates and a fixed set of fields.
We are a registered Ukrainian non-profit (ЄДРПОУ 46034329). If you're a researcher, journalist, fact-checking organization or NGO and the public files don't cover your question, write to us and describe what you need. We can usually prepare a custom extract for a specific date range, a set of channels, a narrative, or a particular field set. For non-commercial research we do it for free.
Please include: what you're researching, the period and channels you need, the fields that matter, and how the result will be published. It helps us say yes faster.
News gets cleaner when somebody checks it
News clusters, fact-check verdicts and a link to the dashboard go out on the Telegram channel every hour. Reading it is free and needs no account. And if you want more of this work to happen, you can join it.
Subscribe to the channel Become a volunteer