Data & Code
Datasets and code from my research. My public repositories are on GitHub.
Datasets
TBMM Plan and Budget Committee Discourse Corpus
A structured, machine-readable corpus of the Turkish Grand National Assembly (TBMM) Plan and Budget Committee (Plan ve Bütçe Komisyonu, PBK) budget proceedings. In Türkiye, the Plan and Budget Committee is the first and most detailed parliamentary stage where the central government budget is deliberated before plenary debate. Each ministry’s budget is discussed in depth; ministers, bureaucrats, and members of parliament engage in extensive technical exchanges.
- Version: v1.1.0, 28 August 2026
- Coverage: budget years 2009–2025
- Size: 231,923 speaker turns; 17.4 million words; 858 MPs
- Format: Parquet
- Licences: data and documentation CC BY 4.0; code MIT
- Zenodo: 10.5281/zenodo.20457565 (concept DOI, always resolves to the latest version); v1.1.0: 10.5281/zenodo.22150634
- GitHub: eozyerden/tbmm-pbk-corpus
Quick start in R
Example from the repository README. The processed data are on Zenodo (tbmm-pbk-corpus-data-v1.1.0.zip); extract the files into data/processed/.
library(arrow)
library(dplyr)
df <- read_parquet("data/processed/konusmalar_metadata.parquet")
# Speeches per budget year
df |> count(butce_yili)
# Word counts by party (party affiliation is in `mv_parti`, not `parti`)
df |>
filter(!is.na(mv_parti)) |>
group_by(mv_parti) |>
summarise(total_words = sum(kelime_sayisi)) |>
arrange(desc(total_words))
# Opposition MP speeches in the 2020 budget hearings
df |>
filter(butce_yili == 2020, rol == "milletvekili", mv_parti %in% c("CHP", "HDP", "İYİP"))How to cite
Özyerden, E. (2026). TBMM Plan and Budget Committee Discourse Corpus (Version 1.1.0) [Data set]. https://doi.org/10.5281/zenodo.20457565
BibTeX
@misc{ozyerden2026tbmm,
author = {Özyerden, Emre},
title = {{TBMM Plan and Budget Committee Discourse Corpus}},
year = {2026},
version = {1.1.0},
doi = {10.5281/zenodo.20457565},
url = {https://github.com/eozyerden/tbmm-pbk-corpus}
}