KorAP: References / Citation Help
If you use the research tools and corpus data provided here in your research, please cite at least one of the listed publications on KorAP and on the underlying corpus and, where applicable, the relevant publications on the annotations, classifications and preparation steps used:
DeReKo
IDS (2026): Deutsches Referenzkorpus / Archiv der Korpora geschriebener Gegenwartssprache 2026-I (Release vom 19.01.2026).
Mannheim: Leibniz-Institut für Deutsche Sprache.
Kupietz, Marc/Lüngen, Harald/Diewald, Nils (2023): Das Gesamtkonzept des Deutschen Referenzkorpus DeReKo. Vom Design bis zur Verwendung und darüber hinaus
In: Deppermann, Arnulf/Fandrych, Christian/Kupietz, Marc/Schmidt, Thomas (Hrsg.): Korpora in der germanistischen Sprachwissenschaft. Mündlich, schriftlich, multimedial. Jahrbuch des Instituts für Deutsche Sprache 2022. (= Jahrbuch des Instituts für Deutsche Sprache 2022). Berlin/Boston: de Gruyter, 1-28.
Kupietz, Marc/Lüngen, Harald/Kamocki, Paweł/Witt, Andreas (2018): The German Reference Corpus DeReKo: New Developments – New Opportunities
In: Proceedings of the 11th International Conference on Language Resources and Evaluation (LREC 2018). Miyazaki/Paris: European Language Resources Association (ELRA), pp. 4353-4360.
Kupietz, Marc/Belica, Cyril/Keibel, Holger/Witt, Andreas (2010): The German Reference Corpus DeReKo: A primordial sample for linguistic research
In: Calzolari, Nicoletta et al. (eds.): Proceedings of the 7th conference on International Language Resources and Evaluation (LREC 2010). Valletta, Malta: European Language Resources Association (ELRA), 1848-1854.
KorAP
Diewald, Nils/Gierke, Marco/Kupietz, Marc/Lüngen, Harald (2024): Das Orthografische Kernkorpus (OKK) in DeReKo: Zusammensetzung, Analyse- und Zugriffsmöglichkeiten über KorAP. In: Krome, Sabine/Habermann, Mechthild/Lobin, Henning/Wöllstein, Angelika (Hgg.): Schriftsystem – Norm – Schreibgebrauch. Berlin, Boston: De Gruyter, S. 329–344. https://doi.org/doi:10.1515/9783111389219-017.
Diewald, Nils/Bodmer, Franck/Harders, Peter/Irimia, Elena/Kupietz, Marc/Margaretha, Eliza/Stallkamp, Helge (2021): KorAP und EuReCo – Recherchieren in mehrsprachigen vergleichbaren Korpora
In: Lobin, Henning/Witt, Andreas/Wöllstein, Angelika (Hrsg.): Deutsch in Europa. Sprachpolitisch, grammatisch, methodisch. Jahrbuch des Instituts für Deutsche Sprache 2020. (= Jahrbuch des Instituts für Deutsche Sprache 2020). Berlin/Boston: de Gruyter, 287-294.
Kupietz, Marc/Diewald, Nils/Margaretha, Eliza/Bodmer, Franck/Stallkamp, Helge/Harders, Peter (2020): Recherche in Social-Media-Korpora mit KorAP
In: Marx, Konstanze/Lobin, Henning/Schmidt, Axel (Hrsg.), Deutsch in Sozialen Medien. Interaktiv, multimodal, vielfältig, Jahrbuch des Instituts für Deutsche Sprache 2019. de Gruyter, Berlin/Boston, pp. 373–378.
Diewald, Nils/Hanl, Michael/Margaretha, Eliza/Bingel, Joachim/Kupietz, Marc/Bański, Piotr/Witt, Andreas (2016): KorAP architecture - Diving in the Deep Sea of Corpus Data
In: Proceedings of the 10th International Conference on Language Resources and Evaluation (LREC 2016). Portorož/Paris: European Language Resources Association (ELRA), pp. 3586–3591.
Corpus preparation, annotation and classification
Kupietz, Marc/Diewald, Nils/Lüngen, Harald/Margaretha Illig, Eliza/Stallkamp, Helge/Tran, Uyen-Nhu/Yaddehige, Rameela (2026): EuReCo, KorAP and DeReKo: Updates on Ingestion and Annotation Pipelines, Backend, Interfaces, Operation, and Corpora
In: Proceedings of the 12th Workshop on the Challenges in the Management of Large Corpora (CMLC-12 2026). Palma de Mallorca, Spain: European Language Resources Association (ELRA), pp. 106–112. https://doi.org/10.63317/2xcbb5knp2mi.
Ecker, Jennifer/Kupietz, Marc (2026): Dynamic Thematic Classification for DeReKo: An Exploratory Study Using Wikipedia's Open Taxonomy.
In: Proceedings of the Conference on Natural Language Processing (KONVENS 2026). Hamburg (accepted). Classifier: https://gitlab.ids-mannheim.de/ecker/wikiTaxonomy.
Ochs, Samira (2026): Die morphosyntaktische Integration neuer Gendersuffixe: Eine korpusbasierte Analyse deutschsprachiger Pressetexte.
In: Gender Linguistics 2. https://doi.org/10.65020/0619d927.
Ochs, Samira/Rüdiger, Jan Oliver (2025): Of stars and colons: A corpus-based analysis of gender-inclusive orthographies in German press texts.
In: Schmitz, Dominic/Stein, Simon David/Schneider, Viktoria (eds.): Linguistic intersections of language and gender. Of gender bias and gender fairness. Berlin/Boston: De Gruyter, pp. 31–62. https://doi.org/10.1515/9783111388694.
Honnibal, Matthew/Montani, Ines/Van Landeghem, Sofie/Boyd, Adriane (2020): spaCy: Industrial-strength Natural Language Processing in Python. https://doi.org/10.5281/zenodo.1212303.
Straka, Milan (2018): UDPipe 2.0 Prototype at CoNLL 2018 UD Shared Task
In: Proceedings of the CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. Brussels, Belgium: Association for Computational Linguistics, pp. 197–207. https://doi.org/10.18653/v1/K18-2020.
Manning, Christopher D./Surdeanu, Mihai/Bauer, John/Finkel, Jenny/Bethard, Steven J./McCloskey, David (2014): The Stanford CoreNLP Natural Language Processing Toolkit
In: Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations. Baltimore, Maryland: Association for Computational Linguistics, pp. 55–60. https://doi.org/10.3115/v1/P14-5010.
Mueller, Thomas/Schmid, Helmut/Schütze, Hinrich (2013): Efficient Higher-Order CRFs for Morphological Tagging
In: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. Seattle, Washington, USA: Association for Computational Linguistics, pp. 322–332. https://www.aclweb.org/anthology/D13-1032.
The Apache Software Foundation (2010): Apache OpenNLP. Machine-learning based toolkit for the processing of natural language text. https://opennlp.apache.org/.
Nivre, Joakim/Hall, Johan/Nilsson, Jens (2006): MaltParser: A Data-Driven Parser-Generator for Dependency Parsing
In: Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06). Genoa, Italy: European Language Resources Association (ELRA).
Weiß, Christian (2005): Die thematische Erschließung von Sprachkorpora.
In: OPAL – Online publizierte Arbeiten zur Linguistik 2005(1). http://pub.ids-mannheim.de/laufend/opal/pdf/opal2005-1.pdf.
Schmid, Helmut (1994): Probabilistic Part-of-Speech Tagging Using Decision Trees.
In: International Conference on New Methods in Language Processing. Manchester, UK, pp. 44–49.