Digital resources in the Social Sciences and Humanities OpenEdition Our platforms OpenEdition Books OpenEdition Journals Hypotheses Calenda Libraries OpenEdition Freemium Follow us

Author: Christof Schöch

Teaching Python programming with DOAJ’s journal dataset

Skip to content Note: This post first appeared on the DOAJ Blog on March 30, 2026 under a Creative Commons Attribution-Non Commercial CC-BY-NC license.  Open science practices are an important aspect of Digital Humanities. Therefore, open access and the crucial role DOAJ plays for the documentation and dissemination of information…

Digital Humanities, Open Science, and Research Data Management: Exploring their intersections in Computational Literary Studies in Germany

Note that this post is the original English version of an editorial first published as《巻頭言》「デジタル人文学、オープンサイエンス、研究データ管理:ドイツの計算文学研究における交点を探る」in: Digital Humanities Monthly 167-1, June 2025 (International Institute for Digital Humanities, Tokyo, Japan). URL: https://w.bme.jp/bm/p/bn/htmlpreview.php?i=dhm&no=all&m=124.) Digital Humanities, Open Science, and Research Data Management – three terms that have become something of a trinity in the…

Teaching and Research, or: a Post-scriptum to ‘Topic Modeling Genre’

Wouldn’t it be nice if teaching and research could always be closely connected? And I don’t just mean students benefitting from their instructors being active researchers. I also mean it the other way around, with researchers benefitting from the work they are doing with the students. The true Einheit von…

Dear fellow stylometrists, let’s drop the dendrogram and cherish the distance matrix

The dendrogram is a classic and beloved visualization in stylometry. One could even say that, ever since it was included in the inevitable and unmatched stylo package for R, it has become an icon and metonymy of stylometry itself. The example shown below (Figure 1) illustrates what such a dendrogram…

Computational Genre Analysis

Note: This text has initially been prepared in 2015-2016, then revised in 2019 and 2021, for an edited volume on Digital Humanities for Literary Studies: Methods, Tools & Practices. Despite the best efforts of all involved, the volume never saw the day. For this reason, the text is now made…

Can Atom replace oXygen?

Virtually anyone working with XML files in the context of the Digital Humanities, and especially in the context of scholarly digital editing, knows the oXygen XML editor. It is mature and packed with useful features, and yet every new version brings even more features and improvements. In fact, it is…

Topic Modeling with MALLET: Hyperparameter Optimization

This is a short technical post about an interesting feature of Mallet which I have recently discovered or rather, whose (for me) unexpected effect on the topic models I have discovered: the parameter that controls the hyperparameter optimization interval in Mallet.[1] Yes, there are parameters, there are hyperparameters, and there…

Follow-up on Simenon and Sentence-Length: Visualization and Hypothesis-Testing

In a conversation about my recent post on sentence length in Georges Simenon’s work, Fotis Jannidis said he thought the post was typical of quite a lot of recent work in digital literary studies in that it is exploratory rather than focused on hypothesis testing. I think this is true…

Detecting Transpositions when Comparing Text Versions using CollateX

The aims of this post are to explain what collation is, why detecting transpositions is special, and how to accomplish it using CollateX. The example I will be using involves comparing two versions of the recent best-selling novel The Martian by Andy Weir, a novel whose publishing history Erik Ketzan…

Does Shorter Sell Better? Belgian author George Simenon’s use of sentence length

Belgian author Georges Simenon is probably most famous for his crime fiction novels in which police detective Maigret investigates serious crimes and elucidates them with intelligence, empathy, a team of inspectors and acquaintances, and, of course, his tobacco pipe as well as sandwiches and beers brought to his office. Simenon…

How good are our texts, really? Quality assurance for literary texts from various sources

by Ulrike Henny and Christof Schöch — this post originally appeared on the CLiGS blog. Some weeks ago, we made our “New Year’s release” of text collections available. We publish the texts in the “CLiGS” group’s GitHub repository called “textbox“ and archive each release on Zenodo where they get a…

Europeana for Quantitative Literary History (Europeana Research Blog, Text Mining #3)

Note: The following post first appeared on the Europeana Research Blog on November 30, 2015, in their “Text Mining” series which also includes posts by Ted Underwood and Gregor Wiedemann. Over the last several years, there has been an increasing interest in large-scale, computational, quantitative investigations into European literary history.…

The Schiller-Kleist Uncertainty Principle (guest post)

This is a guest post, written by Philip Dürholt, a student in the “Digital Humanities” program (BA and MA) at the Department for Literary Computing, University of Würzburg. As always, comments are welcome! I’m a student at the University of Würzburg. My subjects are Philosophy and Digital Humanities. Last winter…

How to Create Lemmatized (French) Text for Topic Modeling

“Gravure représantant Pierre Corneille.” From Wikipedia; source: Bibliothèque nationale de France. http://commons.wikimedia.org/wiki/File:Gravure_Pierre_Corneille.jpg (public domain). It would not, some years ago, have occurred to me that anyone would want to reduce literary texts to the following pitiful state: “me avoir tu faire un rapport bien sincère / ne déguiser tu rien…