| commit | 19fa44868e620946610cdd6560fa9c55c22fa583 | [log] [tgz] |
|---|---|---|
| author | Marc Kupietz <kupietz@ids-mannheim.de> | Sun Jun 07 15:03:56 2026 +0200 |
| committer | Marc Kupietz <kupietz@ids-mannheim.de> | Sun Jun 07 15:03:56 2026 +0200 |
| tree | 875c1ecb7cec2169db9601db3ce904e4924b1f82 | |
| parent | c523981d9536f687473656d46a573ba7d09f5175 [diff] |
Add wikidomain topic-domain classification to the ingestion pipeline Generate Wikipedia top-level topic-domain classifications as stand-off metadata (one <standOff> XML per corpus, via the korap/wiki-taxonomy image) and fold them into the Krill index as a queryable wikidomain keywords field. - New $(BUILD_DIR)/%.wikidomain.meta.xml rule: korapxmltool -t now | docker run korap/wiki-taxonomy. - wikidomain joins the standard ANNOTATIONS list; META_ANNOTATIONS tags which layers emit a stand-off .meta.xml instead of a foundry zip, and an artifact mapper resolves each annotation to the right build target. - krill.tar / KRILL_PREREQS fold the active meta files in; korapxmltool auto-detects the .meta.xml inputs. New `make meta` target builds the classifications without indexing. - Readme: document the classification and META_ANNOTATIONS. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Change-Id: Ibc00aefa4374477f8bc9302b531ad6f6065060d0
Converts TEI XML files to KorAP XML, annotates them with TreeTagger, Marmot, Malt, Spacy, CoreNLP, OpenNLP, indexes them and starts a KorAP instance.
By default, place your .i5.xml files in the I5 directory in the root of the project:
make
Alternatively, you can specify a different source directory containing your .i5.xml files using the SRC_DIR variable:
make SRC_DIR=/path/to/your/files
Then open http://localhost:64543 in your browser.
If port 64543 is already in use on your host (e.g. by another process or IDE port-forwarding helper), you can specify a different host port using the KORAP_PORT variable:
make KORAP_PORT=64544
Then open the corresponding URL (e.g., http://localhost:64544) in your browser.
Standard TEI P5 files (which typically contain one text per file) can be batch-converted together. By default, place your .xml files in the TEI directory in the root of the project:
make
All files in the TEI directory will be packaged into a single KorAP-XML zip archive.
You can specify a different source directory using the TEI_DIR variable:
make TEI_DIR=/path/to/your/tei/files
By default, texts are assigned automatic three-part KorAP/DeReKo sigles starting from:
--auto-textsigle 'TEI/XYZ.00001'
You can customize this or define a mapping from the xml:id of each text to a sigle. For example, if your XML files have IDs like this:
<?xml version="1.0" encoding="UTF-8"?> <TEI xmlns="http://www.tei-c.org/ns/1.0" xml:id="SK.UL.1981.1" version="3.4.0"> <teiHeader>
You can extract the three-part sigle (e.g., SK/UL/19811) by overriding the TEI_FLAGS variable with a matching regular expression:
make TEI_FLAGS='--xmlid-to-textsigle '\''([A-Z]+)\.(.*)\.([0-9]+)\.*([0-9])@$$1/$$2/$$3$$4'\'''
You can specify which annotation layers to run and package by setting the ANNOTATIONS variable. By default, it runs: marmot-malt tree_tagger spacy corenlp opennlp wikidomain
To enable gender annotation (using the conllu-gender tool), add gender to the list:
make ANNOTATIONS="marmot-malt tree_tagger spacy corenlp opennlp gender"
wikidomain is a special annotation: instead of a per-text foundry zip it produces a single stand-off metadata XML per corpus (build/<corpus>.wikidomain.meta.xml) holding Wikipedia top-level topic-domain classifications, generated with the korap/wiki-taxonomy Docker image. korapxmltool auto-detects these .meta.xml files and folds them into the Krill index as a wikidomain keywords field, queryable per text in KorAP.
It is enabled by default. Such metadata-producing layers are listed in META_ANNOTATIONS (default: wikidomain); everything else in ANNOTATIONS is treated as a foundry zip. To build just the classifications without indexing:
make meta
To disable it, drop it from ANNOTATIONS (e.g. make ANNOTATIONS="marmot-malt spacy").