Write the memory mappable model form after training

derekovecs memory maps <model>.vecs and <model>.words next to the model. When
they are missing it builds them on its first start, which takes a while and
needs write access to the directory the models live in. A server in a
container usually has that directory mounted read only and then does not start
at all:

  Converting /models/dereko-2026-i.vecs to memory mappable structures
  Cannot open /models/dereko-2026-i.vecs.words or .../dereko-2026-i.vecs.vecs

dereko2vec writes both files right after saving the vectors now, and
vecs2mmap does it for models that were trained earlier.

The conversion mirrors init_net() of derekovecs including the order of the
floating point operations. Verified against the model of the derekovecs test
data: the vectors come out byte identical, and derekovecs serves the same
neighbours and collocators from the generated files, with the model directory
mounted read only and without converting anything itself.

Words are cut off at 49 characters plus the terminator instead of filling all
50 bytes, so that every entry is terminated within its own record.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Change-Id: I297e2db98bd425aa388e441aca6d2b0b14f0c66c
3 files changed
tree: d59698eebe5848b9fd9635ecb01d5354a6786977
  1. ci/
  2. scripts/
  3. src/
  4. tests/
  5. .gitignore
  6. .gitlab-ci.yml
  7. CMakeLists.txt
  8. LICENSE
  9. README.md
README.md

dereko2vec

Fork of wang2vec with extensions for re-training and count based models, support for tokens with frequencies > 2³² and a more accurate ETA prognosis.

Installation

Dependencies

Build and install

cd dereko2vec
mkdir build
cd build
cmake ..
make && ctest3 --extra-verbose && sudo make install

Run

The command to build word embeddings is exactly the same as in the original version, except that we added type 5 for setting up a purely count based collocation database.

The -type argument is a integer that defines the architecture to use. These are the possible parameters:
0 - cbow
1 - skipngram
2 - cwindow (see below)
3 - structured skipngram(see below)
4 - collobert's senna context window model (still experimental)
5 - build a collocation count database instead of word embeddings

Example

./dereko2vec -train input_file -output embedding_file -type 0 -size 50 -window 5 -negative 10 -nce 0 -hs 0 -sample 1e-4 -threads 1 -binary 1 -iter 5 -cap 0

Generate dereko2vec training input files from KorAP-XML ZIPs

The KorAP-XML-CoNLL-U tool can be used to generate input files for dereko2vec from KorAP-XML ZIPs using its tokenization and setence boundary information, for example:

korapxml2conllu --word2vec wpd19.zip > wpd19.w2vinput

Retrain existing model with new data

For example:

Retrain Vectors:

dereko2vec -train new.traindata -output new.vecs -save-net new.net -type 3 -size 200 -window 5 -negative 10 -threads 44 -binary 1 -iter 100 -read-vocab old.vocab -read-net old.net

Create new RocksDB:

dereko2vec -train new.traindata -output new.rocksdb -type 5 -window 5 -threads 8 -binary 1 -iter 1 -read-vocab old.vocab -sample 0 -min-count 0
dereko2vec -train new.traindata -output .temp.rocksdb -type 5 -window 5 -threads 8 -binary 1 -iter 1 -save-vocab new_focus.vocab -sample 0 -min-count 0
rm -rf .temp.rocksdb
python scripts/merge_vocabs.py old.vocab new_focus.vocab new.vocab

References

@InProceedings{Ling:2015:naacl,  
author = {Ling, Wang and Dyer, Chris and Black, Alan and Trancoso, Isabel},  
title="Two/Too Simple Adaptations of word2vec for Syntax Problems",  
booktitle="Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",  
year="2015",  
publisher="Association for Computational Linguistics",  
location="Denver, Colorado",  
}

@InProceedings{FankhauserKupietz2019,
author    = {Peter Fankhauser and Marc Kupietz},
title     = {Analyzing domain specific word embeddings for a large corpus of contemporary German},
series = {Proceedings of the 10th International Corpus Linguistics Conference},
publisher = {University of Cardiff},
address   = {Cardiff},
year      = {2019},
note      = {\url{https://doi.org/10.14618/ids-pub-9117}}
}