TLDR;
- Remove infrequent words from training data, don’t worry about end of sentences
- Architecture [Embedding x2] -> Dot product -> Sigmoid
- We’re talking about the negative sampling training process
- Make your context window size dynamic !
- Learning rate decrease linearly .005 -> 0.000005
- You can skip the sigmoid backprop, it still works
- Run all you positive/negative samples in one pass
- Code : https://github.com/pierre-wilmot/Word2Vec-CPP
I recently had a chance to play with word embeddings as part of my job at Fetch.ai, more specificaly, I was asked to reimplement the Word2Vec model using Fetch in house machine learning library.
When starting, a lot of questions with no obvious answers came to mind and I made a lot of mistakes.
After a couple of days, I realised it would probably take me forever to get it right and I changed my strategy. I cloned a working implementation from github (https://github.com/chrisjmccormick/word2vec_commented) and starting making incremental changes.
In this post I’m sharing the tricks to get this model to work as a developper. I do not claim to be an expert when it comes to NLP, but I can walk throught the steps to actually get a working product.
1 – Data preprocessing:
I’m using the text8 dataset from Matt Mahoney as a corpus. This is already preprocessed text, with no pontuation, all lowercase, number converted to text and so on.
One of the question I had when I started was « Can you have a context window overlap on two different sentences ? ». Turns out you can, this is how Mikolov does it.
Second thing, you want to keep only the words that appear only a least X times in your training data. In my case, this number is set to 5. It’s seems like it’s prefectly fine if the actual data you feed in model is missing words, the system recovers anyway.
2 – Model Architecture
My first (very naive and full of guesswork) attemp was a looking like that:
Input a [CONTEXT_WINDOW_SIZE x EMBEDDINGS_DIMS] vector that is a concatenation of the word vectors -> FullyConnected layer with output 100 -> RELU -> FullyConnected layer with output VOCAB_SIZE -> Softmax.
⬆ That runs slow and doens’t work.
What I’m describing now is a CBOW (Continious Bag Of Words) model training with negative sampling.
First you need two embeddings matrices that hold VOCAB_SIZE elements of size EMBEDDING_DIMS.
One is a classic embeddings, and the other needs to be what pytorch call an EmbeddingsBag, that means, all the elements you request gets averaged together and the output is always [1xEMBEDDING_DIMS].
The EmbeddingBags module gets as input the context words and return a [1xEMBEDDINGS_DIM] vector.
The the regular embeddings gets as input the target word and return a [1xEMBEDDINGS_DIM] vector.
Then the model is extremely simple, it’s just a dot product of these two vectors followed by a sigmoid activation.
3 – Negative Sampling
As I said, what I’m describing in this post is the negative sampling training process. It’s method where rather then predicting which word is the target, the system predict wheter the pair [context, target] it recieved as input is valid or not. That allow to remove the softmax from the architecture and makes things much faster.
In the original implementation, the negative sampling paramater is set to 25. That is, for every word in the corpus, we train the network to predict 25 « incorrect pairs ».
Negative samples are chosen according to word frequency in the corpus, the more frequent a word, the more likeky it is to be chosen as a negative target.
For performance, sampling is done through a « UnigramTable », that is just a very large array that contain the word ids repeated a certain number of time according to their frequency. Make it easy to sample using a random index sampled from a uniform distribution.
4 – Context window
That one seems to be mandatory in order to get a working system. The context window size has to be dynamic. In this implementation, randomly chosen between 3 & 10 for each sample. A fixed context window size doesn’t seem to work.
5 – Learning rate
Learning rate is gradually decreased as training progress (in a linear fashion). It starts at 0.05 and ends at 0.000005.
6 – Backpropagation
This one is a strange one, on the forward pass, the network computes:
Embeddings Lookup -> Dot product -> Sigmoid -> L1 distance
But on the backward pass, the derivative for the sigmoid is not computed, it goes
L1 -> Dot Product -> Embeddings
The process, would probably works too if the sigmoid derivative was included, but the learning rate would need to be increased.
7 -Make it faster
The C++ code runs much slower than the original C code, this is due to a lot of extra memory copies. These are required because when using a general purpose ML library, each module need to have ownership of it’s own memory buffer, and some extra temporary buffers are added eto keep the interfaces consistent even though in this specific usecase they are not needed.
But one of the way to make the code faster is to turn the 25 Dot product into a matrix multiplication.
Instead of running the Dot product in a loop, it is possible to group all the samples (the positive one and all the negative ones) into a matrix and make a context_vector * target_matrix multiplication. That does help make things a lot faster.
8 – Code
Code is available here : https://github.com/pierre-wilmot/Word2Vec-CPP
I may add a Pytorch version at a later point.











