# Subset of Google model Vector file

**URL:** <https://community.konduit.ai/t/subset-of-google-model-vector-file/256>\
**Category:** DL4J\
**Created:** [March 11, 2020, 8:06pm UTC](https://community.konduit.ai/t/subset-of-google-model-vector-file/256 "2020-03-11T20:06:33Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ethiel](https://avatars.discourse-cdn.com/v4/letter/e/df788c/32.png) [@ethiel](https://community.konduit.ai/u/ethiel)\
**Post date:** [March 11, 2020, 8:06pm UTC](https://community.konduit.ai/t/subset-of-google-model-vector-file/256/1 "2020-03-11T20:06:33Z")

</div>

Hi, people.  
I’m using the widely known Google vector model for sentiment analysis along with a CNN neural network.  
It’s working after a lot of effort and it’s working fine. However, for unit testing, I’d like to use a small part of that file hence a subset. I tried with gensim but the file was not correct, the number of words and the size were not correctly got when I try to use the load static method.  
So, is there a way to use a small subset of that file using DL4J?

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [March 12, 2020, 7:36am UTC](https://community.konduit.ai/t/subset-of-google-model-vector-file/256/2 "2020-03-12T07:36:38Z")

</div>

What exactly have you tried? For the text file based word vectors, you can easily reduce its size by removing the lines that you don’t need.

---

<div class="post-metadata">

**Author:** ![ethiel](https://avatars.discourse-cdn.com/v4/letter/e/df788c/32.png) [@ethiel](https://community.konduit.ai/u/ethiel)\
**Post date:** [March 12, 2020, 8:43am UTC](https://community.konduit.ai/t/subset-of-google-model-vector-file/256/3 "2020-03-12T08:43:42Z")

</div>

Hi, @treo thanks for answering.  
I solved my problem by editing the file manually to add the number of words and the size of the vectors. However, there is something curious about gensim: even although I say to gensim to use only 200 words, it took 184; I don’t understand why, but that was the issue.  
Thanks for your help, @treo
