摘要
针对word2vec模型生成的词向量缺乏语境的多义性以及无法创建集外词(OOV)词向量的问题,引入相似信息与word2vec模型相结合,提出word2vec-ACV模型。该模型首先基于连续词袋(CBOW)和Hierarchical softmax的word2vec模型训练出词向量矩阵即权重矩阵;然后将共现矩阵进行归一化处理得到平均上下文词向量,再将词向量组成平均上下文词向量矩阵;最后将平均上下文词向量矩阵与权重矩阵相乘得到词向量矩阵。为了能同时解决集外词及多义性问题,将平均上下文词向量分为全局平均上下文词向量(global ACV)和局部平均上下文词向量(local ACV)两种,并对两者取权值组成新的平均上下文词向量矩阵,并将word2vec-ACV模型和word2vec模型分别进行类比任务实验和命名实体识别任务实验。实验结果表明,word2vec-ACV模型同时解决了语境多义性以及创建集外词词向量的问题,降低了时间消耗,提升了词向量表达的准确性和对海量词汇的处理能力。
The word2vec model is a neural network model (NNLM) that converts words in text into a word vector. It is widely used in natural language processing tasks such as emotional analysis, question-answering robot and so on. Word vectors generated for the word2vec model lacked the ambiguity of context and the inability to create OOV word vectors. Based on the similarity information of document context and word2vec model, this paper proposed a word vector generation model called the word2vec-ACV model which conformed to the meaning of OOV context. The model was similar to the process of the word vector generated by the word2vec model. First of all, base on the continuous word bag (CBOW) and the Hierarchical softmax, the word2vec model trained the word vector matrix, namely the weight matrix. Secondly, normalized the co-occurrence matrix to get the average context word vector. Then, the word vector consisted of an average context word vector matrix. Finally, it multiplied the average context word vector matrix and the weight matrix to get the word vector matrix. In order to simultaneously solve the ambiguity problem of out of vocabulary words and out of vocabulary words to create, this paper divided the average context word vectors into the global average context word vector (global ACV) and the local average context word vector (local ACV). In addition, the two taken the weight value to form a new average context word vector matrix. The word2vec model could effectively express the word in vector form. Experiments on analogical tasks and named entity recognition (NER) tasks respectively, the results show that the word2vec-ACV model is superior to the word2vec model in the accurate expression of the word vector. It is a word vector representation method to create a contextual context for OOV words.
作者
王永贵
郑泽
李玥
Wang Yonggui;Zheng Ze;Li Yue(College of Software,Liaoning Technical University,Huludao Liaoning 125105,China)
出处
《计算机应用研究》
CSCD
北大核心
2019年第6期1623-1628,共6页
Application Research of Computers
基金
国家自然科学基金青年基金资助项目(61404069)