Enhancing use of BERT information in neural machine translation
preprint
OA: closed
Abstract
Although BERT has achieved excellent results in various natural language processing tasks, it does not exhibit the same high performance in cross-lingual tasks, especially machine translation tasks. We propose a BERT enhanced neural machine translation (BE-NMT) model to improve the use of the information that is contained in BERT by NMT. The model consists of three aspects: (1) A MASKING strategy is applied to alleviate the knowledge forgetting that is caused by the fine-tuning of BERT on the NMT task.(2) Serial and parallel processing are combined for the multi-attention models when incorporating BERT into the NMT model. (3) The multiple hidden layer outputs of BERT are fused to supplement the missing linguistic information of its final hidden layer output. Experiments demonstrate that our method achieves good improvements in various NMT tasks compared with the baseline model.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00
- unpaywall
- last seen: 2026-06-02T02:00:03.124865+00:00