Loading...
Thumbnail Image
Item

Automatic translation of scientific documents in the HAL archive

Lambert, Patrik
Schwenk, Holger
Blain, Frederic
Editors
Other contributors
Affiliation
Epub Date
Issue Date
2012-05-31
Submitted date
Subjects
Alternative
Abstract
This paper describes the development of a statistical machine translation system between French and English for scientific papers. This system will be closely integrated into the French HAL open archive, a collection of more than 100.000 scientific papers. We describe the creation of in-domain parallel and monolingual corpora, the development of a domain specific translation system with the created resources, and its adaptation using monolingual resources only. These techniques allowed us to improve a generic system by more than 10 BLEU points.
Citation
Lambert, P., Schwenk, H. and Blain, F. (2012) Automatic translation of scientific documents in the HAL archive, in Calzolari, N. et al. (eds) Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC’12), 21st-27th May, 2012, Istanbul, Turkey.
Journal
Research Unit
DOI
PubMed ID
PubMed Central ID
Embedded videos
Type
Conference contribution
Language
en
Description
© 2012. Published by ELRA. This is an open access article available under a Creative Commons licence. The published version can be accessed at the following link on the publisher’s website: http://www.lrec-conf.org/proceedings/lrec2012/pdf/703_Paper.pdf
Series/Report no.
ISSN
EISSN
ISBN
ISMN
Gov't Doc #
Sponsors
Rights
Research Projects
Organizational Units
Journal Issue
Embedded videos