Source author record

Romain Bey

Romain Bey appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

2works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

2 published item(s)

preprint2022arXiv

Learning structures of the French clinical language:development and validation of word embedding models using 21 million clinical reports from electronic health records

Background Clinical studies using real-world data may benefit from exploiting clinical reports, a particularly rich albeit unstructured medium. To that end, natural language processing can extract relevant information. Methods based on transfer learning using pre-trained language models have achieved state-of-the-art results in most NLP applications; however, publicly available models lack exposure to speciality-languages, especially in the medical field. Objective We aimed to evaluate the impact of adapting a language model to French clinical reports on downstream medical NLP tasks. Methods We leveraged a corpus of 21M clinical reports collected from August 2017 to July 2021 at the Greater Paris University Hospitals (APHP) to produce two CamemBERT architectures on speciality language: one retrained from scratch and the other using CamemBERT as its initialisation. We used two French annotated medical datasets to compare our language models to the original CamemBERT network, evaluating the statistical significance of improvement with the Wilcoxon test. Results Our models pretrained on clinical reports increased the average F1-score on APMed (an APHP-specific task) by 3 percentage points to 91%, a statistically significant improvement. They also achieved performance comparable to the original CamemBERT on QUAERO. These results hold true for the fine-tuned and from-scratch versions alike, starting from very few pre-training samples. Conclusions We confirm previous literature showing that adapting generalist pre-train language models such as CamenBERT on speciality corpora improves their performance for downstream clinical NLP tasks. Our results suggest that retraining from scratch does not induce a statistically significant performance gain compared to fine-tuning.

preprint2020arXiv

Probing the concept of line tension down to the nanoscale

A novel mechanical approach is developed to explore by means of atom-scale simulation the concept of line tension at a solid-liquid-vapor contact line as well as its dependence on temperature, confinement, and solid/fluid interactions. More precisely, by estimating the stresses exerted along and normal to a straight contact line formed within a partially wet pore, the line tension can be estimated while avoiding the pitfalls inherent to the geometrical scaling methodology based on hemispherical drops. The line tension for Lennard-Jones fluids is found to follow a generic behavior with temperature and chemical potential effects that are all included in a simple contact angle parameterization. Former discrepancies between theoretical modeling and molecular simulation are resolved, and the line tension concept is shown to be robust down to molecular confinements. The same qualitative behavior is observed for water but the line tension at the wetting transition diverges or converges towards a finite value depending on the range of the solid/fluid interactions at play.