自然言語処理学研究室のAdam Nohejlさん(博士課程修了生)らが、21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026)にて、Best Paper in the Shared Task on Vocabulary Difficulty Prediction for English Learnersを受賞しました。(2026/7/4)
|
21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026)は、ACLコミュニティで最大級のワークショップの一つで、2026年7月3日〜4日にサンディエゴ(アメリカ)にて開催されました。 The BEA Workshop is a leading venue for NLP innovation in the context of educational applications. It is one of the largest workshops in the ACL community with over 100 registered attendees in the past several years. The BEA 2026 shared task on vocabulary difficulty drediction, organized by the British Council, aims to advance research into vocabulary difficulty prediction for learners of English with diverse L1 backgrounds, an essential step towards custom content creation, computer-adaptive testing and personalised learning. |
- 受賞者/著者 Awardees /Authors:
Adam Nohejl(RIKEN AIP/doctoral degree from NAIST),Xuanxin Wu(The University of Osaka) Yusuke Ide(博士後期課程3年),Maria Angelica Riera Machin(master's degree from NAIST) Yi-Ning Chang(National Tsing Hua University),Hitomi Yanaka(RIKEN/The University of Tokyo/Tohoku University) -
写真はAdam Nohejlさん - 受賞研究テーマ Research theme:
"Sakura at BEA 2026 Shared Task 1: What Makes Vocabulary Difficult?"
We describe two types of models for vocabulary difficulty prediction: a high-accuracy black-box model, which achieved the top shared task result in the open track, and an explainable model, which outperforms a fine-tuned encoder baseline. As the black-box model, we fine-tuned an LLM using a soft-target loss function for effective application to the rating task, achieving r > 0.91. The explainable model provides insights into what impacts the difficulty of each item while maintaining a strong correlation (r > 0.77). We further analyze the results, demonstrating that the difficulty of items in the British Council's Knowledge-based Vocabulary Lists (KVL) is often affected by spelling difficulty or the construction of the test items, in addition to the genuine production difficulty of the words. - (上記和訳) 本研究では、語彙の難易度予測のための2種類のモデルを提案する。1つは高精度なブラックボックスモデルであり、オープントラックにおいてシェアドタスクの最高成績を達成した。 もう1つは説明可能なモデルであり、ファインチューニングされたエンコーダを用いたベースラインを上回る性能を示した。 ブラックボックスモデルでは、大規模言語モデル(LLM)をソフトターゲット損失関数を用いてファインチューニングし、評価タスクへ効果的に適用することで、 相関係数 r > 0.91 を達成した。一方、説明可能なモデルは、各語彙項目の難易度に影響を与える要因についての洞察を提供するとともに、高い相関(r > 0.77)を維持している。 さらに結果を分析したところ、British Council の Knowledge-based Vocabulary Lists(KVL)に含まれる項目の難易度は、単語そのものの本質的な産出の難しさだけでなく、綴りの難しさやテスト項目の構成によっても大きく左右される場合が多いことが明らかになった。
21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026) https://sig-edu.org/bea/2026
