コロキアムB発表

日時: 07月27日 (Mon) 1限目(9:20 - 10:50)


会場: L1

司会: 松井 智一
藤原 和真 D, 中間発表 光メディアインタフェース 向川 康博, 林 優一, 藤村 友貴, 北野 和哉
title: Optically Encoded Imaging and Restoration for Real-World Light Understanding
abstract: This talk reports my research progress in the first half of the doctoral program and my current research plan. The presentation consists of three parts. First, I introduce a snapshot spectral imaging method based on birefringent materials and a multi-reflection optical system. Second, I present an extension of this idea to a different physical phenomenon and optical system, aiming to recover information that is difficult to separate from ordinary observations. Third, I describe my ongoing research plan on imaging and restoration in real-world scenes containing strong light sources and reflections. Finally, I summarize future challenges and the research direction for the latter half of the doctoral program.
language of the presentation: Japanese
発表題目: 光学符号化に基づく撮像・復元と実環境光理解への展開
発表概要: 本発表では、博士課程前半の研究成果と、現在進めている研究計画について報告する。内容は三つである。第一に、複屈折媒体と多重反射光学系を用いたスナップショット分光撮像手法について述べる。第二に、この考え方を別の物理現象および光学系へ発展させ、単一撮影から通常観測では分離しにくい情報を復元する手法について報告する。第三に、現在取り組んでいる新しい研究計画として、強光源や反射が混在する実環境において、画像中の光の成分をより適切に扱うための撮像・復元手法について述べる。最後に、これらの成果を踏まえ、博士課程後半の研究方針と今後の課題を整理する。
 
中島 夕翔 M, 2回目発表 光メディアインタフェース 向川 康博, 林 優一, 藤村 友貴, 北野 和哉
title: Improving the Robustness of Laser Speckle Authentication Against Axial Displacement by Deep Learning-Based Reference Data Expansion
abstract: Laser speckle authentication is highly sensitive to displacement of the illuminated surface, resulting in a significant degradation in authentication accuracy when the target object is misaligned. In particular, robustness against axial displacement remains a challenge because it cannot be sufficiently improved without sacrificing the laser beam diameter. This study aims to improve the robustness against axial displacement in practical environments by expanding the reference data using deep learning. The proposed method will be evaluated by comparing its authentication performance with that of a phase retrieval-based reference data expansion method. Currently, we are developing a speckle size-based evaluation method to validate the physical consistency of the simulated speckle images used as training data.
language of the presentation: Japanese
発表題目: 深層学習を用いた参照データ拡張によるレーザスペックル認証の光軸方向位置ずれに対するロバスト性の向上
発表概要: レーザスペックル認証はレーザ照射面の移動にセンシティブであるため,認証する物体の位置ずれによって認証精度が著しく低下する.特に,光軸方向の移動に対するロバスト性はレーザビーム径とのトレードオフにより確保できず,依然として課題である.本研究では,深層学習を用いて参照データを拡張することで,実環境における光軸方向の位置ずれに対するロバスト性の向上を目指す.また,位相回復による参照データ拡張手法との比較を通して,提案手法の有効性を評価する.現在は,学習データとして用いるシミュレーションの妥当性を検証するため,スペックルサイズを指標とした評価手法の構築を進めている.
 
宮川 甫哉 M, 2回目発表 光メディアインタフェース 向川 康博, 林 優一, 藤村 友貴, 北野 和哉
title: Object Authentication via Optimization of Speckle Intensity Statistics
abstract: Speckle-based authentication, a type of artifact-metric technique, has a known limitation in that it is not robust to axial displacement of the speckle pattern. This study aims to overcome this limitation by using the statistical properties of speckle intensity, whose axial correlation decays more gradually than the speckle pattern itself, for authentication. For each target object, the object-generated speckle is optimized using a spatial light modulator (SLM), and the object is identified based on the resulting intensity statistics. Simulation experiments confirmed two points: first, that the speckle intensity statistics can be successfully optimized; and second, that the target object can be correctly identified among 200 objects. Future work will further evaluate the robustness of the proposed method as an authentication technique.
language of the presentation: Japanese
発表題目: スペックルの強度統計最適化による物体認証
発表概要: 人工物メトリクス手法の一種であるスペックル認証は、スペックルパターンの軸方向移動に対するロバスト性が低いという既知の課題がある。そこで本研究では、パターンそのものではなく光軸方向の相関減衰がなだらかであるスペックルの強度統計情報を認証に利用することによりこの課題を克服することを目指す。識別対象ごとにSLM(空間光変調機)を用いて物体スペックルを最適化し、その統計情報により識別を行う。シミュレーション実験を用いて、強度統計の最適化が可能であること、200の物体間では正しく対象を識別できることの2点を確認した。今後は、認証手法としての頑健性の検証をさらに進めていく。
 

日時: 07月27日 (Mon) 1限目(9:20 - 10:50)


会場: L2

司会: Monica Perusquia-Hernandez
CHEN YUETING M, 2回目発表 インタラクティブメディア設計学 加藤 博一, Sakriani Sakti, 澤邊 太志, Isidro Butaslac, 藤本 雄一郎
title: Effects of LLM-Generated Mnemonic Stories in a VR Memory Palace on Vocabulary Learning and Long-Term Retention
abstract: The memory palace technique can support vocabulary learning, but constructing both spatial and semantic associations may impose substantial cognitive demands. This study proposes a VR memory palace system that combines a learner-created familiar room, learner-chosen furniture–word mappings, and a continuous mnemonic story. A between-subjects experiment compares three conditions: an LLM-authored story with furniture–word visualization, a learner-authored story with the same visualization, and a baseline condition without furniture–word visualization. Immediate recall, one-week delayed retention, word–furniture associative memory, cognitive load, mnemonic-construction effort, and subjective learning experience will be evaluated. This study examines how using LLM-generated mnemonic stories instead of learner-authored stories affects vocabulary learning, long-term retention, and cognitive load.
language of the presentation: English
 
谷端 真瑠 M, 2回目発表 ヒューマンAIインタラクション Sakriani Sakti, 渡辺 太郎, 大内 啓樹, Faisal Mehmood, Bagus Tris Atmaja

title:
Understanding When and Where Self-Supervised Speech Models Learn Phonemes and Prosodic Cues through Layer-Wise and Time-Resolved Probing

abstract:
Self-supervised speech models (S3Ms) encode a rich hierarchy of acoustic, prosodic, and phonetic information, yet when and where it emerges during pretraining remains unclear.
We train frozen-encoder probes on every Transformer layer across pretraining checkpoints, comparing two architecturally matched models that differ mainly in objective: HuBERT-Base, whose two-stage pseudo-label curriculum offers a natural reorganization boundary, and single-objective wav2vec 2.0-Base.
Our tasks span an acoustic-to-abstract gradient, from absolute pitch (F0), relative pitch, and lexical stress to phoneme boundaries and 38-way phoneme identity.
Low-level acoustic–prosodic cues and phone boundaries are decodable from the earliest checkpoints and barely shift across layers, whereas phoneme identity is visibly constructed during pretraining—improving most and re-localizing to upper-middle layers at HuBERT's stage-2 transition.
Acquisition order thus tracks abstraction from the acoustic surface, not the segmental/suprasegmental distinction; wav2vec 2.0 reorganizes gradually instead.

language of the presentation:
Japanese

発表題目:
自己教師あり音声モデルはいつ・どの層で音素と韻律の手がかりを学習するのか:層別・時間分解プロービングによる分析

発表概要:
自己教師あり音声モデル(S3M)は音響・韻律・音韻情報の豊かな階層構造を内部に符号化することが知られているが、それらの情報が事前学習の「いつ」の時点で、ネットワークの「どこ」(どの層)に現れるのかは明らかでない。
本発表では、事前学習の各チェックポイントにおいて全Transformer層に対して凍結エンコーダのプローブを訓練し、主に学習目的のみが異なる2つのモデル——2段階の擬似ラベルカリキュラムを持つHuBERT-Baseと、単一目的で学習されるwav2vec 2.0-Base——を比較した結果を報告する。
プロービングタスクは、絶対ピッチ(F0)・相対ピッチ・語強勢から音素境界・38クラス音素識別まで、音響的表層から抽象的表現への勾配をなすように設計した。
分析の結果、低次の音響・韻律的手がかりと音素境界は学習のごく初期から解読可能で層方向の変化も小さい一方、音素の同一性は事前学習を通じて段階的に構築され、HuBERTのステージ2移行時に最も大きく改善するとともに上位中間層へ再局在することが分かった。
すなわち、情報の獲得順序は分節/超分節の区別ではなく音響表層からの抽象度に従い、wav2vec 2.0では同様の再編成が離散的な境界を持たず漸進的に生じる。

 
ZHANG YUTING M, 2回目発表 ヒューマンAIインタラクション Sakriani Sakti, 渡辺 太郎, 大内 啓樹, Faisal Mehmood, Bagus Tris Atmaja
title: Correlations Between Dialectal Speech Embedding and Multiple External Factors: An Analysis of Japanese and UK-Irland English
abstract: Although text-level studies have established links between physical geography and linguistic distance, how speech-level embeddings, especially dialectal speech, derived from Self-Supervised Learning (SSL) models encode dialectal languages under the interplay of dynamic and static factors requires deeper investigation. To address this gap, we present an integrated framework combining WavLM-large speech embeddings with multi-factor spatial features. Through comparative latent space analyses using PCA and a PCA-LDA pipeline, we disentangle how geography, mobility, and physical isolation shape dialect variation across languages. This research indicated that while geographic distance dominates Japanese dialectal variation, population mobility plays a decisive role in UK-Ireland English by weakening geographic dependence, with physical isolation exhibiting a subtle trend. This research shows that SSL models could also provide an insight for contemporary dialectology, connecting acoustic dynamics to the social contexts.
language of the presentation: English
 
中畔 彪雅 D, 中間発表 自然言語処理学(ロボット対話知能) 渡辺 太郎☆, 吉野 幸一郎, Angel Garcia Contreras
title: Integrating Paralinguistic Information into Language Models via Modulation
abstract: In spoken dialogue, paralinguistic information such as emotion and prosody can change the meaning of the same text, so how this information is integrated into a language model is a key question. Most existing approaches integrate speech features by concatenating them with text. In contrast, this work hypothesizes that *modulating* linguistic representations, rather than concatenating features, is the effective form of integration, and tests this hypothesis using Feature-wise Linear Modulation (FiLM). In RQ1, we define dialogue failures caused by emotion as "emotional dialogue breakdown" and construct the paraling-dial dataset to evaluate this phenomenon in isolation. While existing speech language models perform near-chance level, our FiLM-based modulation model detects such breakdowns with high accuracy, demonstrating the effectiveness of modulation. In RQ2, we are currently investigating a method that aggregates speech information into a single vector and applies FiLM-based modulation, enabling more efficient processing than concatenation-based methods without increasing the number of speech tokens. This talk reports both studies under a unified framework in which paralinguistic information is integrated as modulation.
タイトル:パラ言語情報の変調による言語モデルへの統合
概要:音声対話では,感情や韻律などのパラ言語情報が同一テキストの意味を変えるため,その言語モデルへの統合が重要である.従来手法の多くは音声特徴をテキストに連結して統合するが,本研究では連結ではなく言語表現を「変調」する統合が有効という仮説を立て,Feature-wise Linear Modulation(FiLM)を用いて検証する.RQ1では,感情に起因する対話の失敗を「感情的対話破綻」と定義し,その分離評価のためデータセットparaling-dialを構築した.既存の音声言語モデルがほぼチャンスレベルに留まる一方,FiLMによる変調モデルは高精度に検出でき,変調の有効性を示した.RQ2では,音声情報を単一ベクトルに集約してFiLMで変調することで,追加音声トークンを増やさず連結型手法より効率的に処理する手法を検討中である.本発表では,これらをパラ言語情報を変調として統合する一貫した枠組みのもとで報告する.