| アイテムタイプ |
Article |
| ID |
|
| プレビュー |
| 画像 |
|
| キャプション |
|
|
| 本文 |
KAKEN_20H04269seika.pdf
| Type |
:application/pdf |
Download
|
| Size |
:691.8 KB
|
| Last updated |
:Dec 11, 2024 |
| Downloads |
: 493 |
Total downloads since Dec 11, 2024 : 493
|
|
| 本文公開日 |
|
| タイトル |
| タイトル |
マルチモーダル言語理解における敵対的データ拡張基盤の構築
|
| カナ |
マルチモーダル ゲンゴ リカイ ニ オケル テキタイテキ データ カクチョウ キバン ノ コウチク
|
| ローマ字 |
Maruchimōdaru gengo rikai ni okeru tekitaiteki dēta kakuchō kiban no kōchiku
|
|
| 別タイトル |
| 名前 |
Adversarial data augmentation for multimodal language understanding
|
| カナ |
|
| ローマ字 |
|
|
| 著者 |
| 名前 |
杉浦, 孔明
 |
| カナ |
スギウラ, コウメイ
|
| ローマ字 |
Sugiura, Komei
|
| 所属 |
慶應義塾大学・理工学部 (矢上) ・教授
|
| 所属(翻訳) |
|
| 役割 |
Research team head
|
| 外部リンク |
科研費研究者番号 : 60470473
|
|
| 版 |
|
| 出版地 |
|
| 出版者 |
|
| 日付 |
| 出版年(from:yyyy) |
2023
|
| 出版年(to:yyyy) |
|
| 作成日(yyyy-mm-dd) |
|
| 更新日(yyyy-mm-dd) |
|
| 記録日(yyyy-mm-dd) |
|
|
| 形態 |
|
| 上位タイトル |
| 名前 |
科学研究費補助金研究成果報告書
|
| 翻訳 |
|
| 巻 |
|
| 号 |
|
| 年 |
2022
|
| 月 |
|
| 開始ページ |
|
| 終了ページ |
|
|
| ISSN |
|
| ISBN |
|
| DOI |
|
| URI |
|
| JaLCDOI |
|
| NII論文ID |
|
| 医中誌ID |
|
| その他ID |
|
| 博士論文情報 |
| 学位授与番号 |
|
| 学位授与年月日 |
|
| 学位名 |
|
| 学位授与機関 |
|
|
| 抄録 |
本研究では,マルチモーダル言語理解,マルチモーダル言語生成,Sim2Real転移学習と介助犬タスクでの評価を行った.理解班では,Vision-and-Language Navigationタスクにおいて,敵対的摂動更新アルゴリズム Momentum-based Adversarial Trainingを構築した.生成班では,動画から将来の状況を説明する文を生成するfuture captioning手法を構築し,既存手法を上回る結果を得た.Sim2Real班では,生活支援ロボット評価フレームワークを構築し,指示文生成とタスク実行を自動化した.
In this study, our objectives are (a) robust multimodal language understanding through adversarial data augmentation, (b) multimodal language generation, and (c) evaluation in the assistance dog tasks.
We first focused on the Vision-and-Language Navigation task and developed the Momentum-based Adversarial Training (MAT) algorithm. We applied MAT to the standard benchmark test, ALFRED, and obtained successful results. We also worked on the task of generating descriptions about future situations. The main novelty of our proposed method lies in the use of Relational Self-Attention as the attention mechanism. Experimental results show that our method outperformed existing methods in standard metrics. We applied the multimodal language understanding and generation methods into a simulator, enabling on-the-fly instruction generation. As a result, we established a robot evaluation framework that does not require manual intervention in task generation, execution, and evaluation.
|
|
| 目次 |
|
| キーワード |
|
| NDC |
|
| 注記 |
研究種目 : 基盤研究 (B) (一般)
研究期間 : 2020~2022
課題番号 : 20H04269
研究分野 : 機械知能, 知能ロボティクス, マルチモーダル言語処理
|
|
| 言語 |
|
| 資源タイプ |
|
| ジャンル |
|
| 著者版フラグ |
|
| 関連DOI |
|
| アクセス条件 |
|
| 最終更新日 |
|
| 作成日 |
|
| 所有者 |
|
| 更新履歴 |
|
| インデックス |
|
| 関連アイテム |
|