Advanced Topics
Multimodal AI
マルチモーダルAI
Multimodal AI processes more than one modality, such as text, images, audio, or video. A vision-language model may use a multimodal embedding to align visual and linguistic information.
Japanese terms
- Multimodal AI — マルチモーダルAI: AI that processes or combines more than one kind of data, such as text, images, audio, or video.
- Modality — モダリティ: A particular form or channel of information, such as text, vision, or audio.
- Vision-language model — 視覚言語モデル: A model trained to represent and reason across visual and linguistic information.
- Multimodal embedding — マルチモーダル埋め込み: A representation that places information from different modalities in a shared or aligned vector space.
Related topics
Computer vision and natural language processing supply important modality-specific foundations.