← All topics

Advanced Topics

Multimodal AI

マルチモーダルAI

Multimodal AI processes more than one modality, such as text, images, audio, or video. A vision-language model may use a multimodal embedding to align visual and linguistic information.

Japanese terms

  1. Multimodal AI — マルチモーダルAI: AI that processes or combines more than one kind of data, such as text, images, audio, or video.
  2. Modality — モダリティ: A particular form or channel of information, such as text, vision, or audio.
  3. Vision-language model — 視覚言語しかくげんごモデル: A model trained to represent and reason across visual and linguistic information.
  4. Multimodal embedding — マルチモーダルみ: A representation that places information from different modalities in a shared or aligned vector space.

Computer vision and natural language processing supply important modality-specific foundations.