Skip to main content

AI initiative registry

Every initiative listed in the registry can be consulted here, without an account. Each record shows its progress, its milestones and the history of its decisions. Requests still under triage do not appear: they join the registry once reviewed.

Perception and Understanding

AI that interprets the world, whether through text, images, or sound.

What is “Perception and Understanding”?

This category includes AI systems that sense, interpret, or perceive data from the environment, emulating human senses like sight, sound, touch, smell, and language comprehension. These capabilities form the foundation of AI awareness, enabling systems to extract meaning from unstructured inputs like text, images, video, audio, or environmental signals.

Perception and Understanding
Subcategory Description Examples
Language understanding — Natural Language Processing (NLP) Understanding and generating human language Chatbots, document summarization
Computer vision Understanding images, video, and spatial data Object detection, satellite imagery analysis, species ID
Speech recognition Converting audio to text Voice-to-text for transcription, voice-based commands for virtual assistants
Multimodal AI Combining input types (text + images, etc.) Image captioning, document understanding tools
Olfactory AI (artificial smell recognition) Detecting and interpreting smells using e-noses Gas leak detection, spoilage sensing, pollutant ID in fish labs
Haptic AI (artificial touch sensing) Interpreting touch and tactile feedback Robotic sampling, pressure detection, texture, temperature
Audio event detection Recognizing non-speech sounds Marine mammal calls (whale song), mechanical alerts, engine noise anomaly detection
Sensor fusion AI Combining multiple sensory inputs Autonomous underwater vehicle navigation, habitat monitoring
Key indicator
If the AI’s primary task is to interpret raw input (e.g. image, text, sound, sensor data), this is where it belongs.
Common input types
  • Visual data (images, video)
  • Audio and speech
  • Natural language text
  • Sensor readings (e.g. haptics, chemical, temperature)
Filter 1
Reset

71 current initiative(s)