AI & Robotics Engineer specializing in Multimodal Perception & Human–Robot Interaction.
I design and deploy intelligent systems that fuse vision, audio, and context to enable machines to understand who is speaking, how humans feel, and when to act — from research prototypes to real-time robotic applications.
From micro-expression analysis to social robotics and edge hardware deployment.
Real-time systems that combine audio cues (voice activity, pitch) and visual signals (lip motion, gaze) to accurately identify the active speaker in multi-person environments.
Developed perception pipelines for Pepper robots enabling:
Context-aware diarization using acoustic features and BERT embeddings to handle overlapping speech in real-world conversations.
From micro-expressions to fatigue detection: CNN-LSTM architectures, facial cues, blink rate, and real-time vision pipelines.
PyTorch, OpenCV, MediaPipe, NVIDIA Jetson, AWS Panorama. Edge-ready, low-latency systems bridging theory and hardware.
Detailed research implementations in computer vision, robotics perception, and multimodal learning.
Built a real-time system that fuses audio cues and visual signals to identify who is speaking in multi-person conversations, improving robustness in noisy and dynamic environments such as social robots and smart meetings.
Developed perception pipelines for a Pepper robot to recognize basic human emotions and respond with context-aware behavior, enabling more natural and socially aware human–robot interaction.
Designed a speaker diarization pipeline combining acoustic features with language embeddings to accurately determine who spoke when, even in overlapping and multi-speaker conversations.
Built a real-time fatigue detection system by fusing visual indicators (blink rate, yawning) and speech patterns to help collaborative robots adapt task pace and improve human safety and comfort.
Developed a system that translates sign language gestures into speech using hand keypoint detection and sequence models, enabling more accessible communication between humans and robots.
Built a deep learning system to detect deception by analyzing micro-expressions in video streams, focusing on subtle and involuntary facial movements relevant to psychology and security research.
Run commands directly or click the preset chips below to explore Amir's bio, skills, and background.
Technologies, frameworks, and hardware utilized across research and production.
Whether you're interested in AI/Robotics research, industrial collaboration, or consulting, feel free to reach out.