Manabu OKUMURA
Professor. Natural language processing, intelligent information presentation, language learning support, text mining
members
Professor. Natural language processing, intelligent information presentation, language learning support, text mining
Associate Professor. Natural language processing, multimodal dialogue systems, human-machine interaction
Assistant Professor. Computer vision, image recognition
Administrative Assistant
Specially Appointed Researcher
Multimodal summarization
Medical Image Analysis, Multimodal Learning, Vision-Language Models
Multimodal AI, AI Agents
Multi-Document Summarization, Graph-based learning, Recommendation
Multimodal Personality Traits Recognition, Mutlimodal Machine Translation
Low-Resource Machine Translation
Multimodal Dialogue System, Human-Machine Interaction
My research focuses on information extraction from materials science literature using large language models. I am currently investigating hallucinations in LLM-based information extraction and methods for their evaluation and mitigation.
My main research interest is training machanism of language models, especially post-training mechanisms towards trustworthy and controllable knowledge manipulation and reasoning behavior for LLMs.
Machine translation, natural language generation, multimodal
Intersection of quantitative finance and artificial intelligence
Noise-Robust Audio-Visual Speech Recognition Using Lip-Reading Video
Semantics, Information Retrieval, Vision-Language Model
Multimodal Learning, LLM Agents, Causal Inference
LLM Agents, Long-term Memory, Persona Consistency, and Reinforcement Learning
Slide-aware Presentation ASR
Development of a proactive dialogue system that imitates speech-language-hearing therapist
Agent-based Semi-Automatic Sign Language Annotation
Information Retieval, LLM Agents
Vision-Language Navigation
MLLM-Powered Agent, Vision-Language Model
Machine Translation, Efficient MBR Decoding
Simultaneous generation of language and action. I work on externalizing a model's reasoning into a visual workspace to improve long-range and spatial consistency in generation.
LLM, RAG, Embedding Model
Designing AI Agent Interaction Across Private and Shared Spaces in Multi-Party Collaboration
Information Retrieval
Visual encoder specialized for human interaction
No members are listed yet.