AI Researcher

Kiana Hooshanfar

Computer Vision · Saliency Prediction · Large Language Models
M.Sc., University of Tehran · ML & Computational Modeling Lab · AVIR AI
Kiana Hooshanfar

About

Hi! I'm Kiana - an AI researcher and developer working on computer vision, saliency, and generative models. I hold a B.Sc. from IUST and recently defended my M.Sc. at the University of Tehran. I collaborate with the Machine Learning & Computational Modeling Lab and build real-world AI systems at AVIR AI.

My research focuses on visual attention modeling, multimodal learning, and human-computer interaction. I'm always happy to discuss research ideas and collaborations.

Research Interests

Computer Vision
Multimodal Learning
Saliency Prediction
Generative Models
Vision-Language Models
Human-Computer Interaction
Large Language Models
AI Agents
Deep Learning

Education

M.Sc. in Electrical Engineering (Control Engineering)
University of Tehran · Tehran, Iran · 2022 – 2025
  • Thesis: Design and Enhancement of Video Saliency Prediction Networks
  • Advisors: Dr. Babak Nadjar Araabi, Dr. Ahmad Kalhor
  • Core courses: Machine Learning, Deep Learning, Deep Generative Models, Game Theory
B.Sc. in Electrical Engineering (Control Engineering)
Iran University of Science & Technology (IUST) · Tehran, Iran · 2018 – 2022
  • Thesis: Intelligent Control of a Robot Arm by Eye Tracking for People with Severe Speech and Motor Impairment (SSMI)

News

Publications

OpenVAM architecture 2026
OpenVAM: Open-World Visual Attention Modeling with VLMs
Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati
Preprint, 2026.
Saliency PredictionVision-Language ModelsMultimodal Learning
DTFSal architecture BMVC'25
DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
Kiana Hooshanfar, Alireza Hosseini, Mona Ahmadian, Ahmad Kalhor, Babak N. Araabi
36th British Machine Vision Conference (BMVC), 2025.
Saliency PredictionComputer VisionMultimodal Learning
Brand visibility schematic Discover'25
Brand Visibility in Packaging: A Deep Learning Approach for Logo Detection, Saliency-Map Prediction, and Logo Placement Analysis
Alireza Hosseini, Kiana Hooshanfar, Pouria Omrani, Reza Toosi, Ramin Toosi, Zahra Ebrahimian, Mohammad Ali Akhaee
Discover Applied Sciences, 2025.
Saliency PredictionComputer VisionDeep Learning
Hybrid RAG architecture ICWR'24
Hybrid Retrieval-Augmented Generation Approach for LLMs Query Response Enhancement
Pouria Omrani, Alireza Hosseini, Kiana Hooshanfar, Zahra Ebrahimian, Ramin Toosi, Mohammad Ali Akhaee
International Conference on Web Research (ICWR), 2024.
Large Language ModelsRAG
Eye-tracking robotics ICRoM'23
Eye-Tracking Based Control of a Robotic Arm and Wheelchair for People with Severe Speech and Motor Impairment
Maryam Asad Samani, Kiana Hooshanfar, Helia Shams Jey, Seyed Majid Esmailzadeh
International Conference on Robotics and Mechatronics (ICRoM), 2023.
Human-Computer InteractionDeep Learning

Projects

Project… Multi-AgentRef: SYS/01

A multi-agent LLM platform: intent routing, tool calling, and auditable workflows, plus a Weaviate RAG pipeline with hybrid retrieval, BGE-M3, and layered safety guardrails.

ADK / Weaviate
Project… YaraRef: DEV/06

A docs support assistant: hybrid RAG for how-to questions, and a multi-step troubleshooting agent when something is broken — it refuses to guess when evidence is insufficient.

ADK / Qdrant Code ↗
Project… OpenVAMRef: PRJ/01

Open-world visual attention with VLMs — predicts where people look, plus the what and why behind it, across natural images, ads, and UIs.

Saliency / VLMs Project ↗ Code
Project… DTFSalRef: PRJ/02

Audio-visual saliency for video: dynamic token fusion predicts where attention lands from what you see and hear — state of the art on six benchmarks.

BMVC 2025 Project ↗ Paper
Project… Brand AttentionRef: PRJ/03

Where do eyes land on a shelf? Logo detection plus packaging-tuned saliency, distilled into a single brand-visibility score.

Discover App. Sci. 2025 Paper ↗ Code
Project… Doc ParserRef: DEV/01

Engineering drawings, structured. YOLO finds the title block, auto-rotates and crops it, then Qwen2.5-VL reads out fields like scale and drawing number.

YOLO / Qwen2.5-VL Code ↗
Project… Influencer AIRef: DEV/04

Turn a few parameters into social content — 15+ formats from hashtags to captions, plus GPT-4o vision for images, in two steps.

OpenAI / GPT-4o Code ↗
Project… Doc IntelligenceRef: DEV/03

GPU OCR, table extraction, and document classification from PDFs and images — Surya OCR plus Qwen, served through a Gradio UI and a REST API.

Surya / vLLM Code ↗
Project… OCR BotRef: DEV/02

Send a document photo to Telegram, get back clean Markdown, LaTeX, or DOCX. Gemini OCR on a FastAPI backend with auth and credits.

Gemini / FastAPI Code ↗

Honors & Awards

Teaching Experience

Machine Learning & AI Control Systems Foundations Laboratory
Teaching Assistant 8 courses
University of Tehran · Tehran, Iran · 2023 – 2025
Machine Learning Neural Networks and Deep Learning Analysis and Design of Deep Neural Networks Linear Algebra System Identification Linear Control Systems Linear Control Systems Laboratory Digital Control Laboratory
Teaching Assistant 6 courses
Iran University of Science & Technology (IUST) · Tehran, Iran · 2021 – 2022
Principles of Mechatronics Signals and Systems Electrical Circuits II Linear Algebra Python Technical Language
Ask about my résumé AI-powered assistant