Research

Machine learning on human signals — face, eye, gait, voice.

Mental health sensing

Care that notices before you have to ask for it.

Mental illness is a leading cause of disability worldwide, affecting an estimated 450 million people. Up to one in five of us will live through major depressive disorder or bipolar disorder. It usually arrives in our early twenties, and it is a meaningful predictor of suicide.

Against that, the instruments are startlingly blunt. A clinician asks how the last two weeks felt. The person answers from memory, in a room they had to travel to. Whatever happened on a particular Tuesday, the day it got bad or the hour it lifted, is gone by the time anyone asks. We are trying to catch something that moves hour to hour with a question that averages a fortnight.

Closing that gap is what this line of work is for. It runs in three directions.

Making the signal collectable at all

FacePsy running in the wild: bright indoor, night, daylight, and mid-movement
FacePsy running in the wild: bright indoor, night, daylight, and mid-movement

Depression leaves marks nobody chooses to make: in facial muscle activity, in how the head moves, in how often the eyes are open, in the shape of a smile, in how the pupil answers the prospect of something good. These facial behavior primitives surface whether or not a person is ready to talk, which is precisely why they matter, because those least able to ask for help are so often the ones who need it most.

None of this is news in a laboratory. What has kept it there is not the science but the bargain it demands. The obvious way to read a face all day is to stream video off someone's phone, and that asks a person to surrender the most intimate data they have, continuously, in exchange for a maybe.

So the first direction is a refusal of that bargain. FacePsy [J.5] and PupilSense [C.2] wake on things a person is already doing, unlocking the phone or opening an app, then read the face or the eye for a moment, extract only geometry and gaze, and destroy the image on the device within seconds. Nothing is stored. Nothing is transmitted. The work is open source, so the claim is inspectable rather than asserted.

FacePsy pipeline: a phone-unlock trigger starts a 10-second on-device capture, facial behavior primitives are extracted, and only the derived features leave the device for the research server
FacePsy pipeline: a phone-unlock trigger starts a 10-second on-device capture, facial behavior primitives are extracted, and only the derived features leave the device for the research server

The door this opens is consent. Sensing that a person would willingly live with is worth more than sensing that is marginally more accurate and switched off in a week.

Reading the whole person, and reading forward

PupilSense: iris and pupil located in everyday photos
PupilSense: iris and pupil located in everyday photos

The second direction is about what the signal is for.

Face and eye turn out to be complementary rather than redundant. Expression carries one part of affect and the pupil carries another, and MoodPupilar [C.3] showed the eye alone speaks to ordinary daily mood, not only clinical episodes but the everyday rise and fall that episodes emerge from. Read together [J.5] [C.2] [C.3], they describe a person more completely than any one channel does alone.

They also revealed something that constrains every system built after them: depression does not wear a single face. The signal differs between people, and a model that assumes otherwise fails hardest on whoever it resembles least. Personalisation is not a refinement here; it is a requirement.

Then MoodCam [C.5] moved the question from the present to the future: not how someone feels now, but how they will feel tomorrow. That shift is the difference between a diary and an intervention. A system that sees a difficult day coming can put something useful in front of a person before it arrives: a check-in, a reminder of what helped last time, a nudge toward someone who cares.

Building something the field can stand on

PupilSense: raw frame and detected pupil
PupilSense: raw frame and detected pupil

The third direction is refusing to keep any of it.

A single lab studying a small group is a beginning, never a conclusion. So the data behind FacePsy became the BeyondSmile Challenge [J.6], and teams worldwide were invited to beat it. They did. The winning entry surpassed the original work outright. That is the outcome worth wanting: the method stops belonging to one lab, the bar ends up higher than any one group would have set it, and the next person starts from the finish line rather than the start.

The same instinct runs into what these systems should do. A phone that recognises distress and stays silent is a surveillance device; the value only appears when sensing becomes care. One strand of this work [W.1] frames conversational agents that draw on affect, physiology and context alongside established therapeutic technique, so that support is offered rather than waited for. The other [Po.3] asked whether it changes anything, putting a chatbot that only reads words against one that also reads the face. Being seen rather than only read measurably changed how much people trusted it and how understood they felt.

That is a result about people rather than models. Empathy that responds to the whole person lands differently, even coming from software.

The future this builds

Every instrument needed to notice a mental health crisis is already in nearly every pocket on earth. Nothing has to be manufactured, purchased or worn. What has been missing is systems designed for a life rather than for a study: private by construction, light enough to run all day, personal enough to fit one person rather than an average, and honest enough that someone would choose to keep them switched on.

Built that way, the shape of care changes. Assessment stops being an appointment and becomes something continuous and quiet. Support arrives on the bad day rather than three weeks after it. Clinicians see the months between visits instead of guessing at them. Research moves at the pace of a shared dataset rather than a single lab. And the person carrying all of it never has to hand over their face to receive it.

That is the work: not better classifiers, but a version of mental health care that reaches people where they already are.

The papers

  1. [J.6] Summary of the BeyondSmile Challenge on Detecting Depression Through Facial Behavior and Head Gestures

    Rahul Islam, Tongze Zhang, Anlan Dong, Melik Ozolcer, Christina Garcia, Milyun Ni'ma Shoumi, Prerna Jamloki, Sozo Inoue, Sang Won Bae

    International Journal of Activity and Behavior Computing (IJABC 2025)

  2. [C.5] MoodCam: Mood Prediction Through Smartphone-Based Facial Affect Analysis in Real-World Settings

    Rahul Islam, Tongze Zhang, Sang Won Bae

    IEEE International Conference on Ubiquitous Intelligence and Computing (UIC 2024)

  3. [C.3] MoodPupilar: Predicting Mood Through Smartphone Detected Pupillary Responses in Naturalistic Settings

    Rahul Islam, Tongze Zhang, Priyanshu Singh Bisen, Sang Won Bae

    IEEE-EMBS International Conference on Wearable and Implantable Body Sensor Networks (BSN 2024)

  4. [J.5] FacePsy: An Open-Source Affective Mobile Sensing System - Analyzing Facial Behavior and Head Gesture for Depression Detection in Naturalistic Settings

    Rahul Islam, Sang Won Bae

    Proceedings of the ACM on Human-Computer Interaction (MobileHCI '24)

  5. [C.2] PupilSense: Detection of Depressive Episodes Through Pupillary Response in the Wild

    Rahul Islam, Sang Won Bae

    Activity and Behavior Computing 2024

  6. [W.1] Revolutionizing Mental Health Support: An Innovative Affective Mobile Framework for Dynamic, Proactive, and Context-Adaptive Conversational Agents

    Rahul Islam, Sang Won Bae

    ACM IMWUT/Ubicomp 2023, GenAI4PC Symposium

  7. [Po.3] Towards Designing Empathetic and Trustworthy AI Chatbots: an Exploratory Study

    Melik Ozolcer, Rahul Islam, Abdullah Mohammed, Tongze Zhang, Sang Won Bae, Ting Liao

    CESUN 2023

Substance use sensing

Knowing someone is impaired, in the moment it matters, without asking them.

There is no breathalyser for cannabis. Blood, urine and saliva can show THC for days after use, long past any impairment, so a positive test says almost nothing about whether a person is affected right now. Alcohol has a reliable chemical test, but by the time it reads positive the drinking has already happened.

Both failures are the same failure. Our tests measure the presence of a substance, and what actually matters to a clinician, to a driver, or to someone trying to cut down is a state: is this person impaired, and is this the moment to say something. Binge drinking alone, defined as four or more drinks for women or five for men on one occasion, carries consequences that arrive within hours: traffic fatalities, injuries, violence.

Closing the gap between substance and state is what this work is about. It runs in two directions.

Reading the state instead of the substance

The first direction abandons chemistry and reads behaviour and physiology instead.

Cannabis intoxication is visible in how a person moves and travels, and in how their body responds. Heart rate rises sharply within minutes of smoking and falls again, a physiological trace of the episode itself. Mobile phone sensor-based detection of subjective cannabis intoxication [J.3] established that everyday phone signals, movement from the accelerometer and travel from GPS, track the subjective experience of being high in daily life, not in a laboratory. The follow-on work [J.7] added a consumer wearable to the phone, pairing behaviour with physiology, and found that the two together describe the state far better than either alone.

The word that matters in both is subjective. These systems are not looking for a molecule. They are looking for the state a person would recognise in themselves, which is the only thing an intervention can usefully respond to.

The door this opens is a measure that works in real time, requires no lab, no blood draw and no roadside device, and runs on hardware people already own.

Getting ahead of the event

The second direction is a shift in tense.

Detection tells you drinking has started. That is useful for a diary and nearly useless for prevention: the decision it might have influenced is already several drinks behind you. So the question became whether the hours before an event look different from ordinary ones.

They do. Predicting imminent same-day binge-drinking events [J.4] showed that a heavy drinking episode is visible in phone data hours before the first drink, from where a person has been and when: the rhythm of a day bending toward a particular evening. Related work on alcohol interventions [Po.4] takes the same location signal and asks how it should shape the support that gets delivered.

That earlier work on cannabis [J.3] found the same thing from another angle: simple routine, day of the week and time of day, already predicts use, because substance use lives inside habit rather than outside it.

Prediction is what makes just-in-time support possible at all. A message that lands during a difficult hour, before a decision rather than after it, is an entirely different intervention from a summary read a week later.

The future this builds

Substance use research has been shaped by what could be measured: a substance in a sample, at a clinic, after the fact. Everything downstream inherited that shape: treatment as episodic, relapse understood in retrospect, help arriving after the harm.

Sensing that reads state instead of substance, and anticipates rather than records, changes what care is able to do. Support can arrive in the hour it matters instead of at the next appointment. Impairment can be assessed where it actually matters, before someone drives, without a test that cannot tell last Tuesday from right now.

That is the future worth building: not better surveillance of people who use substances, but timely support that reaches them at the only moment it can still change the outcome.

The papers

  1. [J.7] Enhancing Interpretable, Transparent, and Unobtrusive Detection of Acute Marijuana Intoxication in Natural Environments: Harnessing Smart Devices and Explainable AI to Empower Just-In-Time Adaptive Interventions: Longitudinal Observational Study

    Sang Won Bae, Tammy Chung, Tongze Zhang, Anind K Dey, Rahul Islam

    JMIR AI 2025

  2. [Po.4] Moving Toward Personalized Behavioral Medicine: Integrating Smartphone-based GPS Data into a Digital Alcohol Intervention

    Tammy Chung, Sang Won Bae, Tongze Zhang, Melik Ozolcer, Rahul Islam, Anind Dey, Yiyi Ren, Brian Suffoletto, Aidan GC Wright, Trishnee Bhurosy

    Society of Behavioral Medicine/Annals of Behavioral Medicine 2024

  3. [J.4] Leveraging Mobile Phone Sensors, Machine Learning, and Explainable Artificial Intelligence to Predict Imminent Same-Day Binge-drinking Events to Support Just-in-time Adaptive Interventions: Algorithm Development and Validation Study

    Sang Won Bae, Brian Suffoletto, Tongze Zhang, Tammy Chung, Melik Ozolcer, Rahul Islam, Anind K Dey

    JMIR Formative Research 2023

  4. [J.3] Mobile phone sensor-based detection of subjective cannabis intoxication in young adults: A feasibility study in real-world settings

    Sang Won Bae, Tammy Chung, Rahul Islam, Brian Suffoletto, Jiameng Du, Serim Jang, Yuuki Nishiyama, Raghu Mulukutla, Anind K Dey

    Drug and Alcohol Dependence 2021

Clinical movement

Gait and pose analysis for cerebral palsy and cognitive decline, run on a phone.

  1. [J.9] AIGaitor: Privacy-preserving and cloud-free motion analysis for everyone, using edge computingTo appear

    Lauhitya Reddy, Rahul Islam, Trisha M. Kesar, Hyeokhyen Kwon

    PLOS ONE 2026 (Under Review)

  2. [J.8] Wearable Sensing for Quantifying Cognitive and Balance Functions in Naturalistic Movements of Older Adults with Mild Cognitive Impairment in Therapeutic EnvironmentsTo appear

    Rahul Islam*, Jangwon Lim*, Dharini Raghavan, Bolaji Omofojoye, Amy D. Rodriguez, Yashar Kiarashi, Rachel Hershenberg, Gari D. Clifford, Hyeokhyen Kwon

    PLOS Digital Health 2026 (* = Equal Contribution, Under Review)

  3. [C.6] Clinically Accessible 2D Video Analysis Accurately Captures 3D Knee Gait KinematicsTo appear

    Jeremy Bauer, Susan Sienko, Lauhitya Reddy, Vedant Kulkarni, Seth Donahue, Ross Chafetz, Anita Bagley, Joseph Krzak, Rahul Islam, Hyeokhyen Kwon

    Annual Meeting of the American Academy for Cerebral Palsy and Developmental Medicine (AACPDM 2026, Under Review)

  4. [Po.5] Evaluating the Feasibility of Deploying Quantized Human Pose Estimation on Smartphone for Gait Analysis for Children with Cerebral Palsy

    Rahul Islam, Seth Donahue, Jeremy Bauer, Susan Sienko, Anita Bagley, Joseph Krzak, Maura Eveld, Karen Kruger, Ross Chafetz, Vedant Kulkarni, Hyeokhyen Kwon

    Annual Meeting of the Gait & Clinical Movement Analysis Society (GCMAS 2026)

Eye & biometrics

The most revealing surface on the body, read by a camera that is already pointed at it.

The eye gives away more than any other part of us that can be seen without touching. The iris carries a pattern unique to a person and stable for life. The sclera and the skin around it, the periocular region, carry more. Pupils widen and narrow to light, to effort, to feeling, without anyone deciding they should. It is the densest signal on the human body, and for most of the day it is pointed directly at a camera.

For decades that fact was largely unusable. Iris recognition meant dedicated infrared hardware, controlled illumination and a person deliberately presenting themselves to a scanner. Eye tracking meant a specialised rig costing hundreds or thousands. The signal was rich; the equipment kept it in the laboratory and at the border.

This work is about what becomes possible once the equipment is just a phone.

Identity in ordinary light

The first direction moved ocular biometrics out of the infrared and into the visible spectrum: the light a normal camera already sees.

That sounds like a small change and is not. Infrared gives a clean, high-contrast iris under controlled conditions. Ordinary light gives you reflections, shadow, whatever colour the room happens to be, and an iris partly hidden by lashes and lids. A preliminary study of CNNs for iris and periocular verification [C.1] and the journal work that grew from it [J.1] asked whether learned representations could hold up under those conditions, using images captured on the phones people actually carry.

They could. And the more interesting finding was about the hardware itself: verification behaves differently when two images come from the same model of phone than when they come from different ones. Every camera leaves its own signature, and a biometric system that ignores that is quietly relying on it. Naming that constraint matters more for real deployment than any single result, because in the real world nobody guarantees both photographs came from the same device.

The door this opens is authentication that needs no scanner, no infrared, no dedicated sensor, only the camera in the phone already in someone's hand.

The eye as an input, not just an identifier

EyeSpyVR: what the phone camera sees inside the headset
EyeSpyVR: what the phone camera sees inside the headset

The second direction turned the same camera outward as an interface.

Cheap smartphone VR headsets put a phone a few centimetres from someone's face, with the front camera aimed squarely at one eye, a sensing opportunity that existed in millions of devices and was going entirely unused. EyeSpyVR [J.2] took it, using the phone's camera as the sensor and its own screen as the illuminator, with no added hardware at all.

From that single unglamorous view, through a plastic lens and under shifting screen light, came four distinct capabilities: whether the headset is being worn, when the wearer blinks, who the wearer is, and roughly where they are looking. Capabilities that otherwise required a purpose-built eye tracker arrived instead as a software update.

That is the pattern worth noticing across both directions: the sensor was already there. What was missing was the willingness to work with a difficult signal instead of demanding a clean one.

Where the eye leads

A commodity smartphone VR headset
A commodity smartphone VR headset

There is a thread running from this work into everything after it. The same camera that can verify who someone is can also read how they are: pupils responding to emotional weight, to reward, to fatigue. That is exactly what the later mental health work does with the eye [C.2] [C.3].

Which makes this the place to be honest about the double edge. A phone that can identify a person from their eye in ordinary light is also a phone that can do it without their knowledge. The same properties that make ocular sensing accessible (no special hardware, works at a distance, needs no cooperation) are what make it worth being careful with. That is not an argument against building it. It is an argument for building it deliberately, and for the people doing the building to be the ones thinking hardest about consent.

The future this points to is quiet: authentication that stops being a ritual, interfaces that know where you are looking without a rig, and health signals read from a glance, all from hardware people already own, and all designed so the person looking into the camera is the one it answers to.

The papers

  1. [J.2] EyeSpyVR: Interactive Eye Sensing Using Off-the-Shelf, Smartphone-Based VR Headsets

    Karan Ahuja, Rahul Islam, Varun Parashar, Kuntal Dey, Chris Harrison, Mayank Goel

    ACM IMWUT/Ubicomp 2018

  2. [J.1] Convolutional neural networks for ocular smartphone-based biometrics

    Karan Ahuja, Rahul Islam, Ferdous Barbhuiya, Kuntal Dey

    Pattern Recognition Letters, Volume 91. 2017

  3. [C.1] A Preliminary Study of CNNs for Iris and Periocular Verification in the Visible Spectrum

    Karan Ahuja, Rahul Islam, Ferdous Barbhuiya, Kuntal Dey

    ICPR 2016

HCI & systems

Making demanding sensing cheap enough to leave running.

A capability that only works on a workstation is a demonstration. The same capability running on a mid-range phone, all day, without draining the battery is a product, and the distance between those two things is where most sensing research quietly dies.

That gap is what this strand of work is about. Not new signals, but making known signals affordable enough that an ordinary device can carry them, and shaping them into interactions people never have to be taught.

Finding the cheapest honest representation

MicroFlow: face geometry to affective state
MicroFlow: face geometry to affective state

The instinct in machine perception is to reach for more: more resolution, more parameters, more compute. Often the better move is a representation that throws away everything except what carries the signal.

SenTion [A.1] took that route with facial expression. Rather than treat the face as pixels to be learned, it described it as geometry: the angles between facial landmarks, which it called Inter Vector Angles. Angles are indifferent to how large a face appears and to whose face it is, so the representation transfers across people and distances without retraining. Combined with appearance features, it ran in real time on a laptop CPU with no GPU at all, at low resolution.

That was 2016, and the choice kept paying. Seven years later MicroFlow [Po.2] reached for the same features to do something quite different: reading micro-expressions, the involuntary flickers lasting under half a second that are far harder to suppress than an ordinary expression, to detect whether a student was in flow, bored or anxious while working through a programming assignment. The goal was educational rather than clinical: give an instructor a way to see engagement, and pitch difficulty to the person in front of them.

A representation good enough to survive that jump, from expression recognition to affect in a classroom, was a better investment than any amount of extra compute would have been.

Spending the right resource

MotionTrace: predicted hand trajectories
MotionTrace: predicted hand trajectories

The second direction is systems thinking about what a capability costs to run.

Smartphone augmented reality needs to know precisely where the phone is in space, and the obvious way to get that is the camera. But the camera is the most expensive sensor on the device: it drains the battery, heats the phone until performance throttles, and fails exactly when the light is poor or the view is blocked.

MotionTrace [C.4] asked whether the inertial sensors could carry that load instead. The IMU costs almost nothing to run, works in the dark, and does not care what is in front of it. Predicting where the phone is heading, rather than continuously measuring where it is, lets the system prepare what the wearer is about to look at, so content is ready before the view arrives instead of loading after it.

The gain here is not accuracy, it is endurance. A tracking method the device can afford to keep running changes what an AR application can be, in a way that a more accurate method it must ration does not.

Interfaces built on what people already do

EyamKayo: the CAPTCHA asks for an expression
EyamKayo: the CAPTCHA asks for an expression

The third direction is about the interaction itself: the best input is something a person does naturally and a machine finds hard to fake.

EyamKayo [Po.1] applied that to a familiar irritation. CAPTCHAs ask humans to prove humanity by doing something machines are good at: reading distorted text, spotting objects in photographs. That is why they keep falling. So instead it asks a person to look where they are told and to produce a sequence of expressions. Easy, wordless, and nothing to decipher; for a bot, it means convincingly presenting a human face in three dimensions and animating it on demand. The security comes from human capability rather than from human effort.

The most recent work, XDubber [C.7], applies the same principle to language. Creators who work in a second language lose their audience at the border of it, and automatic dubbing normally hands back a result with no way to see or correct how it was made. XDubber is built around explaining its own decisions to the creator, so a non-native speaker keeps authorship of the result instead of accepting whatever the model produced.

The future this builds

There is a straight line through all of it. Choose the representation that survives contact with cheap hardware. Spend the sensor budget where it buys the most. Build the interaction out of something the person is already doing.

Follow that line and capabilities stop being features of expensive devices and become properties of ordinary ones: affect sensing that runs on a school laptop, AR that lasts a whole afternoon, security that asks for a glance instead of a puzzle, and creative tools that do not require you to work in English to be understood.

That last one is what I am building now. XDubber is the research; AstraClips is the attempt to put it in front of the people it was written for.

The papers

  1. [C.7] XDubber: Supporting Non-Native Creators with Explainable Cross-Lingual Video DubbingTo appear

    Naoto Nishida, Rahul Islam, Karan Ahuja

    ACM UIST 2026 (Under Review)

  2. [C.4] MotionTrace: IMU-based Trajectory Prediction for Smartphone AR Interactions

    Rahul Islam, Vasco Miguel Liang Xu, Karan Ahuja

    IEEE-EMBS International Conference on Wearable and Implantable Body Sensor Networks (BSN 2024)

  3. [Po.2] MicroFlow: Advancing Affective States Detection in Learning Through Micro-Expressions

    Rahul Islam, Sang Won Bae

    CESUN 2023

  4. [Po.1] EyamKayo: Interactive Gaze and Facial Expression Captcha

    Utkarsh Dwivedi, Karan Ahuja, Rahul Islam, Ferdous Barbhuiya, Seema Nagar, Kuntal Dey

    ACM IUI 2017

  5. [A.1] SenTion: A framework for Sensing Facial Expressions

    Rahul Islam, Karan Ahuja, Sandip Karmakar, Ferdous Barbhuiya

    arXiv preprint arXiv:1608.04489