Research
Machine learning on human signals — face, eye, gait, voice.
Mental health sensing
Care that notices before you have to ask for it.
Mental illness is a leading cause of disability worldwide, affecting an estimated 450 million people. Up to one in five of us will live through major depressive disorder or bipolar disorder. It usually arrives in our early twenties, and it is a meaningful predictor of suicide.
Against that, the instruments are startlingly blunt. A clinician asks how the last two weeks felt. The person answers from memory, in a room they had to travel to. Whatever happened on a particular Tuesday, the day it got bad or the hour it lifted, is gone by the time anyone asks. We are trying to catch something that moves hour to hour with a question that averages a fortnight.
Closing that gap is what this line of work is for. It runs in three directions.
Making the signal collectable at all

Depression leaves marks nobody chooses to make: in facial muscle activity, in how the head moves, in how often the eyes are open, in the shape of a smile, in how the pupil answers the prospect of something good. These facial behavior primitives surface whether or not a person is ready to talk, which is precisely why they matter, because those least able to ask for help are so often the ones who need it most.
None of this is news in a laboratory. What has kept it there is not the science but the bargain it demands. The obvious way to read a face all day is to stream video off someone's phone, and that asks a person to surrender the most intimate data they have, continuously, in exchange for a maybe.
So the first direction is a refusal of that bargain. FacePsy [J.5] and PupilSense [C.2] wake on things a person is already doing, unlocking the phone or opening an app, then read the face or the eye for a moment, extract only geometry and gaze, and destroy the image on the device within seconds. Nothing is stored. Nothing is transmitted. The work is open source, so the claim is inspectable rather than asserted.

The door this opens is consent. Sensing that a person would willingly live with is worth more than sensing that is marginally more accurate and switched off in a week.
Reading the whole person, and reading forward

The second direction is about what the signal is for.
Face and eye turn out to be complementary rather than redundant. Expression carries one part of affect and the pupil carries another, and MoodPupilar [C.3] showed the eye alone speaks to ordinary daily mood, not only clinical episodes but the everyday rise and fall that episodes emerge from. Read together [J.5] [C.2] [C.3], they describe a person more completely than any one channel does alone.
They also revealed something that constrains every system built after them: depression does not wear a single face. The signal differs between people, and a model that assumes otherwise fails hardest on whoever it resembles least. Personalisation is not a refinement here; it is a requirement.
Then MoodCam [C.5] moved the question from the present to the future: not how someone feels now, but how they will feel tomorrow. That shift is the difference between a diary and an intervention. A system that sees a difficult day coming can put something useful in front of a person before it arrives: a check-in, a reminder of what helped last time, a nudge toward someone who cares.
Building something the field can stand on

The third direction is refusing to keep any of it.
A single lab studying a small group is a beginning, never a conclusion. So the data behind FacePsy became the BeyondSmile Challenge [J.6], and teams worldwide were invited to beat it. They did. The winning entry surpassed the original work outright. That is the outcome worth wanting: the method stops belonging to one lab, the bar ends up higher than any one group would have set it, and the next person starts from the finish line rather than the start.
The same instinct runs into what these systems should do. A phone that recognises distress and stays silent is a surveillance device; the value only appears when sensing becomes care. One strand of this work [W.1] frames conversational agents that draw on affect, physiology and context alongside established therapeutic technique, so that support is offered rather than waited for. The other [Po.3] asked whether it changes anything, putting a chatbot that only reads words against one that also reads the face. Being seen rather than only read measurably changed how much people trusted it and how understood they felt.
That is a result about people rather than models. Empathy that responds to the whole person lands differently, even coming from software.
The future this builds
Every instrument needed to notice a mental health crisis is already in nearly every pocket on earth. Nothing has to be manufactured, purchased or worn. What has been missing is systems designed for a life rather than for a study: private by construction, light enough to run all day, personal enough to fit one person rather than an average, and honest enough that someone would choose to keep them switched on.
Built that way, the shape of care changes. Assessment stops being an appointment and becomes something continuous and quiet. Support arrives on the bad day rather than three weeks after it. Clinicians see the months between visits instead of guessing at them. Research moves at the pace of a shared dataset rather than a single lab. And the person carrying all of it never has to hand over their face to receive it.
That is the work: not better classifiers, but a version of mental health care that reaches people where they already are.
The papers
Substance use sensing
Knowing someone is impaired, in the moment it matters, without asking them.
There is no breathalyser for cannabis. Blood, urine and saliva can show THC for days after use, long past any impairment, so a positive test says almost nothing about whether a person is affected right now. Alcohol has a reliable chemical test, but by the time it reads positive the drinking has already happened.
Both failures are the same failure. Our tests measure the presence of a substance, and what actually matters to a clinician, to a driver, or to someone trying to cut down is a state: is this person impaired, and is this the moment to say something. Binge drinking alone, defined as four or more drinks for women or five for men on one occasion, carries consequences that arrive within hours: traffic fatalities, injuries, violence.
Closing the gap between substance and state is what this work is about. It runs in two directions.
Reading the state instead of the substance
The first direction abandons chemistry and reads behaviour and physiology instead.
Cannabis intoxication is visible in how a person moves and travels, and in how their body responds. Heart rate rises sharply within minutes of smoking and falls again, a physiological trace of the episode itself. Mobile phone sensor-based detection of subjective cannabis intoxication [J.3] established that everyday phone signals, movement from the accelerometer and travel from GPS, track the subjective experience of being high in daily life, not in a laboratory. The follow-on work [J.7] added a consumer wearable to the phone, pairing behaviour with physiology, and found that the two together describe the state far better than either alone.
The word that matters in both is subjective. These systems are not looking for a molecule. They are looking for the state a person would recognise in themselves, which is the only thing an intervention can usefully respond to.
The door this opens is a measure that works in real time, requires no lab, no blood draw and no roadside device, and runs on hardware people already own.
Getting ahead of the event
The second direction is a shift in tense.
Detection tells you drinking has started. That is useful for a diary and nearly useless for prevention: the decision it might have influenced is already several drinks behind you. So the question became whether the hours before an event look different from ordinary ones.
They do. Predicting imminent same-day binge-drinking events [J.4] showed that a heavy drinking episode is visible in phone data hours before the first drink, from where a person has been and when: the rhythm of a day bending toward a particular evening. Related work on alcohol interventions [Po.4] takes the same location signal and asks how it should shape the support that gets delivered.
That earlier work on cannabis [J.3] found the same thing from another angle: simple routine, day of the week and time of day, already predicts use, because substance use lives inside habit rather than outside it.
Prediction is what makes just-in-time support possible at all. A message that lands during a difficult hour, before a decision rather than after it, is an entirely different intervention from a summary read a week later.
The future this builds
Substance use research has been shaped by what could be measured: a substance in a sample, at a clinic, after the fact. Everything downstream inherited that shape: treatment as episodic, relapse understood in retrospect, help arriving after the harm.
Sensing that reads state instead of substance, and anticipates rather than records, changes what care is able to do. Support can arrive in the hour it matters instead of at the next appointment. Impairment can be assessed where it actually matters, before someone drives, without a test that cannot tell last Tuesday from right now.
That is the future worth building: not better surveillance of people who use substances, but timely support that reaches them at the only moment it can still change the outcome.
The papers

[Po.4] Moving Toward Personalized Behavioral Medicine: Integrating Smartphone-based GPS Data into a Digital Alcohol Intervention
Society of Behavioral Medicine/Annals of Behavioral Medicine 2024
.avif)
Clinical movement
Gait and pose analysis for cerebral palsy and cognitive decline, run on a phone.
[J.9] AIGaitor: Privacy-preserving and cloud-free motion analysis for everyone, using edge computingTo appear
PLOS ONE 2026 (Under Review)
[J.8] Wearable Sensing for Quantifying Cognitive and Balance Functions in Naturalistic Movements of Older Adults with Mild Cognitive Impairment in Therapeutic EnvironmentsTo appear
PLOS Digital Health 2026 (* = Equal Contribution, Under Review)
[C.6] Clinically Accessible 2D Video Analysis Accurately Captures 3D Knee Gait KinematicsTo appear
Annual Meeting of the American Academy for Cerebral Palsy and Developmental Medicine (AACPDM 2026, Under Review)
[Po.5] Evaluating the Feasibility of Deploying Quantized Human Pose Estimation on Smartphone for Gait Analysis for Children with Cerebral Palsy
Annual Meeting of the Gait & Clinical Movement Analysis Society (GCMAS 2026)
Eye & biometrics
The most revealing surface on the body, read by a camera that is already pointed at it.
The eye gives away more than any other part of us that can be seen without touching. The iris carries a pattern unique to a person and stable for life. The sclera and the skin around it, the periocular region, carry more. Pupils widen and narrow to light, to effort, to feeling, without anyone deciding they should. It is the densest signal on the human body, and for most of the day it is pointed directly at a camera.
For decades that fact was largely unusable. Iris recognition meant dedicated infrared hardware, controlled illumination and a person deliberately presenting themselves to a scanner. Eye tracking meant a specialised rig costing hundreds or thousands. The signal was rich; the equipment kept it in the laboratory and at the border.
This work is about what becomes possible once the equipment is just a phone.
Identity in ordinary light
The first direction moved ocular biometrics out of the infrared and into the visible spectrum: the light a normal camera already sees.
That sounds like a small change and is not. Infrared gives a clean, high-contrast iris under controlled conditions. Ordinary light gives you reflections, shadow, whatever colour the room happens to be, and an iris partly hidden by lashes and lids. A preliminary study of CNNs for iris and periocular verification [C.1] and the journal work that grew from it [J.1] asked whether learned representations could hold up under those conditions, using images captured on the phones people actually carry.
They could. And the more interesting finding was about the hardware itself: verification behaves differently when two images come from the same model of phone than when they come from different ones. Every camera leaves its own signature, and a biometric system that ignores that is quietly relying on it. Naming that constraint matters more for real deployment than any single result, because in the real world nobody guarantees both photographs came from the same device.
The door this opens is authentication that needs no scanner, no infrared, no dedicated sensor, only the camera in the phone already in someone's hand.
The eye as an input, not just an identifier

The second direction turned the same camera outward as an interface.
Cheap smartphone VR headsets put a phone a few centimetres from someone's face, with the front camera aimed squarely at one eye, a sensing opportunity that existed in millions of devices and was going entirely unused. EyeSpyVR [J.2] took it, using the phone's camera as the sensor and its own screen as the illuminator, with no added hardware at all.
From that single unglamorous view, through a plastic lens and under shifting screen light, came four distinct capabilities: whether the headset is being worn, when the wearer blinks, who the wearer is, and roughly where they are looking. Capabilities that otherwise required a purpose-built eye tracker arrived instead as a software update.
That is the pattern worth noticing across both directions: the sensor was already there. What was missing was the willingness to work with a difficult signal instead of demanding a clean one.
Where the eye leads

There is a thread running from this work into everything after it. The same camera that can verify who someone is can also read how they are: pupils responding to emotional weight, to reward, to fatigue. That is exactly what the later mental health work does with the eye [C.2] [C.3].
Which makes this the place to be honest about the double edge. A phone that can identify a person from their eye in ordinary light is also a phone that can do it without their knowledge. The same properties that make ocular sensing accessible (no special hardware, works at a distance, needs no cooperation) are what make it worth being careful with. That is not an argument against building it. It is an argument for building it deliberately, and for the people doing the building to be the ones thinking hardest about consent.
The future this points to is quiet: authentication that stops being a ritual, interfaces that know where you are looking without a rig, and health signals read from a glance, all from hardware people already own, and all designed so the person looking into the camera is the one it answers to.
The papers
HCI & systems
Making demanding sensing cheap enough to leave running.
A capability that only works on a workstation is a demonstration. The same capability running on a mid-range phone, all day, without draining the battery is a product, and the distance between those two things is where most sensing research quietly dies.
That gap is what this strand of work is about. Not new signals, but making known signals affordable enough that an ordinary device can carry them, and shaping them into interactions people never have to be taught.
Finding the cheapest honest representation

The instinct in machine perception is to reach for more: more resolution, more parameters, more compute. Often the better move is a representation that throws away everything except what carries the signal.
SenTion [A.1] took that route with facial expression. Rather than treat the face as pixels to be learned, it described it as geometry: the angles between facial landmarks, which it called Inter Vector Angles. Angles are indifferent to how large a face appears and to whose face it is, so the representation transfers across people and distances without retraining. Combined with appearance features, it ran in real time on a laptop CPU with no GPU at all, at low resolution.
That was 2016, and the choice kept paying. Seven years later MicroFlow [Po.2] reached for the same features to do something quite different: reading micro-expressions, the involuntary flickers lasting under half a second that are far harder to suppress than an ordinary expression, to detect whether a student was in flow, bored or anxious while working through a programming assignment. The goal was educational rather than clinical: give an instructor a way to see engagement, and pitch difficulty to the person in front of them.
A representation good enough to survive that jump, from expression recognition to affect in a classroom, was a better investment than any amount of extra compute would have been.
Spending the right resource

The second direction is systems thinking about what a capability costs to run.
Smartphone augmented reality needs to know precisely where the phone is in space, and the obvious way to get that is the camera. But the camera is the most expensive sensor on the device: it drains the battery, heats the phone until performance throttles, and fails exactly when the light is poor or the view is blocked.
MotionTrace [C.4] asked whether the inertial sensors could carry that load instead. The IMU costs almost nothing to run, works in the dark, and does not care what is in front of it. Predicting where the phone is heading, rather than continuously measuring where it is, lets the system prepare what the wearer is about to look at, so content is ready before the view arrives instead of loading after it.
The gain here is not accuracy, it is endurance. A tracking method the device can afford to keep running changes what an AR application can be, in a way that a more accurate method it must ration does not.
Interfaces built on what people already do

The third direction is about the interaction itself: the best input is something a person does naturally and a machine finds hard to fake.
EyamKayo [Po.1] applied that to a familiar irritation. CAPTCHAs ask humans to prove humanity by doing something machines are good at: reading distorted text, spotting objects in photographs. That is why they keep falling. So instead it asks a person to look where they are told and to produce a sequence of expressions. Easy, wordless, and nothing to decipher; for a bot, it means convincingly presenting a human face in three dimensions and animating it on demand. The security comes from human capability rather than from human effort.
The most recent work, XDubber [C.7], applies the same principle to language. Creators who work in a second language lose their audience at the border of it, and automatic dubbing normally hands back a result with no way to see or correct how it was made. XDubber is built around explaining its own decisions to the creator, so a non-native speaker keeps authorship of the result instead of accepting whatever the model produced.
The future this builds
There is a straight line through all of it. Choose the representation that survives contact with cheap hardware. Spend the sensor budget where it buys the most. Build the interaction out of something the person is already doing.
Follow that line and capabilities stop being features of expensive devices and become properties of ordinary ones: affect sensing that runs on a school laptop, AR that lasts a whole afternoon, security that asks for a glance instead of a puzzle, and creative tools that do not require you to work in English to be understood.
That last one is what I am building now. XDubber is the research; AstraClips is the attempt to put it in front of the people it was written for.
The papers
[C.7] XDubber: Supporting Non-Native Creators with Explainable Cross-Lingual Video DubbingTo appear
ACM UIST 2026 (Under Review)










