Describe an activity, generate wearable sensor data
Labelled wearable data is scarce. IMUGPT generates it from language: an LLM writes many different descriptions of an activity, a text-to-motion model turns each one into 3D human motion, and the motion becomes virtual IMU data for training activity recognition models. No cameras, no recording sessions.
Pick an activity and one of the LLM-written descriptions. The generated motion plays in 3D with virtual sensors at the dataset's body locations; the panel shows what each sensor records. Drag to orbit.
The LLM varies who is moving, where and how. Each description becomes a different motion, and that variety is what makes the virtual data useful.
Text-to-motion models sometimes get it wrong: a "lying down" prompt becomes sitting. IMUGPT 2.0 captions each motion with MotionGPT and asks an LLM whether the caption matches the description. Below are real generated clips with GPT-4's verdict and the human label.
Incorrect motions in the generated set: 35.5% before filtering.
Verdicts are GPT-4's majority vote over five runs; "correct" is a human's visual judgement of the motion.
Text diversity and motion diversity move together (Pearson r = 0.87β0.92 across datasets). Once new descriptions stop adding diversity, generation can stop: the saturation point.
IMUGPT 1.0: with only 50 descriptions per activity, adding virtual data improves a Random Forest on all three benchmarks. On RealWorld the virtual-only model beats the real-only one.
Macro F1 as the amount of real data grows, with the virtual data held fixed (IMUGPT 1.0). Hover for values.
Real + virtual data with descriptions from each LLM (T2M-GPT motion), against real data only. Macro F1, mean over runs.
Best value per dataset in bold. Generating with a saturation point uses at least 50% fewer descriptions with equal or better accuracy (e.g. Random Forest RealWorld ).
@article{leng2024imugpt,
title = {IMUGPT 2.0: Language-Based Cross Modality Transfer for
Sensor-Based Human Activity Recognition},
author = {Leng, Zikang and Bhattacharjee, Amitrajit and Rajasekhar, Hrudhai
and Zhang, Lizhe and Bruda, Elizabeth and Kwon, Hyeokhyen
and Pl{\"o}tz, Thomas},
journal = {Proceedings of the ACM on Interactive, Mobile, Wearable and
Ubiquitous Technologies},
volume = {8}, number = {3}, articleno = {112}, year = {2024},
doi = {10.1145/3678545}
}
@inproceedings{leng2023generating,
title = {Generating Virtual On-Body Accelerometer Data from Virtual
Textual Descriptions for Human Activity Recognition},
author = {Leng, Zikang and Kwon, Hyeokhyen and Pl{\"o}tz, Thomas},
booktitle = {Proceedings of the 2023 ACM International Symposium on
Wearable Computers (ISWC '23)},
pages = {39--43}, year = {2023},
doi = {10.1145/3594738.3611361}
}