Describe an activity, generate wearable sensor data

Zikang LengAmitrajit BhattacharjeeHrudhai RajasekharLizhe Zhang Elizabeth BrudaHyeokhyen KwonThomas PlΓΆtz
Georgia Institute of Technology
IMWUT 2024IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based HAR ACMarXiv
ISWC 2023Generating Virtual On-body Accelerometer Data from Virtual Textual Descriptions for HAR πŸ† Best Paper Honorable MentionACMarXiv
5HAR datasets
5LLMs compared
4motion models
β‰₯50%less generation with diversity-based stopping

Labelled wearable data is scarce. IMUGPT generates it from language: an LLM writes many different descriptions of an activity, a text-to-motion model turns each one into 3D human motion, and the motion becomes virtual IMU data for training activity recognition models. No cameras, no recording sessions.

Interactive

From a sentence to a sensor signal

Pick an activity and one of the LLM-written descriptions. The generated motion plays in 3D with virtual sensors at the dataset's body locations; the panel shows what each sensor records. Drag to orbit.

β€”
Dataset
Activity
LLM-written descriptions
Sensor location
Accelerometer (m/sΒ²)
Gyroscope (rad/s)
xyzsensor frame, 20 Hz
Diversity from language

One activity, a thousand descriptions

The LLM varies who is moving, where and how. Each description becomes a different motion, and that variety is what makes the virtual data useful.

IMUGPT 2.0 Β· Motion filter

An LLM checks every generated motion

Text-to-motion models sometimes get it wrong: a "lying down" prompt becomes sitting. IMUGPT 2.0 captions each motion with MotionGPT and asks an LLM whether the caption matches the description. Below are real generated clips with GPT-4's verdict and the human label.

Filter quality against human labels (RealWorld, published)

Incorrect motions in the generated set: 35.5% before filtering.

Show clips

Verdicts are GPT-4's majority vote over five runs; "correct" is a human's visual judgement of the motion.

IMUGPT 2.0 Β· When to stop generating

Diverse text gives diverse motion, until it saturates

Text diversity and motion diversity move together (Pearson r = 0.87–0.92 across datasets). Once new descriptions stop adding diversity, generation can stop: the saturation point.

Generated RealWorld motions, embedded in 2D (hover for the description)
Descriptions kept at the saturation point (RealWorld, of ~1,000 generated)

Results

Virtual data improves recognition

IMUGPT 1.0: with only 50 descriptions per activity, adding virtual data improves a Random Forest on all three benchmarks. On RealWorld the virtual-only model beats the real-only one.

RealVirtual onlyReal + VirtualMacro F1, Random Forest

How much real data does virtual data replace?

Macro F1 as the amount of real data grows, with the virtual data held fixed (IMUGPT 1.0). Hover for values.

Real onlyReal + VirtualVirtual only

IMUGPT 2.0 across 5 LLMs and 5 datasets

Real + virtual data with descriptions from each LLM (T2M-GPT motion), against real data only. Macro F1, mean over runs.

Best value per dataset in bold. Generating with a saturation point uses at least 50% fewer descriptions with equal or better accuracy (e.g. Random Forest RealWorld ).

Renders

Generated motions, rendered

Cite

BibTeX

@article{leng2024imugpt,
  title     = {IMUGPT 2.0: Language-Based Cross Modality Transfer for
               Sensor-Based Human Activity Recognition},
  author    = {Leng, Zikang and Bhattacharjee, Amitrajit and Rajasekhar, Hrudhai
               and Zhang, Lizhe and Bruda, Elizabeth and Kwon, Hyeokhyen
               and Pl{\"o}tz, Thomas},
  journal   = {Proceedings of the ACM on Interactive, Mobile, Wearable and
               Ubiquitous Technologies},
  volume    = {8}, number = {3}, articleno = {112}, year = {2024},
  doi       = {10.1145/3678545}
}

@inproceedings{leng2023generating,
  title     = {Generating Virtual On-Body Accelerometer Data from Virtual
               Textual Descriptions for Human Activity Recognition},
  author    = {Leng, Zikang and Kwon, Hyeokhyen and Pl{\"o}tz, Thomas},
  booktitle = {Proceedings of the 2023 ACM International Symposium on
               Wearable Computers (ISWC '23)},
  pages     = {39--43}, year = {2023},
  doi       = {10.1145/3594738.3611361}
}