Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference. Zhu, Y., Liang, Z., Wu, Y., Li, T., & Wang, Y. Volume 1 , Association for Computing Machinery, 2025.
Paper doi abstract bibtex Cybersickness remains a major obstacle to the widespread adoption of immersive virtual reality (VR), particularly in consumer-grade environments. While prior methods rely on invasive signals such as electroencephalography (EEG) for high predictive accuracy, these approaches require specialized hardware and are impractical for real-world applications. In this work, we propose a scalable, deployable framework for personalized cybersickness prediction leveraging only non-invasive signals readily available from commercial VR headsets, including head motion, eye tracking, and physiological responses. Our model employs a modality-specific graph neural network enhanced with a Difference Attention Module to extract temporal-spatial embeddings capturing dynamic changes across modalities. A cross-modal alignment module jointly trains the video encoder to learn personalized traits by aligning video features with sensor-derived representations. Consequently, the model accurately predicts individual cybersickness using only video input during inference. Experimental results show our model achieves 88.4% accuracy, closely matching EEG-based approaches (89.16%), while reducing deployment complexity. With an average inference latency of 90ms, our framework supports real-time applications, ideal for integration into consumer-grade VR platforms without compromising personalization or performance. The code will be relesed at https://github.com/U235-Aurora/PTGNN.
@book{
title = {Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference},
type = {book},
year = {2025},
source = {MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025},
keywords = {consumer-grade deployment,cross-modal alignment,cybersickness prediction,difference attention},
pages = {6859-6867},
volume = {1},
issue = {1},
publisher = {Association for Computing Machinery},
id = {7fdbe78d-5bda-3836-8f74-28b58490b47f},
created = {2026-06-09T08:12:19.137Z},
file_attached = {true},
profile_id = {4b66b327-35ad-3956-a9a2-307331dd9988},
last_modified = {2026-06-12T06:38:55.158Z},
read = {false},
starred = {false},
authored = {true},
confirmed = {true},
hidden = {false},
private_publication = {false},
abstract = {Cybersickness remains a major obstacle to the widespread adoption of immersive virtual reality (VR), particularly in consumer-grade environments. While prior methods rely on invasive signals such as electroencephalography (EEG) for high predictive accuracy, these approaches require specialized hardware and are impractical for real-world applications. In this work, we propose a scalable, deployable framework for personalized cybersickness prediction leveraging only non-invasive signals readily available from commercial VR headsets, including head motion, eye tracking, and physiological responses. Our model employs a modality-specific graph neural network enhanced with a Difference Attention Module to extract temporal-spatial embeddings capturing dynamic changes across modalities. A cross-modal alignment module jointly trains the video encoder to learn personalized traits by aligning video features with sensor-derived representations. Consequently, the model accurately predicts individual cybersickness using only video input during inference. Experimental results show our model achieves 88.4% accuracy, closely matching EEG-based approaches (89.16%), while reducing deployment complexity. With an average inference latency of 90ms, our framework supports real-time applications, ideal for integration into consumer-grade VR platforms without compromising personalization or performance. The code will be relesed at https://github.com/U235-Aurora/PTGNN.},
bibtype = {book},
author = {Zhu, Yitong and Liang, Zhuowen and Wu, Yiming and Li, Tangyao and Wang, Yuyang},
doi = {10.1145/3746027.3755115}
}
Downloads: 0
{"_id":"DpNDzbYHm3W34EK84","bibbaseid":"zhu-liang-wu-li-wang-towardsconsumergradecybersicknesspredictionmultimodelalignmentforrealtimevisiononlyinference-2025","author_short":["Zhu, Y.","Liang, Z.","Wu, Y.","Li, T.","Wang, Y."],"bibdata":{"title":"Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference","type":"book","year":"2025","source":"MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025","keywords":"consumer-grade deployment,cross-modal alignment,cybersickness prediction,difference attention","pages":"6859-6867","volume":"1","issue":"1","publisher":"Association for Computing Machinery","id":"7fdbe78d-5bda-3836-8f74-28b58490b47f","created":"2026-06-09T08:12:19.137Z","file_attached":"true","profile_id":"4b66b327-35ad-3956-a9a2-307331dd9988","last_modified":"2026-06-12T06:38:55.158Z","read":false,"starred":false,"authored":"true","confirmed":"true","hidden":false,"private_publication":false,"abstract":"Cybersickness remains a major obstacle to the widespread adoption of immersive virtual reality (VR), particularly in consumer-grade environments. While prior methods rely on invasive signals such as electroencephalography (EEG) for high predictive accuracy, these approaches require specialized hardware and are impractical for real-world applications. In this work, we propose a scalable, deployable framework for personalized cybersickness prediction leveraging only non-invasive signals readily available from commercial VR headsets, including head motion, eye tracking, and physiological responses. Our model employs a modality-specific graph neural network enhanced with a Difference Attention Module to extract temporal-spatial embeddings capturing dynamic changes across modalities. A cross-modal alignment module jointly trains the video encoder to learn personalized traits by aligning video features with sensor-derived representations. Consequently, the model accurately predicts individual cybersickness using only video input during inference. Experimental results show our model achieves 88.4% accuracy, closely matching EEG-based approaches (89.16%), while reducing deployment complexity. With an average inference latency of 90ms, our framework supports real-time applications, ideal for integration into consumer-grade VR platforms without compromising personalization or performance. The code will be relesed at https://github.com/U235-Aurora/PTGNN.","bibtype":"book","author":"Zhu, Yitong and Liang, Zhuowen and Wu, Yiming and Li, Tangyao and Wang, Yuyang","doi":"10.1145/3746027.3755115","bibtex":"@book{\n title = {Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference},\n type = {book},\n year = {2025},\n source = {MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025},\n keywords = {consumer-grade deployment,cross-modal alignment,cybersickness prediction,difference attention},\n pages = {6859-6867},\n volume = {1},\n issue = {1},\n publisher = {Association for Computing Machinery},\n id = {7fdbe78d-5bda-3836-8f74-28b58490b47f},\n created = {2026-06-09T08:12:19.137Z},\n file_attached = {true},\n profile_id = {4b66b327-35ad-3956-a9a2-307331dd9988},\n last_modified = {2026-06-12T06:38:55.158Z},\n read = {false},\n starred = {false},\n authored = {true},\n confirmed = {true},\n hidden = {false},\n private_publication = {false},\n abstract = {Cybersickness remains a major obstacle to the widespread adoption of immersive virtual reality (VR), particularly in consumer-grade environments. While prior methods rely on invasive signals such as electroencephalography (EEG) for high predictive accuracy, these approaches require specialized hardware and are impractical for real-world applications. In this work, we propose a scalable, deployable framework for personalized cybersickness prediction leveraging only non-invasive signals readily available from commercial VR headsets, including head motion, eye tracking, and physiological responses. Our model employs a modality-specific graph neural network enhanced with a Difference Attention Module to extract temporal-spatial embeddings capturing dynamic changes across modalities. A cross-modal alignment module jointly trains the video encoder to learn personalized traits by aligning video features with sensor-derived representations. Consequently, the model accurately predicts individual cybersickness using only video input during inference. Experimental results show our model achieves 88.4% accuracy, closely matching EEG-based approaches (89.16%), while reducing deployment complexity. With an average inference latency of 90ms, our framework supports real-time applications, ideal for integration into consumer-grade VR platforms without compromising personalization or performance. The code will be relesed at https://github.com/U235-Aurora/PTGNN.},\n bibtype = {book},\n author = {Zhu, Yitong and Liang, Zhuowen and Wu, Yiming and Li, Tangyao and Wang, Yuyang},\n doi = {10.1145/3746027.3755115}\n}","author_short":["Zhu, Y.","Liang, Z.","Wu, Y.","Li, T.","Wang, Y."],"urls":{"Paper":"https://bibbase.org/service/mendeley/4b66b327-35ad-3956-a9a2-307331dd9988/file/2e2a4050-2758-1cd7-124b-f0871088490f/37460273755115.pdf.pdf"},"biburl":"https://bibbase.org/service/mendeley/4b66b327-35ad-3956-a9a2-307331dd9988","bibbaseid":"zhu-liang-wu-li-wang-towardsconsumergradecybersicknesspredictionmultimodelalignmentforrealtimevisiononlyinference-2025","role":"author","keyword":["consumer-grade deployment","cross-modal alignment","cybersickness prediction","difference attention"],"metadata":{"authorlinks":{}}},"bibtype":"book","biburl":"https://bibbase.org/service/mendeley/4b66b327-35ad-3956-a9a2-307331dd9988","dataSources":["PW3eQRZmFcK6vLuar","2252seNhipfTmjEBQ","jGKtB5DgeBbTCLaTG"],"keywords":["consumer-grade deployment","cross-modal alignment","cybersickness prediction","difference attention"],"search_terms":["towards","consumer","grade","cybersickness","prediction","multi","model","alignment","real","time","vision","inference","zhu","liang","wu","li","wang"],"title":"Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference","year":2025}