TY - GEN
T1 - A comparison of AI models for viewport prediction in immersive 360° video streaming
AU - Farooq, Muhammad
AU - Manfredi, Gioacchino
AU - Martini, Maria
AU - Mascolo, Saverio
AU - Cicco, Luca De
PY - 2026/8/7
Y1 - 2026/8/7
N2 - Streaming of 360° videos has gained popularity in recent years due to its immersive viewing experience and the diffusion of inexpensive Head Mounted Displays (HMDs). Such videos allow users consuming the content through an HMD to freely change the point of view of a scene by simply turning their head. Streaming these videos over the Internet is challenging due to bandwidth constraints and latency requirements. Only roughly 1/6th of the whole omnidirectional scene falls in the field of view – or viewport – of the user. Therefore, efficient viewport prediction plays a key role as it helps identifying user’s regions of interest for both short-term and long-term periods. However, as the prediction horizon increases, the accuracy of viewport prediction tends to decrease. This paper presents a comparative analysis of Machine Learning (ML) models for viewport prediction, using a public dataset. Results show that with the selected configurations the Convolutional Neural Network (CNN) model performs better than all other models and achieves a viewport prediction accuracy of 93.93% for the next frame, while the Long Short-Term Memory (LSTM) model performs better when the viewport prediction horizon is increased up to 5 seconds.
AB - Streaming of 360° videos has gained popularity in recent years due to its immersive viewing experience and the diffusion of inexpensive Head Mounted Displays (HMDs). Such videos allow users consuming the content through an HMD to freely change the point of view of a scene by simply turning their head. Streaming these videos over the Internet is challenging due to bandwidth constraints and latency requirements. Only roughly 1/6th of the whole omnidirectional scene falls in the field of view – or viewport – of the user. Therefore, efficient viewport prediction plays a key role as it helps identifying user’s regions of interest for both short-term and long-term periods. However, as the prediction horizon increases, the accuracy of viewport prediction tends to decrease. This paper presents a comparative analysis of Machine Learning (ML) models for viewport prediction, using a public dataset. Results show that with the selected configurations the Convolutional Neural Network (CNN) model performs better than all other models and achieves a viewport prediction accuracy of 93.93% for the next frame, while the Long Short-Term Memory (LSTM) model performs better when the viewport prediction horizon is increased up to 5 seconds.
U2 - 10.1109/CoDIT70676.2026.11630935
DO - 10.1109/CoDIT70676.2026.11630935
M3 - Conference contribution
T3 - International Conference on Control, Decision and Information Technologies (CoDIT)
SP - 1349
EP - 1354
BT - 2026 12th International Conference on Control, Decision and Information Technologies (CoDIT)
PB - Institute of Electrical and Electronics Engineers
CY - Piscataway, U.S.
T2 - 2026 12th International Conference on Control, Decision and Information Technologies (CoDIT)
Y2 - 13 July 2026 through 16 July 2026
ER -