Skip to main content

Adaptemy research accepted for presentation at IEEE Telepresence 2026

New collaborative research explores how Vision Transformers could make quality assessment for immersive video significantly more computationally efficient.

We are delighted to share that new research by Adaptemy and its academic partners has been accepted for presentation as a Special Session Short Paper at the 2026 IEEE Conference on Telepresence.

The paper, “On the Use of ViT Embeddings for Blind Omnidirectional Video Quality Assessment Based on Saliency-guided Viewport Extraction,” is authored by Jiantao Wu and Ioana Ghergulescu of Adaptemy, together with Sara Baldoni and Federica Battisti of the University of Padova, and Gabriel-Miro Muntean of Dublin City University.

The challenge: assessing quality without the original video

Omnidirectional—or 360-degree—video is central to many virtual reality and immersive telepresence experiences. However, its extremely high resolution creates substantial bandwidth and computational demands.

Assessing the quality of this video is particularly difficult when the original, undistorted source is unavailable. This is known as no-reference—or “blind”—video quality assessment: estimating what the viewer is experiencing using only the video they actually receive.

There is another complication. A viewer wearing a headset sees only a limited portion of the complete 360-degree environment at any one time. Analysing every part of every frame can therefore consume significant computational resources while giving equal attention to areas the viewer may never see.

Using AI to focus attention

The research explores whether embeddings from pretrained Vision Transformers—or ViTs—can identify the visually significant regions within an immersive scene without requiring labelled training data.

The proposed framework uses these embeddings to create an unsupervised saliency map: an estimate of which areas are most likely to attract the viewer’s attention. It then groups these areas into a smaller number of viewports for quality assessment.

In simple terms, the approach asks:

Can an AI model help a quality-assessment system concentrate its effort on the parts of a 360-degree video that matter most to the viewer?

The study compares embeddings from two pretrained models: CLIP, which has learned relationships between images and language, and an ImageNet-trained vision model.

A promising efficiency result

The preliminary results show that CLIP produced more concentrated saliency regions. This reduced the average number of retained viewports from 7.6 to 3.2—an approximate 60% reduction in the regions requiring evaluation.

That could ultimately help make quality monitoring more computationally efficient in environments such as live immersive streaming and resource-constrained edge devices.

The researchers are careful not to overstate the finding. Correlation with human quality ratings remained low for both models, and CLIP did not improve prediction accuracy over the ImageNet approach. The current framework should therefore be viewed as evidence of potential computational efficiency, rather than a deployment-ready predictor of perceived quality.

Future research will explore stronger no-reference video-quality models, temporal and motion features, and viewport-selection methods that more closely reflect individual viewing behaviour.

Collaborative research through HEAT

This research was conducted in collaboration between Adaptemy, the University of Padova and Dublin City University, and was partially funded through the European Union’s Horizon Europe programme as part of the HEAT project.

Although the study focuses on immersive video, it reflects a wider research principle that is also central to our work at Adaptemy: using AI to identify what matters, reduce unnecessary processing and support more intelligent, responsive digital experiences.

The paper will be presented at the 2026 IEEE Conference on Telepresence, taking place in Bristol from 9–13 November 2026. Following presentation, it is expected to appear in the conference proceedings and be submitted for inclusion in IEEE Xplore.

Congratulations to our own Jiantao Wu and Ioana Ghergulescu alongside Sara Baldoni, Federica Battisti and Gabriel-Miro Muntean!