OA18973A

OA18973A - Quality estimation of adaptive multimedia streaming

Info

Publication number: OA18973A
Application number: OA1201800536
Authority: OA
Inventors: Tomas Lundberg; Junaid SHAlKH; Jing Fu; Gunnar HEIKKILÄ; David Lindegren
Original assignee: Telefonaktiebolaget Lm Ericsson (Publ)
Priority date: 2016-06-29
Filing date: 2017-06-29
Publication date: 2019-10-28

Abstract

There are provided mechanisms for predicting a multimedia session MOS. The multimedia session comprises a video session and an audio session, wherein video quality is represented by a vector of per-time-unit scores of video quality and wherein audio quality is represented by is a vector of pertime-unit scores of audio quality. The multimedia session is represented by a vector of rebuffering start times of each rebuffering event, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration being the time between an initiation of the multimedia session and a start time of the multimedia session. The method comprises generating audiovisual quality features from the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality. The audiovisual quality features comprise: a vector of per-time-unit scores of audiovisual quality, calculated as a polynomial function of the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality; a weighted combination of the per-time-unit scores of audiovisual quality, wherein the weights are exponential functions of a time since the start time of multimedia session and a multimedia session duration; a negative bias representing how a sudden drop in per-timeunit scores of audiovisual quality affects the multimedia session MOS; and a term representing a degradation due to oscillations in the per-timeunit-scores of audiovisual quality. The method comprises generating buffering features from the vector of rebuffering start times of each rebuffering event, calculated from the start time of multimedia session, and the vector of rebuffering durations of each rebuffering event. The method comprises estimating a multimedia session MOS from the generated audiovisual quality features and the generated buffering features.

Description

QUALITY ESTIMATION OF ADAPTIVE MULTIMEDIA

STREAMING

TECHNICAL FIELD

This invention relates to a method, a MOS estimator, a computer program and a computer program product for predicting multimedia session MOS (Mean Opinion Score).

BACKGROUD

Streaming media is more popular than ever, as both consumer and enterprise users increase content consumption. It is used on social media such as YouTube, Twîtter, and Facebook, and of course also by the providers of on-demand video services such as Netflix. According to some reports, Netflix and YouTube together make up half of peak Internet traffic in North America. Moreover, the number of subscription video on demand homes is forecast to reach 306 million across 200 countries by 2020.

When the transmission capacity in a network fluctuâtes, for instance for a wireless connection, the media player can often select to adapt the bitrate, so that the video can still be delivered, albeit with sometimes worse quality (lower bitrate, lower resolution etc.). An example is shown in Figurai A for a 60-second video, where the segment heights represents the bitrate, and each segment is 5 second long. In almost ail cases, the quality will vary in a corresponding way, i.e. higher bitrate will give a higher quality, and lower bitrate will give a lower quality.

It is therefore of vital importance for providers to estimate the users’ Quality of Expérience (QoE), which is fundamentally the subjective opinion of the quality of a service. For this purpose, subjective test may be used, where a panel of viewers are asked to evaluate the perceived quality of streaming media, Typically, the quality is given on a scale from 1 (“bad”) to 5 (“excellent”), and is then averaged over ail viewers, forming a Mean Opinion Score (MOS). However, these subjective tests ara costly, both in time and money, and, to circumvent this, objective QoE estimation methods (objective quality models) hâve been developed.

Mean Opinion Score (MOS) is a measure of subjective opinion of users about a service or application performance. It has been widely used to evaluate the quality of multimedia applications. The ITU-T Recommendation P. 800 has standardized the use of MOS on a 5point Absolute Category Rating (ACR) scale for évaluation of the audio-visual test sequences. The ACR scale ranges from 5 (Excellent) to 1 (Bad). This method is particularly relevant in scénarios where a user is presented with one test sequence at a time and then asked to rate it.

Different objective quality models are normally used for audio and video. The models estimate the quality dégradation due to the coding itself, taking into account parameters such as bitrate (audio and video), sampling rate (audio), number of channels (audio), resolution (video), frame rate (video), GOP size (video, a parameter related to video coding), etc. The output from the audio or video quality model for a complété session (as in the picture above) is typically a list of objective MOS scores, where each score represents the quality for an individual media segment (i.e. each score represents the quality during 5 seconds in the figure above). Examples of the audio and video coding quality models can be found in the ITU-T P.1201 recommendation.

When created, the audio and video quality models are trained on a set of subjective tests. This is accomplished in the following manner: a spécifie number of parameters are varied and multimedia clips are produced using these parameters. These clips are then graded by viewers during a subjective test, and the quality models are then made to as closely as possible (in some sense) match the results from the subjective tests.

Typically, the models are trained on shorter signal segments, typically around 5 to 10 seconds, where the media quality is more or less constant during the clip. This means that the models in principle only give accurate results when presented with segments of corresponding durations, and where no major quality variations are present. To obtain an objective score for a multimedia clip that is much longer than this, an aggregation model is needed. Due to nonlinear human perception processing it is not just possible to e.g. average the individual segment scores.

An aggregation model also combines the audio and video model quality scores into combined media scores, representing the total perception of the media. Another task for the aggregation model is to take into account dégradations due to buffering. Buffering occurs when the transmission speed in the network is not high enough so that more data is consumed in the media player than what is delivered by the network. This will cause “gaps” in the media playout during which the media player fills up its data buffer, as exemplified in Figure 1B. The aggregation model will consequently in the end need to take both these effects into account, both a varying intrinsic audio and video quality, and dégradations due to bufferings, as in the more complex example shown in Figure 1C.

The buffering can be either initial buffering (before any media is presented to the user) or possible rebufferings during play-out.

SUMMARY

Existing buffer aggregation models, e.g. as in ITU-T P.1201, hâve so far been limited to session lengths of up to one minute, which is much too short for a typical video session, e.g. YouTube. With longer sequences, human memory effects also start to be noticeable, meaning that people remember less of a what they saw longer back in time, and thus mostly rate the quality of the video after the last parts. This is not handled in existing models. To accurately mimic the total effect of quality adaptations, different resolutions, buffering and longer session times, a more complex model is needed.

It is an object to improve how Mean Opinion Scores are predicted.

A first aspect of the embodiments defines a method, performed by a Mean Opinion Score, MOS, estimator, for predicting a multimedia session MOS. The multimedia session comprises a video session and an audio session, wherein video quality is represented by a vector of pertime-unit scores of video quality and wherein audio quality is represented by is a vector of pertime-unit scores of audio quality. The multimedia session is represented by a vector of rebuffering start times of each rebuffering event, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration being the time between an initiation of the multimedia session and a start time of the multimedia session. The method comprises generating audiovisual quality features from the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality. The audiovisual quality features comprise: a vector of per-time-unit scores of audiovisual quality, calculated as a polynomial function of the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality; a weighted combination of the per-time-unit scores of audiovisual quality, wherein the weights are exponential functions of a time since the start time of multimedia session and a multimedia session duration; a négative bias representing how a sudden drop in per-time-unit scores of audiovisual quality affects the multimedia session MOS; and a term representing a dégradation due to oscillations in the pertime-unit-scores of audiovisual quality. The method comprises generating buffering features from the vector of rebuffering start fîmes of each rebuffering event, calculated from the start time of multimedia session, and the vector of rebuffering durations of each rebuffering event. The method comprises estimating a multimedia session MOS from the generated audiovisual quality features and the generated buffering features.

A second aspect of the embodiments defines a Mean Opinion Score, MOS, estimator, for predicting a multimedia session MOS. The multimedia session comprises a video session and an audio session, wherein video quality is represented by a vector of per-time-unit scores of video quality and wherein audio quality is represented by is a vector of per-time-unit scores of audio quality. The multimedia session is represented by a vector of rebuffering start times of each rebuffering event, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration being the time between an initiation of the multimedia session and a start time of the multimedia session. The MOS estimator comprises processing means operative to generate audiovisual quality features from the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality. The audiovisual quality features comprise a vector of per-time-unit scores of audiovisual quality, calculated as a polynomial function of the vector of per-time-unit scores of video quality and the vector of pertime-unit scores of audio quality; a weighted combination of the per-time-unit scores of audiovisual quality, wherein the weights are exponential functions of a time since the start time of multimedia session and a multimedia session duration; a négative bias representing how a sudden drop in per-time-unit scores of audiovisual quality affects the multimedia session MOS; and a term representing a dégradation due to oscillations in the per-time-unitscores of audiovisual quality. The MOS estimator comprises processing means operative to generate buffering features from the vector of rebuffering start times of each rebuffering event, calculated from the start time of multimedia session, and the vector of rebuffering durations of each rebuffering event. The MOS estimator comprises processing means operative to estimate a multimedia session MOS from the generated audiovisual quality features and the generated buffering features.

A third aspect of the embodiments defines a computer program for a Mean Opinion Score, MOS, estimator, for predicting a multimedia session MOS. The multimedia session comprises a video session and an audio session, wherein video quality is represented by a vector of pertime-unit scores of video quality and wherein audio quality is represented by is a vector of pertime-unit scores of audio quality. The multimedia session is represented by a vector of rebuffering start times of each rebuffering event, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration being the time between an initiation of the multimedia session and a start time of the multimedia session. The computer program comprises a computer program code which, when executed, causes the computer program to generate audiovisual quality features from the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality. The audiovisual quality features comprise a vector of per-time-unit scores of audiovisual quality, calculated as a polynomial function of the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality; a weighted combination of the per-time-unit scores of audiovisual quality, wherein the weights are exponential functions of a time since the start time of multimedia session and a multimedia session duration; a négative bias representing how a sudden drop in per-time-unit scores of audiovisual quality affects the multimedia session MOS; and a terrn representing a dégradation due to oscillations in the per-time-unit-scores of audiovisual quality. The computer program comprises a computer program code which, when executed, causes the computer program to generate buffering features from the vector of rebuffering start times of each rebuffering event, calculated from the start time of multimedia session, and the vector of rebuffering durations of each rebuffering event. The computer program comprises a computer program code which, when executed, causes the computer program to estimate a multimedia session MOS from the generated audiovisual quality features and the generated buffering features.

A fourth aspect of the embodiments defines a computer program product comprising computer readable means and a computer program according to the third aspect, stored on the computer readable means.

Advantageously, at least some of the embodiments provide a MOS estimator that handles both short and long video sessions, and gives a more accurate MOS score The MOS estimator according to at least some of the embodiments is relatively low-complex in terms of computational power and can easily be implemented in ail environments.

It is to be noted that any feature of the first, second, third and fourth aspects may be applied to any other aspect, whenever appropriate. Likewise, any advantage of the first aspect may equally apply to the second, third and fourth aspect respectively, and vice versa. Other objectives, features and advantages of the enclosed embodiments will be apparent from the following detailed disclosure, from the attached dépendent claims and from the drawings.

Generally, ail terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise herein. Ail référencés to a/an/the element, apparatus, component, means, step, etc. are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not hâve to be performed in the exact order disclosed, unless explicitly stated.

BRIEF DESCRIPTION OF THE DRAWINGS

The invention is now described, by way of example, with reference to the accompanying drawings, in which:

Figures 1A-C are schematic graphs illustrating buffering and bitrate over time.

Figure 2 illustrâtes the steps performed by a MOS estimator according to the embodiments of the present invention.

Figure 3 illustrâtes the weight factor as a function of a sample âge according to the embodiments of the present invention.

Figure 4 shows an initial buffering impact as a function of initial buffering duration according to the embodiments of the present invention.

Figure 5 shows a forgetness factor impact as a function of time since the start time of multimedia session, according to the embodiments of the present invention.

Figure 6 illustrâtes a rebuffering duration impact as a function of rebuffering duration, according to the embodiments of the present invention.

Figure 7 illustrâtes a rebuffering répétition impact as a function of rebuffering répétition number, according to the embodiments of the present invention.

Figure 8 illustrâtes a forgetting factor impact as a function of time since the last rebuffering, according to the embodiments of the présent invention.

Figure 9 is an aggregation module according to the embodiments of the present invention.

Figure 10 depicts a schematic block diagram illustrating functional units of a MOS estimator for predicting a multimedia session MOS according to the embodiments of the present invention.

Figure 11 illustrâtes a schematic block diagram illustrating a computer comprising a computer program product with a computer program for predicting a multimedia session MOS, according to embodiments of the present invention.

DETAILED DESCRIPTION OFTHE PROPOSED SOLUTION

The invention will now be described more fully hereinafter with reference to the accompanying drawings, in which certain embodiments ofthe invention are shown. This invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complété, and will fully convey the scope of the invention to those skilled in the art. Like numbers refer to like éléments throughout the description.

The subjective MOS is how humans rate the quality of a multimedia sequence. Objective MOS estimation is using models to predict/estimate how humans will rate it. in general, parametric based methods are usually used to predict the multimedia MOS. This kind of parametric based methods usually results in quite a large prédiction error.

The basic idea of embodiments presented herein is to predict the multimedia session MOS. The multimedia session comprises a video session and an audio session, wherein video quality is represented by a vector of per-time-unit scores of video quality and wherein audio quality is represented by a vector of per-time-unit scores of audio quality. The multimedia session is further represented by a vector of rebuffering start times of each rebuffering event, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration being the time between an initiation of the multimedia session and a start time of the multimedia session.

A time unit may be a second. Thus, the lists of per time unit scores of the video and audio quality may be obtained per second. For example, a 300 second clip has audio and video vectors with 300 éléments each.

Initial buffering duration may also be expressed in seconds. For example, an 8-second initial buffering (which has a start time at 0 seconds) has a duration of 8 seconds. Rebuffering duration and location may also be expressed in seconds. Start times are in media time, so it doesn't dépend on a duration of any previous buffering.

According to one aspect, a method, performed by a MOS, Mean Opinion Score, estimator, for predicting a multimedia session MOS is provided, as described in Figure 2. The method comprises a step S1 of generating audiovisual quality features from the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality.

The audiovisual quality features comprise a vector of per-time-unit scores of audiovisual quality, calculated as a polynomial function of the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality. That is, video quality and audio quality are “merged” to a measure of a combined quality, mosBoth. This merge is known from ITU-T P. 1201. For example, as given in a source code below, a per-time-unit score of audiovisual quality may be calculated as:

mosBoth[i] = (mosV[i] - 1) + c[17] (mosA[i] - 1) + c[18] · (mosV[i] - 1) ^m°^sA^-----) + c[17] + c[18] wherein mosV and mosA respectively are vectors of per-time-unit scores of video and audio quality, and c[17] and c[18] are audio and video merging weights. For example, c[17] may be set to 0.16233, and c[18] to-0.013804, but the present invention is by no means limited to these spécifie values.

The audiovisual quality features further comprise a weighted combination of the per-time-unit scores of audiovisual quality, wherein the weights are exponential functions of the time since the start time of multimedia session and the multimedia session duration. Namely, due to memory effects, media played longer back in time and thus longer back in memory is slightly forgotten, and is thus weighted down. The weighted combination of the per-time-unit scores of audiovisual quality is referred to as mosBasic. An example of the weights as functions of a différence between the multimedia session duration and the time since the start time (depicted as a sample âge here) of multimedia session is shown in Fig. 3. A source code below demonstrates how mosBasic may be calculated:

for i in range(mosLength):

mosBoth[i] = (1 * (mosV[i] - 1) + c[17] * (mosA[i] - 1) + c[18] * (mosV[i] - 1) * (mosA[i] -1)/4)/(1+ c[17] + c[18]) + 1 mosTime = mosLength - i - 1 mosWeight = exponential([1, c[l], 0, c[2]], mosTime) suml += mosBoth[i] * mosWeight sum2 += mosWeight mosBasic = suml / sum2 wherein mosLength corresponds to the multimedia session duration, mosTime corresponds to the différence between the multimedia session duration and the time since the start time of multimedia session, and c[1] and c[2] are memory adaptation weights. For example, c[1] may be set to 0.2855, and c[2] to 10.256, but the present invention is by no means limited to these spécifie values.

The audiovisual quality features further comprise a négative bias. The négative bias represents how a sudden drop in per-time-unit scores of audiovisual quality affects the multimedia session MOS. When media quality varies, one is more affected by a sudden drop in quality, as compared to a similar sudden improvement. This effect is captured by the négative bias. The négative bias may be modelled by calculating the offsets for each per-time-unit (e.g., onesecond) quality score towards mosBasic. These offsets may also be sealed by the forgettîng factor weight, so that media longer back in memory gets less impact.

From this vector of weighted per-time-unit (i.e., one-second) offsets, a certain percentile can be calculated. For example, it may be an —10th percentile, but it could be a different percentile as well. This is usually a négative number, as the lowest quality scores in the vectors should normally be lower than mosBasic, so the resuit is negated into a positive value, meaning a higher value now indicates a higher impact of the négative bias. This is then sealed linearly to the right range. An example of a source code for calculating the négative bias is as:

mosOffset = list(mosBoth) for i in range(mosLength):

mosTime = mosLength-i-1 mosWeight = exponential([1, c[l], 0, c[2]], mosTime) mosOffset[i] = (mosOffset[i] - mosBasic)*mosWeight mosPerc = np.percentile(mosOffset, c[22], interpolation='1inear') negBias = np.maximum(0, -mosPerc) negBias = negBias*c[23]

Equivalent^, the négative bias is calculated as follows:

negBias = max I 0, -lOt/ι percentile of per — time - unit scores of audiovisual quality [t] / (T-f) log(0.5) c[l] + (1 - c[l]) e =^2] •c[23] wherein t is time since the start time of multimedia session and T is the multimedia session duration. Here c[22] and c[23] represent négative bias coefficients. For example, c[22] may be set to 9.1647, and c[23] to 0.74811, but the present invention is by no means limited to these spécifie values.

The audiovisual quality features comprise a term representing a dégradation due to oscillations in the per-time-unit-scores of audiovisual quality. Nameîy, when media quality fluctuâtes this is annoying, and the effect of quality fluctuation is caught by counting the number of tops and dips where the unweighted one-second media quality scores (mosBoth) goes above or below mosBasic. In other words, the term representing a dégradation due to oscillations in the pertime-unit scores of audiovisual quality may be calculated as the number of occurrences when the absolute différence between the per-time-unit scores of the audiovisual quality and the weighted combination of the per-time-unit scores of audiovisual quality exceeds a given threshold value, divided by the multimedia session duration. The threshold value may be used to disregard small variations that may not be perceivable. An example for the threshold value is 0.1, i.e., a hystérésis of 0.1 is used.

The term representing a dégradation due to oscillations, oscDeg, in the per-time-unit-scores of audiovisual quality may also be truncated so that the maximum value is 0.2 oscillations per second. This may then multiplied by a standard déviation of the per-time unit (i.e., per-second) audiovisual quality values, so that higher level of oscillations gets a higher impact. The following source code illustrâtes how the term representing a dégradation due to oscillations can be calculated: ose = 0 offset = 0.1 State = 0 for i in range(mosLength): if State != 1: if mosBoth[i] > mosBasic + offset: ose += 1

State elif state != -1:

if mosBoth[i] < mosBasic - offset:

ose += 1 state =-l oscRel = ose / mosLength oscRel = np.minimum(oscRel, 0.2) # Limit to one change per 5 sec oscDeg = np.power(oscRel * np.std(mosBoth, ddof=l), c[19]) * c[20]

The resuit may then be scaled non-linearly (approximately squared), and finally linearly scaled to the right range.

The method comprises a step S2 of generating buffering features from the vector of rebuffering start times of each rebuffering event, calculated from the start time of multimedia session, and the vector of rebuffering durations of each rebuffering event.

The generated buffering features may comprise a term representing a dégradation due to initial buffering, initDeg, and a term representing a dégradation due to rebuffering, bufDeg.

The term representing dégradation due to initial buffering may be modeled as a product of a term representing an initial buffering impact and a term representing a forgetness factor impact.

The initial buffering impact may be a sigmoid function ofthe initial buffering duration. For example, the sigmoid function may basically give a zéro impact below 5 seconds and an impact of 4 if the initial buffering duration is longer than that, as shown in Figure 4. The source code for calculating inlDeg may be as follows:

lengthDeg = sigmoid([0, 4, c(10], c[10] + c[llj], buflnit) memoryDeg = exponential([1, c[4], 0, c[5]], mosLength) initDeg = lengthDeg*memoryDeg

Here c[10] and c[11] are constants related to initial buffering and c[4] and c[5] are memory weights related to initial buffering. For example, c[10] = 4.5327, c[11] = 1.0054, c[4] = 0.054304 and c[5] = 10.286, but the present invention is by no means limited to these spécifie values.

However, the impact from initial buffering is only annoying during the initial buffering itself or close after. If the media continues to stream, this problem isforgotten quite soon. Thus, the second modelling is to weight the initial buffering impact with a forgetness factor. The forgetness factor may be an exponential function of the time since the start time of multimedia session, as shown in Fig. 5.

The term representing dégradation due to rebuffering, bufDeg, may be modeled as a sum, over ail rebuffering events, of products of a rebuffering duration impact, a rebuffering répétition impact, and an impact of time since the last rebuffering. For each rebuffering instance, first the impact of the rebuffering is calculated. The rebuffering duration impact may be a sigmoid function of a rebuffering duration, as shown in Figure 6.

However, the rebuffering duration impact only models a single rebuffering, evaluated close to the time when the rebuffering happened. If there are more rebufferings, one gets more annoyed for each additional one. This is modeled by the rebuffering répétition impact. The rebuffering répétition impact may be a sigmoid function of a rebuffering répétition number, as shown in Figure 7. For example, a weight of up to 5 is assigned when the number of rebufferings becomes 4 or more.

Finally, as the time since the last rebuffering passes, one tends to forget about it. The impact of time since the last rebuffering, or a so-called forgetting factor, may be modelled as an exponential function of the time since the last rebuffering, as shown in Figure 8.

To get the final effect of a single rebuffering, the rebuffering duration impact, the rebuffering répétition impact and the impact of time since the last rebuffering are multiplied. This resuit is then added to the total impact resuit for ail rebufferings, as shown in the following source code;

bufDeg = 0;

for j in range [len(bufLength)) : lengthDeg = sigmoid([0, 4, c[12], c[12]+c(13 J] , bufLength[j]) repeatDeg = sigmoid([l, c[14], c[15], c [ 15)+c[16]], j) memoryDeg = exponential([1, c[7], 0, c[8]] , mosLength - bufStart[j]) bufDeg = bufDeg + lengthDeg * repeatDeg * memoryDeg bufDeg = bufDeg/4 * (mosBasic-1)

Here lengthDeg, repeatDeg and memoryDeg dénoté impacts due to rebuffering duration, rebuffering répétition and the impact of time since the last rebuffering respectively, and bufStart[j] dénotés the time since the last rebuffering. In addition, c[12] and c[13] are rebuffering impact constants, c[14]-c[16] are constants related to rebuffering répétition, and c[7] and c[8] are time-since-the-last rebuffering impact (also referred to as rebuffering memory weights). For example, one may set c[12] = -67.632, c[13] = 158.18, c[14] = 4.9894, c[15] = 2.1274, c[16] = 2.0001, c[7] = 0.17267 and c[8] = 10, but the present invention is by no means limited to these spécifie values.

Finally, the resulting term representing dégradation due to rebuffering may be rescaled relative to mosBasic. This may be done since people are more annoyed by a rebuffering if they otherwise hâve good quality, while if the quality is poor, a rebuffering does not dégradé peoples' perception so much.

The method comprises a step S3 of estimating a multimedia session MOS from the generated audiovisual quality features and the generated buffering features, as illustrated in Figure 9. The multimedia session MOS may be estimated as the différence between the weighted combination of the per-time-unît scores of audiovisual quality and the sum of: the négative bias, the term representing dégradation due to oscillations in the per-time-unit-scores of audiovisual quality, the term representing dégradation due to initial buffering, and the term representing dégradation due to rebuffering. The score is also truncated to be between 1 and

5. In other words, the multimedia session MOS may be estimated according to a source code below:

mos = mosBasic - initDeg - bufDeg - oscDeg - negBias if mos < 1: mos = 1 if mos > 5: mos = 5 return (mos)

Figure 10 is a schematic block diagram of a MOS estimator 100, for predicting a multimedia session MOS, wherein the multimedia session comprises a video session and an audio session. The video quality is represented by a vector of per-time-unit scores of video quality and the audio quality is represented by is a vector of per-time-unit scores of audio quality. The multimedia session is represented by a vector of rebuffering start times of each rebuffering event, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration being the time between an initiation of the multimedia session and a start time of the multimedia session.

The MOS estimator 100 comprises, according to this aspect, a generating unit 170, configured to generate audiovisual quality features from the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality. The audiovisual quality features comprise:

- a vector of per-time-unit scores of audiovisual quality, calculated as a polynomial function of the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality;

- a weighted combination of the per-time-unit scores of audiovisual quality, wherein the weights are exponentîal functions of a time since the start time of multimedia session and a multimedia session duration;

- a négative bias representing how a sudden drop in per-time-unit scores of audiovisual quality affects the multimedia session MOS; and

- a term representing a dégradation due to oscillations in the per-time-unit-scores of audiovisual quality.

The generating unit 170 is further configured to generate buffering features from the vector of rebuffering start times of each rebuffering event, calculated from the start time of multimedia session, and the vector of rebuffering durations of each rebuffering event.

The MOS estimator 100 comprises, according to this aspect, an estimating unit 180, configured to estimate a multimedia session MOS from the generated audiovisual quality features and the generated buffering features.

The generating 170 and estimating 180 units may be hardware based, software based (in this case they are called generating and estimating modules respectively) or may be a combination of hardware and software.

The generating unit 170 may calculate the négative bias as:

negBias = max l 0, -lOth percentile of per - time — unit scores of audiovisual quality [t] c[l] + (1 — c[l]) · e (T-Qlog(0.5)· •c[23] wherein t is time since the start time of multimedia session, T is the multimedia session duration and c[1], c[2] and c[23] are constants.

The generating unit 170 may calculate the dégradation due to oscillations in the per time unit scores of audiovisual quality as the number of occurrences when the absolute différence between the per time unit scores of the audiovisual quality and the weighted combination of the per time unit scores of audiovisual quality exceeds a given threshold value, divided by the multimedia session duration. The threshold value may be e.g. 0.1. The dégradation due to oscillations in the per time unit scores of audiovisual quality may also be truncated so that the maximum value is 0.2 oscillations per second.

The generated buffering features comprise a term representing a dégradation due to initial buffering and a term representing a dégradation due to rebuffering. Thus, the generating unit 170 may model the term representing dégradation due to initial buffering as a product of a term representing an initial buffering impact and a term representing a forgetness factor impact. The initial buffering impact may be a sigmoid fonction of the initial buffering duration, and the forgetness factor may be an exponential fonction of the time since the start time of multimedia session.

The generating unit 170 may model the term representing dégradation due to rebuffering as a sum, over ali rebuffering events, of products of a rebuffering duration impact, a rebuffering répétition impact, and an impact of time since the last rebuffering ended. The rebuffering duration impact may be a sigmoid fonction of a rebuffering duration. The rebuffering répétition impact may be a sigmoid fonction of a rebuffering répétition number. The impact of time since the last rebuffering ended may be an exponential fonction of the time since the last rebuffering ended.

The MOS estimator 100 may estimate the multimedia session MOS as the différence between the weighted combination of the per-time-unit scores of audiovisual quality and the sum of the négative bias, the term representing dégradation due to oscillations in the per-time-unit-scores of audiovisual quality, the term representing dégradation due to initial buffering, and the term representing dégradation due to rebuffering.

The MOS estimator 100 can be implemented in hardware, in software or a combination of hardware and software. The MOS estimator 100 can be implemented in user equipment, such as a mobile téléphoné, tablet, desktop, netbook, multimedia player, video streaming server, settop box or computer. The MOS estimator 100 may also be implemented in a network device in the form of or connected to a network node, such as radio base station, in a communication network or system.

Although the respective units disclosed in conjunction with Figure 10 hâve been disclosed as physically separate units in the device, where ail may be spécial purpose circuits, such as ASICs (Application Spécifie Integrated Circuits), alternative embodiments of the device are possible where some or ail of the units are implemented as computer program modules running on a general-purpose processor. Such an embodiment is disclosed in Figure 11.

Figure 11 schematically illustrâtes an embodiment of a computer 150 having a processing unit 110 such as a DSP (Digital Signal Processor) or CPU (Central Processing Unit). The processing unit 110 can be a single unit or a plurality of units for performing different steps of the method described herein. The computer also comprises an input/output (I/O) unit 120 for receiving a vector of per-time-unit scores of video quality, a vector of per-time-unit scores of audio quality, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration. The I/O unit 120 has been illustrated as a single unit in Figure 11 but can likewise be in the form of a separate input unit and a separate output unit.

Furthermore, the computer 150 comprises at least one computer program product 130 in the form of a non-volatile memory, for instance an EEPROM (Electrically Erasable Programmable Read-Only Memory), a flash memory or a disk drive. The computer program product 130 comprises a computer program 140, which comprises code means which, when run on the computer 150, such as by the processing unit 110, causes the computer 150 to perform the steps of the method described in the foregoing in connection with Figure 2.

The embodiments described above are to be understood as a few illustrative examples of the present invention. It will be understood by those skilled in the art that various modifications, combinations and changes may be made to the embodiments without departing from the scope of the present invention. In particular, different part solutions in the different embodiments can be combined in other configurations, where technically possible.

Aggregation code

The Python code below summarizes the algorithm for estimating MOS, according to the embodiments of the present invention:

def aggregationll(mosV, mosA, buflnit, bufLength, bufStart):

# mosV and mosA are vectors of 1-sec scores, index 0 is start of video or audio # buflnit is seconds of initial buffering # bufLength is a vector of rebuffering lengths # bufStart is a vector of rebuffering start times # cO - Dummy # cl-c3 - Adaptation memory weights il c4-c6 - Initbuf memory weights # c7-c9 - Buffering memory weights fl clO-cll - Initbuf impact

K cl2-cl3 - Rebuf impact # c!4-cl6 - Répétition annoyance # C17-C1S - Audio/video merging weights # cl9-c20 - Oscillation weights

Il c21 - Last part bias (not used) # c22-23 - Négative bias coefs c = (0, 0.2055, 10.256, 17.85, 0.054304, 10.206, 9.0766, 0.17267, 10, 17.762, 4.5327, 1.0054, -67.632, 156.10, 4.9094, 2.1274, 2.0001, 0.16233, -0.013804, 2.1944, 43.565, 0.13025, 9.1647, 0.74811) mosLength = np.minimum(len(mosV), len(mosA)) suml =0 sum2 -0 mosBoth = list(mosV) for i in range(mosLength): mosBoth[iJ * (1 * (mosVfi] - 1) + c[17] * (mosAfi] - 1) + c[1Θ] * (mosV[i] 1) * (mosA[i] -1) /4) / (1 + c[17] + c(18 ) ) +1 mosTime = mosLength - i- 1 mosWeight = exponentiel([1, c[l], 0, c[2j], mosTime) suml += mosBoth[i] * mosWeight sum2 += mosWeight mosBasic = suml / sum2 ose = 0 offset = 0.1 state - 0 for i in range(mosLength):

if State != 1: # State = unknown or dip if mosBoth[i] > mosBasic + offset:

ose +~ 1 state = 1 elif state != -1: # State = unknow or top if mosBoth[i] < mosBasic - offset:

ose += 1 state =-1 oscRel = ose / mosLength oscRel = np.minimum(oscRel, 0.2) # Limit to one change per 5 sec oscDeg - np.power(oscRel * np.std(mosBoth, ddof=l), c[19]) * c[20] mosOffset = list(mosBoth) for i in range(mosLength):

mosTime = mosLength-i-1 mosWeight = exponential([1, c[l], 0, c[2] ] , mosTime) mosOffset[i] = (mosOffset[i] - mosBasic)‘mosWeight mosPerc = np.percentile(mosOffset, c[22], interpolation='linear') # Should normally be négative negBias = np.maximum(0, -mosPerc) negBias = negBias*c[23] lengthDeg = sigmoid([0, 4, c[10], c[10] + c [ 11 ] ] , buflnit) memoryDeg = exponential([1, c[4], 0, c[5]), mosLength) initDeg = lengthDeg‘memoryDeg bufDeg = 0;

for j in range(len(bufLength)):

lengthDeg = sigmoid([0, 4, c[12J, c [ 12]+c [ 13]), bufLength [ j]) repeatDeg = sigmoid((l, c[14], c(15], c [ 15]+c [ 16]], j) memoryDeg = exponential((1, c [ 1], 0, c[8]], mosLength - bufStartfj]) bufDeg = bufDeg + lengthDeg * repeatDeg * memoryDeg bufDeg = bufDeg/4 * (mosBasic-1) # Convert to relative change mos = mosBasic - initDeg - bufDeg - oscDeg - negBias i f mos < 1 :

mos = 1 if mos > 5:

mos = 5 return (mos) def sigmoid{par, x) :

scalex = 10 / (par[3] - par[2]} midx = (par[2] + par[3]) / 2 y = par[0] + (par[l] - par[OJ) / (1 + np.exp(-scalex * (x - midx})) return y def exponential(c, x):

z = np.log(0.5) / (-[c[3] - c[2))) y - c[1] + <c[0] - c[l]) * np.exp(-(x - c(2]) * z) return y

Claims

1. A method, performed by a Mean Opinion Score, MOS, estimator, for predicting a multimedia session MOS, wherein the multimedia session comprises a video session and an audio session, wherein video quality is represented by a vector of per-time-unit scores of video quality and wherein audio quality is represented by a vector of per-time-unit scores of audio quality, and wherein the multimedia session is represented by a vector of rebuffering start times of each rebuffering event, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration being the time between an initiation of the multimedia session and a start time of the multimedia session, the method comprising:

- generating audiovisual quality features from the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality, the audiovisual quality features comprising:

- a weighted combination ofthe per-time-unit scores of audiovisual quality, wherein the weights are exponential fonctions of a time since the start time of multimedia session and a multimedia session duration;

- a négative bias representing how a sudden drop in per-time-unit scores of audiovisual quality affects the multimedia session MOS, the négative bias being calculated as:

negBias = max(o, — lOt/ι percentile of [per - time - unit scores of audiovisual quality [t] · (c[l] + (1 - c[l]) e -Φ) J])c[23], wherein c[1 ], c[2] and c[23] are given coefficients, t îs time since the start time of multimedia session and T is the multimedia session duration; and

- a term representing a dégradation due to oscillations in the per-time-unit-scores of audiovisual quality;

- generating buffering features from the vector of rebuffering start times of each rebuffering event, calculated from the start time of multimedia session, and the vector of rebuffering durations of each rebuffering event;

estimating a multimedia session MOS from the generated audiovisual quality features and the generated buffering features.

2. The method according to claim 1, wherein the term representing a dégradation due to oscillations in the per time unit scores of audiovisual quality is calculated as the number of occurrences when the absolute différence between the per time unit scores of the audiovisual quality and the weighted combination of the per time unit scores of audiovisuai 5 quality exceeds a given threshold value, divided by the multimedia session duration.

3. The method according to any of the previous claims, wherein the generated buffering features comprise a term representing a dégradation due to initial buffering and a term representing a dégradation due to rebuffering.

4. The method according to claim 3, wherein the term representing dégradation due to

10 initial buffering is modeled as a product of a term representing an initial buffering impact and a term representing a forgetness factor impact, wherein the initial buffering impact is a sigmoid function of the initial buffering duration, and the forgetness factor is an exponential function of the time since the start time of multimedia session.

5. The method according to claim 3, wherein the term representing dégradation due to rebuffering is modeled as a sum, over ail rebuffering events, of products of a rebuffering duration impact, a rebuffering répétition impact, and an impact of time since the last rebuffering ended, wherein the rebuffering duration impact is a sigmoid function of a rebuffering duration, the rebuffering répétition impact is a sigmoid function of a rebuffering répétition number, and the impact of time since the last rebuffering ended is an exponential function of the time since the last rebuffering ended.

6. The method according to any of the preceding claims, wherein the multimedia session MOS is estimated as the différence between the weighted combination of the per-time-unit scores of audiovisual quality and the sum of: the négative bias, the term representing dégradation due to oscillations in the per-time-unit-scores of audiovisual quality, the term

25 representing dégradation due to initial buffering, and the term representing dégradation due to rebuffering.

7. A Mean Opinion Score, MOS, estimator, for predicting a multimedia session

MOS, wherein the multimedia session comprises a video session and an audio session, wherein video quality is represented by a yqctor of per-time-unit scores of video quality and wherein audio quality is represented by a vector of per-time-unit scores of audio quality, and wherein the multimedia session is represented by a vector of rebuffering start times of each rebuffering event, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration being the time between an initiation ofthe multimedia session and a start time of the multimedia session, the MOS estimator comprising processing means operative to:

generate audiovisual quality features from the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality, the audiovisual quality features comprising:

- a weighted combination of the per-time-unit scores of audiovisual quality, wherein the weights are exponential functions of a time since the start time of multimedia session and a multimedia session duration;

negBias = max i 0, —lOt/ι percentile of per — time — unit scores of audiovisual quality [t] c[l] + (1 - c[l]) · e (T—t) log(0.5)\ 1 c[23], wherein c[1 ], c[2] and c[23] are given coefficients, t is time since the start time of multimedia session and T is the multimedia session duration; and

generate buffering features from the vector of rebuffering start times of each rebuffering event, calculated from the start time of multimedia session, and the vector of rebuffering durations of each rebuffering event;

estimate a multimedia session MOS from the generated audiovisual quality features and the generated buffering features.

8. The MOS estimator according to claim 7, wherein the term representing a dégradation due to oscillations in the per time unit scores of audiovisual quality is calculated as the number of occurrences vyhen the absolute différence between the per time unit scores of the audiovisual quality and the weighted combination of the per time unit scores of audiovisual quality exceeds a given threshold value, divided by the multimedia session duration.

9. The MOS estimator according to any of claims 7-8, wherein the generated buffering features comprise a term representing a dégradation due to initial buffering and a term representing a dégradation due to rebuffering.

10. The MOS estimator according to claim 9, wherein the term representing dégradation q due to initial buffering is modeled as a product of a term representing an initial buffering impact and a term representing a forgetness factor impact, wherein the initial buffering impact is a sigmoid function ofthe initial buffering duration, and the forgetness factor is an exponential function of the time since the start time of multimedia session.

11. The MOS estimator according to claim 9, wherein the term representing dégradation

- due to rebuffering is modeled as a sum, over ail rebuffering events, of products of a rebuffering duration impact, a rebuffering répétition impact, and an impact of time since the last rebuffering ended, wherein the rebuffering duration impact is a sigmoid function of a rebuffering duration, the rebuffering répétition impact is a sigmoid function of a rebuffering répétition number, and the impact of time since the last rebuffering ended is an exponential

0 function of the time since the last rebuffering ended.

12. The MOS estimator according to any ofthe preceding claims, wherein the multimedia session MOS is estimated as the différence between the weighted combination of the pertime-unit scores of audiovisual quality and the sum of: the négative bias, the term representing dégradation due to oscillations in the per-time-unit-scores of audiovisual

- quality, the term representing dégradation due to initial buffering, and the term representing dégradation due to rebuffering.

13. A computer program for a Mean Opinion Score, MOS, estimator, for predicting a multimedia session MOS, wherein the multimedia session comprises a video session and an audio session, wherein video quality is represented by a vector of per-time-unit scores of video quality and wherein audio quality is represented by a vector of per-time-unit scores of audio quality, and wherein the multimedia session is represented by a vector of rebuffering start times of each rebuffering event, a vector of rebuffering durations of each rebuffering event, and an initial buffering duration being the time between an initiation of the multimedia session and a start time of the multimedia session, the computer program comprising a computer program code which, when run on a computer, causes the computer to:

- generate audiovisual quality features from the vector of per-time-unit scores of video quality and the vector of per-time-unit scores of audio quality, the audiovisual quality features comprising:

- a terrn representing a dégradation due to oscillations in the per-time-unit-scores of audiovisual quality;

14. A computer program product for a MOS estimator comprising a computer program for a MOS estimator according to claim 13 and a computer readable means on which the computer program for a MOS estimator is stored.