EP4604114A1 - Brightness range adaptation for computers - Google Patents

Brightness range adaptation for computers

Info

Publication number
EP4604114A1
EP4604114A1 EP24157639.6A EP24157639A EP4604114A1 EP 4604114 A1 EP4604114 A1 EP 4604114A1 EP 24157639 A EP24157639 A EP 24157639A EP 4604114 A1 EP4604114 A1 EP 4604114A1
Authority
EP
European Patent Office
Prior art keywords
brightness
color
output
operating system
code
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24157639.6A
Other languages
German (de)
French (fr)
Inventor
Danny BERENDSE
Mark Jozef Willem Mertens
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Koninklijke Philips NV
Original Assignee
Koninklijke Philips NV
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Koninklijke Philips NV filed Critical Koninklijke Philips NV
Priority to EP24157639.6A priority Critical patent/EP4604114A1/en
Priority to PCT/EP2025/053542 priority patent/WO2025172273A1/en
Publication of EP4604114A1 publication Critical patent/EP4604114A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09GARRANGEMENTS OR CIRCUITS FOR CONTROL OF INDICATING DEVICES USING STATIC MEANS TO PRESENT VARIABLE INFORMATION
    • G09G5/00Control arrangements or circuits for visual indicators common to cathode-ray tube indicators and other visual indicators
    • G09G5/10Intensity circuits
    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09GARRANGEMENTS OR CIRCUITS FOR CONTROL OF INDICATING DEVICES USING STATIC MEANS TO PRESENT VARIABLE INFORMATION
    • G09G5/00Control arrangements or circuits for visual indicators common to cathode-ray tube indicators and other visual indicators
    • G09G5/14Display of multiple viewports
    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09GARRANGEMENTS OR CIRCUITS FOR CONTROL OF INDICATING DEVICES USING STATIC MEANS TO PRESENT VARIABLE INFORMATION
    • G09G2320/00Control of display operating conditions
    • G09G2320/02Improving the quality of display appearance
    • G09G2320/0271Adjustment of the gradation levels within the range of the gradation scale, e.g. by redistribution or clipping
    • G09G2320/0276Adjustment of the gradation levels within the range of the gradation scale, e.g. by redistribution or clipping for the purpose of adaptation to the characteristics of a display device, i.e. gamma correction
    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09GARRANGEMENTS OR CIRCUITS FOR CONTROL OF INDICATING DEVICES USING STATIC MEANS TO PRESENT VARIABLE INFORMATION
    • G09G2320/00Control of display operating conditions
    • G09G2320/06Adjustment of display parameters
    • G09G2320/0673Adjustment of display parameters for control of gamma adjustment, e.g. selecting another gamma curve
    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09GARRANGEMENTS OR CIRCUITS FOR CONTROL OF INDICATING DEVICES USING STATIC MEANS TO PRESENT VARIABLE INFORMATION
    • G09G2340/00Aspects of display data processing
    • G09G2340/06Colour space transformation
    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09GARRANGEMENTS OR CIRCUITS FOR CONTROL OF INDICATING DEVICES USING STATIC MEANS TO PRESENT VARIABLE INFORMATION
    • G09G2340/00Aspects of display data processing
    • G09G2340/10Mixing of images, i.e. displayed pixel being the result of an operation, e.g. adding, on the corresponding input pixels
    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09GARRANGEMENTS OR CIRCUITS FOR CONTROL OF INDICATING DEVICES USING STATIC MEANS TO PRESENT VARIABLE INFORMATION
    • G09G2352/00Parallel handling of streams of display data
    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09GARRANGEMENTS OR CIRCUITS FOR CONTROL OF INDICATING DEVICES USING STATIC MEANS TO PRESENT VARIABLE INFORMATION
    • G09G2370/00Aspects of data communication
    • G09G2370/20Details of the management of multiple sources of image data

Definitions

  • the invention relates to coordinating the brightnesses of pixels in various visual assets for display, in particular specifically for computer environments, in which various visual assets of different maximum luminance, or maximum brightness relative to a reference level, are combined in a total view. These assets may lie in different windows, and be generated or managed by different applications which run concurrently.
  • LDR Low Dynamic Range
  • SDR Standard Dynamic Range
  • videos i.e. temporally successive sequences of images
  • still images such as e.g. graphics, e.g. in games.
  • Colorimetrically i.e. regarding the specification of the pixel colors, it was based on the technology which already worked fine decades before for photographic materials and paintings: one merely needed to be able to define, and display, most of the colors projecting out of an axis of achromatic colors (a.k.a.
  • the Munsell tree can be used to characterize the color of an object one finds in the world, e.g. a colored stone. The real world is different: from the darkest corner at night, till the middle of a supernova, colors (and their brightnesses) can be almost anything, which would be represented in an infinite cylinder rather than a limited gamut like e.g. a diamond-shaped gamut.
  • the earliest television standards (NTSC, PAL) communicated the color components as three voltage signals (which defined the amount of a color component between 0 and 700mV), where the time positions along the voltage signal corresponded by using a scan path with pixels on the screen.
  • the control signals generated at the creation side directly instructed what the display should make as proportion (apart from their being an accidental fixed gamma pre-correction at the transmitter, because the physics of the cathode ray tube took approximately a square power of the input voltage, which would have made the dark colors much blacker than they were intended e.g. as seen by a camera) at the creation side). So a 60%, 30%, 25% color (which is a dark red) in the scene being captured, would look substantially similar on the display, since it would be re-generated as a 60%, 30%, 25% color (note that the absolute brightness didn't matter much, since the eye of the viewer would adapt to the white and the average brightness of the colors being displayed on the screen).
  • RGB and YCbCr are easy ones, namely they can be calculated into each other by using a simple fixed 3x3 matrix (the coefficients of which depend on the emission spectra of the three primaries, and are standardized, i.e. also to be emulated electronically internally by LCDs which actually may have different optical characteristics, so that from image communication point of view all SDR displays are alike).
  • relative brightness as a percentage of something, which may be undefined until a choice is made, e.g. by the consumer buying a certain display, and setting its brightness setting to e.g. 120%, which will make e.g. the backlight emit a certain amount of light, and so also the white and colored pixels
  • absolute brightness on the other hand absolute brightness.
  • the latter can be characterized by the universal physical quantity luminance (which is measured technically in the unit nit , which is also candela per square meter).
  • the luminance can be stated as an amount of photons coming out of a patch on an object, such as a pixel on the screen, towards the eye (and it is related to the lighting concept of illuminance, since such patch will receive a certain illuminance, and send some fraction of it towards the viewer).
  • a color space is the mathematical 3D space to represent colors (as geometric positions of coordinate numbers), the base of which being defined by the 3 primaries; for the technical discussion we may better use the word color gamut , which is the set of all colors that can be technically defined or displayed (i.e. a space may be e.g. a 3D coordinate system going to infinity whereas the gamut may be a cube of some size in that space); for the brightness aspect only, be will talk about brightness or luminance range (more commonly worded as " dynamic range ", spanning from some minimum brightness to its maximum brightness).
  • the present technologies will not primarily be about chromatic (i.e. color per se, such as more specifically its saturation) aspects, but rather about brightness aspects, so the chromatic aspects will only be mentioned to the extent needed for the relevant embodiments.
  • HDR High Dynamic Range
  • HDR images can represent brighter colors than SDR images, so in particular brighter than white colors (glowing whites e.g.). Or in other words, the dynamic range will be larger.
  • the SDR signal can represent a dynamic range of 1000: 1 (how much dynamic range is actual visible when displaying, will depend inter alia on the amount of surround light reflecting on the front of the display). So if one wants to represent e.g. 10,000: 1, one must resort to making a new HDR image format definition (we may in general use the word signal if the image representation is being or to be communicated rather than e.g. merely existing in the creating IC, and in general signaling will also have its own formatting and packaging, and may employ further techniques depending on the communication mechanism such as modulation).
  • a HDR image may be associated with a metadatum called mastering display white point luminance (MDWPL), a.k.a. ML_V .
  • MDWPL mastering display white point luminance
  • ML_V ML_V
  • This value which is typically communicated in metadata of the signal, and is a characterizer of the HDR video images (rather than of a specific display, as it is an element of a virtual display associated specifically with the video, being some ideal intended display for which the video pixel colors have been optimized to be conforming).
  • This is an electable parameter of the video, which can be contemplated as similar to the election of the painting canvas aspect ratio by a painter: first the painter chooses an appropriate AR, e.g.
  • the creator will then, after having established that the MDWPL is e.g. 5000 nit, make his secondary elections that in a specific scene this lamp shade should be at 700 nit, the flames in the hearth distributed around 500 nit, etc.
  • the MDWPL is e.g. 5000 nit
  • HDR image creation can also involve deeper blacks, up to as deep as e.g. 0.0001 nit (although that is mostly relevant for dark viewing environments, such as in cinema theatres).
  • the other objects e.g. the objects which merely reflect the scene light, will be coordinated to be e.g. at least 40x darker in a 5000 nit MDWPL graded video, and e.g. at least 20x darker in a 2000 nit video, etc. So the distribution of all image pixel luminances will typically depend on the MDWPL value (not making most of the pixels very bright).
  • Relative brightness systems can code brighter pixels compared to some reference relative brightness or luminance level. E.g., one may use the level 100% to indicate the classical Lambertian reflection base lighting white of legacy LDR image representations, and denote brighter whites or colors with higher percentages (e.g. 1000% white is 10x brighter than the normal LDR white; and if one were to map the normal LDR white to 100 nit, one would typically map the 1000% white to 1000 nit).
  • B_relative is a float number ranging from 0 to 1.0, so will Y_float.
  • signal value Y_float is quantized, because we want 8 bit digital representations, ergo, the Y_dig value that is communicated to receivers over e.g. airways DVB (or internet-supplied video on demand, or blu-ray disk, etc.) has a value between 0 and 255 (i.e. power(2;8)-1).
  • the vertical axis represents normalized luminances (in a linear gamut representation), or normalized lumas (in a non-linear representation (coding) of those luminances, e.g via a psychovisually uniformized OETF or its inverse the EOTF). Since after normalization (i.e. division by the respective MDWPL values, e.g. 2000 for an HDR image of a particular video and 100 for an SDR image), the common representation will become easy, and one can define luminance (or luma) mapping functions on normalized axes as shown in Fig. 1D (i.e.
  • the receiving side in the old days, or today will know it has an SDR video, if it gets this format.
  • the maximum white (of SDR) will be by definition the brightest color that SDR can define. So if one now wants to make brighter image colors (e.g of real luminous lamps), that should be done with a different codec (as one can show the math of the Rec. 709 OETF allows only a coding of up to 1000:1 and no more).
  • HDR codecs start by defining an Electro-optical transfer function instead of its inverse, the OETF. Then one can at least basically define brighter (and darker) colors. That as such is not enough for a professional HDR coding system, since because it is different from SDR, and there are even various flavors, one wants more (new compared to SDR coding) technical information relating to the HDR images, which will be metadata.
  • HDR EOTFs are much steeper, to encode a much larger range of needed to be coded HDR luminances, and a significant part of that range coding specifically darker colors (relatively darker, since although one may be coding absolute luminances with e.g. the Perceptual Quantizer (PQ) EOTF (standardized in SMPTE 2084) , one applies the function after normalization).
  • PQ Perceptual Quantizer
  • EOTFs standardized in SMPTE 2084
  • the EOTF will be able to decode the pixel lumas in the plane of lumas spanning the image (i.e. having a width of e.g. 4000 pixels and a height of 2000), which will simply be binary numbers.
  • HDR images will also have a larger word length, e.g. 10 bit.
  • a linear (bit-represented) code e.g. a DMD pixel, indeed to reach e.g.
  • the receiving side may in both situations get as input a coded pixel color (luma and Cb, Cr; or in some systems by matrixing equivalent non-linear R'G'B' component values) which lie between 0 and 255, or 0 and 1023, but it will know the kind of signal it is getting (hence what ought to be displayed) from the metadata, such as the metadata (e.g. MPEG VUI metadata) co-communicated EOTF (e.g.
  • Fig. 1 for a typical nice HDR scene image, of a monster being fought in a cave with a flame thrower, the master grading (Mstr_HDR) of which is shown spatially in Fig. 1A , and the range of occurring pixel luminances on the left of Fig. 1C ).
  • the master grading or master graded image is where the image creator can make his image look as impressive (e.g. realistic) as desired.
  • the image creator can make his image look as impressive (e.g. realistic) as desired.
  • a baker's shop window look somewhat illuminated by making the yellow walls somewhat brighter than paper white e.g.
  • the light bulbs can be made 900 nit (which will give a really lit Christmas-like look to the image, instead of a dull one in which all lights are clipped white, and not much more bright than the rest of the image objects, such as the green of the Christmas tree).
  • SDR (and its coding and signaling) was designed to be able to communicate any Lambertian reflecting color (i.e. a typical object, like your blue jeans pants, which absorbs some of the infalling light, e.g. the red and green wavelengths, to emit only blue light to the viewer or capturing camera) under good uniform lighting (of the scene where the action is camera-captured).
  • any Lambertian reflecting color i.e. a typical object, like your blue jeans pants, which absorbs some of the infalling light, e.g. the red and green wavelengths, to emit only blue light to the viewer or capturing camera
  • uniform lighting of the scene where the action is camera-captured
  • a chromaticity is composed of a certain (rotation angle) hue h (e.g. bluish-green e.g. "teal”), and a saturation sat, which is the amount of pure color mixed in a grey, e.g.
  • the two dotted horizontal lines represent the limitations of the SDR codable image, when associating 100 nit with the 100% of SDR white.
  • the monster will be strongly illuminated by the light of the flames, so we will give it an (average) luminance of 300 nit (with some spread, due to the square power law of light dimming, skin texture, etc.).
  • the soldier may be 20 nit, since that is a nicely slightly dark value, still giving some good basic visibility.
  • a vehicle may be hidden in some shadowy corner, and therefore in a archetypical good impact HDR scene of a cave e.g. have a luminance of 0.01 nit.
  • the camera operator would open his iris so that the soldier comes out at "20 nit", or in fact more precisely 20%. Since the flames are much brighter (note: we didn't actually show the real world scene luminances, since master HDR video Mstr_HDR is already an optimal grading to have best impact in a typical living room viewing scenario, but also in the real world the flames would be quite brighter than the soldier, and certainly the vehicle), they would all clip to maximum white.
  • HDR maximum luminance i.e. ML_V
  • EOTF e.g. PQ for coding
  • the problem is that, unless the receiving side has a display which can display pixels at least as bright as 5000 nit, there is still a question of how to display those pixels.
  • Some (DR adaptation) luminance down-mapping must be performed in the TV, to make darker pixels which are displayable.
  • the display has a (end-user) display maximum luminance ML_D of 1500 nit, one could somehow try to calculate 1200 nit yellow pixels for the flame (potentially with errors, like some discoloration, e.g. changing the oranges into yellows).
  • This luminance down-mapping is not really an easy task, especially to do very accurately instead of sufficiently well, and therefore various technologies have been invented (also for the not necessarily similar task of luminance up-mapping, to create an output image of larger dynamic range and in particular maximum luminance than the input image).
  • mapping function (generically, i.e. used for simplicity of elucidation) of a convex shape in a normalized luminance (or brightness) plot, as shown in Fig. 1D .
  • Both input and output luminances are defined here on a range normalized to a maximum equaling one, but one must mind that on the input axis this one corresponds to e.g. 5000 nit, and on the output axis e.g. 200 nit (which to and for can be easily implemented by division respectfully multiplication).
  • the darkest colors will typically be too dark for the grading with the lower dynamic range of the two images (here for down-conversion shown on the vertical output axis, of normalized output luminances L_out, the horizontal axis showing all possible normalized input luminances L_in).
  • Ergo to have a satisfactory output image corresponding to the input image, we must relatively boost those darkest luminances, e.g. by multiplying by 3x, which is the slope of this luminance compression function F_comp for its darkest end. But one cannot boost forever if one wants no colors to be clipped to maximum output, ergo, the curve must get an increasingly lower slope for brighter input luminances, e.g. it may typically map input 1.0 to output 1.0. In any case the luminance compression function F_comp for down-grading will lie above the 45 degree diagonal (diag) typically.
  • the general desired shape for the brightening of the colors may still be the function F_comp (e.g. determined by the video creator, when grading a secondary image corresponding to his master HDR image already optimally graded), one wants a more savvy down-mapping.
  • F_comp e.g. determined by the video creator, when grading a secondary image corresponding to his master HDR image already optimally graded
  • Fig. 1B for many scenarios one may desire a re-grading which merely changes the brightness of the normalized luminance component (L), but now the innate type of color, i.e. its chromaticity (hue and saturation). If both SDR and HDR are represented with the same red, green and blue color primaries, they will have a similarly shaped gamut tent, only one being higher than the other in absolute luminance representation.
  • both gamuts will exactly overlap.
  • the desired mapping from a HDR color C_H to a corresponding output SDR color CL (or vice versa) will simply be a vertical shifting, whilst the projection to the chromaticity plane circle stays the same.
  • communication image Im_comm instead of just making some final secondary grading from the master image, e.g. in a television, one can make a lower dynamic range image version for communication , communication image Im_comm .
  • this image was defined with its communication image maximum luminance ML_C equal to 200 nit.
  • the original 5000 nit image can then be reconstructed (a.k.a. decoded) as a reconstructed image Rec_HDR (i.e.
  • the proxy image for communicating actually an image a higher dynamic range (DR_H, e.g. spanning from 0.001 nit to 5000 nit) is an image of a different, lower dynamic range (DR_L).
  • the communication image can even elect the communication image to be a 100 nit LDR (i.e. SDR) image , which is immediately ready (without further color processing) to be displayed on legacy LDR images (which is a great advantage, because legacy displays don't have HDR knowledge on board).
  • SDR 100 nit LDR
  • legacy LDR images which is immediately ready (without further color processing) to be displayed on legacy LDR images (which is a great advantage, because legacy displays don't have HDR knowledge on board).
  • the legacy TV doesn't recognize the MDWPL metadatum (cos that didn't exist in the SDR video standard, so the TV is also not arranged to go look for it somewhere in the signal, e.g. in a Supplemental Enhancement Information message, which is MPEG's mechanism to introduce all kinds of pre-agreed new technical information). It is also not going to look for the function. It just looks at the YCbCr e.g.
  • the image can be tuned for any possible connected tv, i.e. any ML_D, because one can double the function of the coding function FL_enc as some guidance function for the up-mapping from 100 nit Im_comm not to a 5000 nit reconstructed image, but to e.g. a 1500 nit image.
  • the concave function which is substantially the inverse of F_comp (note, for display tuning there is no requirement of exact inversion as there is for reconstruction), will now have to be scaled to be somewhat less steep (i.e.
  • a display adapted luminance mapping function FL_DA will be calculated), since we expand to only 1500 nit instead of 5000 nit.
  • an image of tertiary dynamic range (DR_T) can be calculated, e.g. optimized for a particular display in that the maximum luminance of that tertiary dynamic range is typically the same as the maximum displayable luminance of a particular display.
  • Fig. 2 shows -in general, without desiring to be limiting- a few typical creations of video where the present teachings may be usefully deployed.
  • an intermediate dynamic range format calculate and format all the needed metadata, convert to some broadcasting format like DVB or ATSC, packetize in chunks for distribution, etc. (the etc. indicating there may be tables added for signaling available content, sub-titling, encryption, but at least some of that will be of lesser interest to understand the details of the present technical innovations).
  • the video (a television broadcast in the example) is communicated via a television satellite 250 to a satellite dish 260 and a satellite signal capable set-top-box 261. Finally it will be displayed on an end-user display 263.
  • This display may be showing this first video, but it may also show other video feed, potentially even at the same time, e.g. in Picture-in-Picture windows (or some data of the first HDR video program may come via some distribution mechanism and other data via another).
  • a second production is typically an off-line production.
  • a Hollywood movie but it can also be a show of somebody having a race through a jungle.
  • Such a production may be shot with other optimal cameras, e.g. steadicam 211 and drone 210.
  • the camera feeds (which may be raw, or already converted to some HDR production format like HLG) are stored somewhere on network 212, for later processing.
  • the video may be uploaded to some internet-based video service 251. For professional video distribution this may be e.g. Netflix.
  • a third example is consumer video production.
  • the user will have e.g. when making a vlog a ring lighter 221, and will capture via a mobile phone 220, but (s)he may also be capturing in some exterior location without supplementary lighting. She/he will typically also upload to the internet, but now maybe to YouTube, or TikTok, etc.
  • the display 263 In case of reception via the internet, the display 263 will be connected via a modem, or router 262 or the like (more complicated setups like in-house Wi-Fi and the like are not shown in this mere elucidation).
  • Another user may be viewing the video content on a portable display (271), such as a laptop (or similarly other users may use a non-portable desktop PC), or a mobile phone etc.
  • The may access the content over a wireless connection (270), such as Wi-Fi, 5G, etc.
  • Fig. 3 shows an example of an absolute (nit-level-defined) dynamic range conversion circuit 300 for a (HDR) image or video decoder shown in a video processing circuit chain in Fig. 3C .
  • the encoder would typically work similarly but with inverted functions typically, i.e. the function to be applied being the function of the other side mirrored over the diagonal). It is based on coding a primary image (e.g. a master HDR grading) with a primary luminance dynamic range (DR_Prim) as another (so-called proxy) image with a different secondary range of pixel luminances (DR_Sec).
  • a primary image e.g. a master HDR grading
  • DR_Prim primary luminance dynamic range
  • proxy secondary range of pixel luminances
  • the encoder and all its supply-able decoders have pre-agreed or know that the proxy image has a maximum luminance of 100 nit, this need not be communicated as an SDR_WPL metadatum.
  • the proxy image is e.g. a 200 nit maximum image, this will be indicated by filling its proxy white point luminance P_WPL with the value 200, or similarly for 80 nit etc.
  • the various pixel lumas will typically come in as a luma image plane, i.e. the sequential pixels will have first luma Y11, second Y21, etc. (typically these will be scanned, and the dynamic range conversion circuit will convert pixel by pixel to output pixel color triplets (Y_out, Cb_out, Cr_out).
  • Y_out, Cb_out, Cr_out pixel color triplets
  • Various dynamic range conversion circuits may internally work differently, to achieve basically the same thing: a correctly reconstructed output luminance L_out for all image pixels (the actual details don't matter for this innovation, and the embodiments will focus on teaching only aspects as far as needed).
  • the mapping of luminances from the secondary dynamic range to the primary dynamic range may happen on the luminances themselves, but also on any luma representation (i.e. according to any EOTF, or OETF), provided it is done correctly, e.g. not separately on non-linear R'G'B' components.
  • the internal luma representation need not even be the one of the input (i.e. of Y_in), or for that manner of whatever output the dynamic range conversion circuitry or its encompassing decoder may deliver (e.g. a format luma Y_sigfm for a particular communication format or communication system, "communicating" including storage to a memory, e.g. inside a PC, a hard disk, an optical storage medium, etc.).
  • a luma conversion circuit 301 which turns the input lumas Y_in into perceptionally uniformized lumas Y_pc.
  • Y _ pc log _ 10 1 + RHO WPL _ inrep ⁇ 1 ⁇ power Ln _ in ; 1 / 2.4 / log _ 10 RHO WPL _ inrep
  • RHO WPL _ inrep 1.32 ⁇ power WPL _ inrep / 10000 ; 1 / 2.4
  • the value WPL_inrep is the maximum luminance of the range that needs to be converted to psychovisually uniformized lumas, so for the 100 nit SDR image this value would be 100, and for the to be reconstructed output image (or the originally coded image at the creation side) the value would be 1000.
  • this mapping function had been specifically chosen by the encoder of the image (at least for yielding good quality reconstructability, and maybe also a reduced amount of needed bits when MPEG compressing, but sometimes also fulfilling further criteria like e.g. the SDR proxy image being of correct luminance distribution for the particular scene -a dark cave, or a daytime explosion- on a legacy SDR display, etc.).
  • this function F_dec (or its inverse) will be extracted from metadata of the input image signal or representation, and supplied to the dynamic range conversion circuit for doing the actual per pixel luma mapping.
  • the function F_dec directly specifies the needed mapping in the perceptual luma domain, but other variants are of course possible, as the various conversions can also be applied on the functions.
  • the lower and higher dynamic range image will in general have object pixels of different brightness, and oftentimes at least some of the pixels will have different saturation, but ideally the hue of the pixels in both image versions will be the same).
  • This color function will typically specify a multiplier which has a value dependent on a brightness code Y (e.g. the Y_pc, or other codes in other variants).
  • a multiplier establishment circuit 305 will yield the correct multiplier m for the brightness situation of the pixel being processed.
  • a formatting circuit 310 so that the output color triplet (Y_out, Cb_out, Cr_out) can be converted to whatever needed output format (e.g. an RGB format, or a communication YCbCr format, Y_sigfm, Cb_sigfm, Cr_sigfm).
  • a communication channel 379 which is an HDMI cable
  • such cables typically use PQ-based YCbCr pixel color coding, ergo, the lumas will again be converted from the perceptual domain to the PQ domain by the formatting circuit.
  • a display tuning circuit 380 which calculates ultimate pixel colors and luminances to be displayed at the screen of some display, e.g. a 450 nit tv which some consumer has at home.
  • Fig. 3B for some typical HDR image being an indoors/outdoors image (the geometry and comprised image objects of which are shown in Fig. 3A ).
  • the outdoors luminances may typically be 100 times brighter than the indoors luminances
  • in an actual master graded HDR image it may be better to make them e.g. 10x brighter, since the viewer will be viewing all together on a screen, in a fixed viewing angle, even typically in a dimly illuminated room in the evening, and not in the real world.
  • a convex function as shown in Fig. 1 , or inside luma mapper 302, is used which squeezes in the brighter luminances, due to the limitations of the smaller luminance dynamic range.
  • the indoors objects will display (assuming for the moment an 100 or 200 nit display would faithfully display those luminances as coded, and not e.g. do some arbitrary beautification processing which brightens them) darker, darker than ideally desired, i.e.
  • the outdoors colors may also be somewhat pastellized, i.e. of lowered saturation. But of course if the grader at the creation side has control over all the functions (F_enc, FCOL), hey may balance those features, so that some have a lesser deviation at the detriment of others. E.g. if the outdoors shows a plain blue sky, the grader may opt for making it brighter, yet less blue.
  • the middle graph shows what the lumas would look like for the proxy luminances, and that may typically give a more uniform histogram, with e.g. approximately the same span for the indoors and outdoors image object luminances.
  • the lumas are however only relevant to the extent of coding the luminances, or in case some calculations are actually performed in the luma domain (which has advantages for the size of the word length on the processing circuitry). Note that whereas the absolute formalism can allocate luminances on the input side between zero and 100 nit, one can also treat the SDR luminances as relative brightnesses (which is what a legacy display would do, when discarding all the HDR knowledge and communicated metadata, and looking merely at the 0-255 luma and chroma codes).
  • Fig. 4 elucidates how the computer world looked at HDR, in particular for the calculation of HDR scenes, e.g. by ray-tracing (i.e. the equivalent of actual camera capturing).
  • This technology which can be used in e.g. gaming, in which a different view on an environment has to be re-calculated each time a player moves, potentially with a specular reflection appearing that wasn't in view when the player was positioned one game meter to the left, was not primarily geared for communication, e.g. from a broadcaster to a receiver, any receiver, with any of various kinds of displays with different maximum brightness characteristics.
  • Computer systems may be more complex, as there may be various unrelated visual assets, and the viewer may also be using different displays (e.g. an old SDR one, and a new one, and show at least some of the content on one of the displays, but he may also change the display for that content by dragging e.g. its window to the other display).
  • a video processor/processing is usually provided by one manufacturer, and usually follows well-standardized principles
  • a computer may be running applications from just about anybody, and those manufacturers (e.g. of software and its look and feel) may have different visions about higher luminance representation and/or display in more than just minor details.
  • a computer may also be seen as a "kit of parts", and that may be good for its general usability for a myriad of tasks, but that doesn't necessarily mean that these parts would work together in a stable well-defined manner.
  • graphics processing unit is a circuit (or maybe in some systems circuits plural, if there are two or more separate GPUs being supplied with the visual assets to be displayed in totality) which contain the final buffering for such pixels that should be shown on a display, i.e. be scanned out to at least one display (and not processors which do not have this scanout buffer and circuitry for communicating the video signals out to display(s); this GPU is sometimes also called Display Processing Unit DPU).
  • the indicated problems are handled by a computer (700) arranged to run an operating system (504), wherein the operating system is arranged to manage the color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Ass1) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness, wherein the visual asset is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness;
  • TM1 first brightness mapping function
  • the new technical effect hitherto unmanageable is that this approach to in particular the operating system construction and its configurable communication with applications and their visual color asset desiderata, by means of the new color transformation control data (CTctrl), makes it possible that the operating system creates better coordinated colors for the one or more visual assets, for ultimate display. So the final displayed look of all the visual assets together will be more controllable towards a better looking total output picture.
  • CTctrl new color transformation control data
  • Unit in terms of software code may typically mean a sub-process, which will run on an electronic circuit (it will have access to memory for temporarily storing the data, and may have access to reconfigurable processing, such as different algorithms for color transformation, and different manners of deciding the geometric positioning of the various transformed pixels on a canvas, such as a windowing system, the details of the latter being of lesser relevance for understanding the color coordination technology of the present technical system).
  • the brightness mapping unit (710) will receive the pixel color codes of a visual asset (from the geometrical management unit 711) to be brightness mapped to a different brightness range, and apply e.g.
  • This suitable (final) brightness range may be determined based on various parameters, e.g. for which kind of display the composited canvas is generated, user preferences, the kind of total presentation (e.g. a user-interface-focused presentation), etc., but those details are not the core aspects of the present innovation.
  • the original ranges of the assets may be determined by the various applications based on very different criteria. E.g., a game which expects lots of bright explosions to be displayed on the end-user display, may make its graphics (e.g. a notation of successful shots) equally super-bright.
  • a set of lumas respectively colors may for many situations simply means one or more arrays. There will be an array of e.g. 200x100 luma code values (e.g. for a graphics to put on the top-left of the screen), and typically two more arrays for the chromas. In other situations the set of lumas (respectively colors) will be specified by drawing instructions (e.g. draw a line from beginning position (x1,y1) to end position (x2,y2) with luma 422 out of 1023).
  • a color code representing an input luma code means that, whichever the color specification/model used in any detailed embodiment, the coding of the color is specifying colors in the respective gamut (e.g. an SDR gamut, or some HDR gamut). So it may be that the color model directly comprises the input luma code, as in the model which is popular for video: Y'CbCr, in which Y' is the luma code of a pixel. However, all these colors (i.e. the e.g. SDR gamut) can also be represented in another additive color model. E.g., in the computer world the color code R',G',B' is a popular coding.
  • Y' a*R'+b*G'+c*B', in which a,b, and c are fixed constants, depending on the primaries of the gamut.
  • a suitable embodiment of a common format will be based on an SDR color gamut, e.g. a normalized Rec. 709 primaries gamut in which the brightest white luma code may represent 100 nit (e.g. the 8 bit luma code 255 dictates that pixels having these values, are to be displayed at 100 nit on any display, if not transformed into a tertiary brightness).
  • SDR SDR color representation for communication
  • the OS has an actual maximum luminance as a nit number.
  • Tm 1 is able to allocate (during the reconstruction) an original maximum luminance (ML_V) to the largest possibly occurring code of the pixels in Im 1 (or equivalently a percentage of that largest HDR white if the asset only goes as bright as e.g. 40% grey).
  • ML_V original maximum luminance
  • the dynamic range preferably in absolute luminances, or alternatively relative brightnesses
  • the asset resides with its HDR white if the asset actually has white pixels, or compared to its HDR white if it is darker (which can be achieved by giving it lower sRGB values in an Im 1 which forms a proxy for a e.g.
  • an averagely lit piece of paper, casu quo display on a display as a typical SDR-brightness range white, e.g. 200 nit on a display than can go as high as 600 nit) may be coded in the sRGB communicated image e.g. at luma 60%, or 30% (and that level can be communicated as the reference white level WL1).
  • both the original and communicated representation may be using the same amount of bits per color component, e.g. 10 or 12 (or the original may be 3x12 bit and the communicated image Im1 3x8 bit), the difference in luminance range will be apparent from the squeezing together of at least some of the object colors, e.g.
  • TM1 as a function which maps the largest possible communicated pixel luma, i.e.
  • a flag may indicate whether the direct or already pre-inverted form of the brightness mapping function is put in the metadata (or the system may work in a pre-agreed manner).
  • the original colors and their lumas and the luminances (as an amount of nit a.k.a. Cd/m2) or relative brightnesses they code (relative to e.g. the 100% level of SDR white), i.e. of the visual asset as the creating/communicating application ideally wants to see it displayed, may be considerably different than the coded colors as communicated: i.e. the normally decoded colors of the communicated coded proxy colors (using the normal sRGB definition instead of the brightness mapping function TM1 for the decoding) will usually be in a much smaller secondary brightness range, which ends at e.g. 100 nit instead of the original e.g. 20,000 nit.
  • the secondary range as communicated may also be larger than the original range of the visual asset, in particular it may end at a larger maximum luminance than the original maximum luminance of the asset pixels, but usually a smaller brightness range ending at a lowered maximum luminance will perform sufficiently well for communicating assets to the operating system.
  • the operating system can perform suitable tertiary transformations in line with what the original asset's colors were, i.e. are supposed to be displayed as.
  • the operating system may chose the tertiary brightness range to be identical to the primary brightness range, and invert the brightness mapping function (TM1), and use inverted first brightness mapping function (ITM1) on the input luma codes to obtain output luma codes which are a reconstruction of the original lumas of the asset of the application.
  • TM1 brightness mapping function
  • ITM1 inverted first brightness mapping function
  • the operating system may want to coordinate the output luma of the geometric composition of various assets together in a total campus, into a tertiary brightness range which may end lower than a primary brightness range of at least one of the assets of at least one of the applications.
  • the operating system may also take into account what display will be served by the GPU with the composite images of its scanout buffer (ScOBff), and if that display does not have a high displayable brightness range, or the viewer wants to see everything bright near the upper end of the displayable brightness range, the OS may take this into account when firstly electing a tertiary brightness range, and secondly deriving optimal tertiary brightness mapping functions (FL_op) for mapping the various luma code values of the various assets to that common range.
  • ScOBff composite images of its scanout buffer
  • any luma code gets which can equivalently be described as a multiplication by a multiplier which depends on the value of that luma code (a multiplier larger than 1 indicating a -relative if performed in a normalized to 1.0 representation of the lumas or absolute- brightness boost, and a multiplier smaller than one performing a brightness dimming), will depend on where in the input range the luma code falls.
  • the amount of boost a luma say halfway gets may depend on how much the darker lumas get brightened in their mapping.
  • each pixel luma will get an optimal mapping so that the range of input brightnesses (e.g. luminances in some embodiments) is well-represented in the output range, but the many shapes of brightness mapping function that any application or the operating system may chose is also a detail we need not dive into, since the new technical system construction must be able to function which substantially each desired function.
  • range of input brightnesses e.g. luminances in some embodiments
  • many shapes of brightness mapping function that any application or the operating system may chose is also a detail we need not dive into, since the new technical system construction must be able to function which substantially each desired function.
  • some embodiments of the operating system's common mapping may depend only on the communicated at least one brightness mapping function, whilst others may work solely on the communicated at least one reference white level (WL1), whereas other more sophisticated tertiary mappings may design their tertiary mapping function (or algorithm, e.g. taking into account the spatial nature of an asset, such as geometrically non-uniform shading) on both of those communicated color transformation control data elements.
  • WL1 reference white level
  • the brightness mapping function essentially communicates how many luminances respectively lumas of the original asset's colors are squeezed into the typically smaller brightness range of the pixel color array (Im1) that gets communicated, one may know where the brightest street lamp falls in that lower brightness range (namely near 1.0, or 255 in an 8 bit coding of the brightnesses respectively luminances), but one doesn't know yet where the reference level of the uniformly lit (under the average base lighting of the scene, usually the bigger area of the geometrical frame of the images) Lambertian diffusive white object falls. That reference white level (WL1) of the asset might (depending on how much brighter the brightest HDR objects in the original asset representation are) e.g.
  • the first mapping (with TM1) will map the original pixel colors and their lumas, originally lying in the first luminance dynamic range which is determined by whatever the asset was created to be (e.g. a games designer may create a blue laser beam which is as bright as 8000 nit).
  • the secondary brightnesses may be coded as lumas which either code absolute nit secondary color lumas, but which may pragmatically end at a lower maximum luminance, e.g. 100 nit for SDR ( reversible ) common asset communication, or relative/percentual brightnesses.
  • This first mapping establishes the relationship between the original asset colors, and the common interface colors (in the pixel color component arrays) which actually get communicated to the OS.
  • the OS So conversely, for the OS only getting the interface colors, it establishes what the original colors were, and were supposed to be in a displaying. So the maximum brightness of the secondary colors will typically be lower than that of the original colors of the various assets (transforming HDR assets into SDR assets, but in a smart invertible manner by using typically invertible or largely invertible first mapping function(s) TM1). But as regards the tertiary mapping the OS can basically do what it wants (except for it should normally try to fulfil the desired look of the original colors, by at least taking into account the original first mapping function and or reference white level, and typically using functions which keep the color look of most of the colors, such as their differences, still reasonably similar as far as the OS-side, e.g.
  • output-side desiderata enable (which can be elegantly realized e.g. by giving the tertiary mapping function a shape which is similar to the shape of the first brightness mapping, e.g. a weakened down version of that function). E.g. it may lift primarily the darkest colors, and shift the whole secondary brightness range (or even when considering from the original brightness range) upwards to basically much brighter colors for display.
  • geometrical management unit (711) is a unit which manages the geometrical aspects of the assets, and their basic ingestion. So e.g. it will determine the position and possibly scale of assets in the total canvas (but not solely as a set of parameters, such as a top-left (x,y) coordinate pair, but generically also with its content filling, e.g. an image (i.e. the which asset to go where)), potentially in windows and the like. In this innovation, it will also take care of the control of luma processing on demand (to the brightness mapping unit functionality), so they become already of the correct color, and only the geometric aspects of pixel placement are in order. E.g.
  • Brightness mapping unit means any unit that can do brightness processing for one of more pixels of an asset, so that ultimately these pixels will have the desired brightness, either relative to some intermediate or maximum value, or absolute in nits.
  • the application will, for this asset, communicate two brightness mapping function, and instructions where to apply which function (e.g. with a bitmap, where 1 means first function and 0 means the second function should be applied by the operating system, or more precisely, its tertiary function should be based on that second function). Any geometric information needed may be communicated by the application in a geometric data section (Geo1).
  • the problem is less difficult. It may determine an output pixelated image format (ImFinFmt) for writing the geometric composition in the total canvas of the one or more assets into one or more portions of the scanout buffer, e.g. first memory portion 720 and second memory portion 721.
  • ImFinFmt output pixelated image format
  • a common format which is also useful for further communication by the GPU to a display such as a perceptual quantizer luma based format (such a 10 bit format may represent pixel luminances up to 10,000 nit, and even if the original asset had some higher pixel luminances, this may be satisfactory; for relative brightness communications one can still use this coding, by assuming, or explicitly communicating to a display by a further metadatum, that some value is the 100% value, e.g. 100 or 200 nit).
  • a software application is a computer program designed to carry out a specific task other than one relating to the operation of the computer itself, i.e. other than the basic control of the computer hardware. It may implement various functionalities for the user, e.g. online shopping, presentation of visual media, etc.
  • apps for portable apparatuses (the shorthand app can mean all of those).
  • the operating system will work in a manner enabling software applications to specify and communicate the original brightness of any pixel of any asset specifying a luminance of a pixel to be displayed as an amount of nits.
  • This means that such an embodiment will communicate, even if the image array of Im1 itself codes only relative (0-100%) lumas or normalized brightnesses, or only relative non-linear R'G'B' values are communicated, the totality with the color transformation control data allows the receiving operating system to establish for each pixel (i.e. its reconstruction, or some derived re-graded image of pixels) an absolute luminance. This can be performed e.g.
  • the brightness mapping function TM1 in a format which established or allows to establish absolute nit outputs, such as Perceptual Quantizer EOTF luma output domain values for the function (or equivalently in other embodiments one could add additional data to the color transformation control data CTctrl enabling e.g. an absolute scaling of the normalized to 1.0 brightnesses, such as a common multiplier, typically a maximum luminance value for the image or video ML_V).
  • At least one visual asset comprising pixels wherein a pixel has an original color code which specifies an original brightness and to communicate such visual asset to an operating system (504), wherein the original brightness lies in a first brightness range which has a first maximum brightness (ML_V);
  • the application may elect what original HDR format it will use for its asset's colors, but PQ luma-based colors will be a useful manner.
  • the original asset pixel colors may be 0-2000 nit absolute luminance colors, represented as equivalent PQ Y'CbCr or PQ R'G'B' values.
  • the color representation for unified communication is e.g. sRGB, since then the receiving operating system need only understand the original colors per se from the proxy sRGB color, and it may internally generate its output colors (of the output intermediate pixelated image format ImFinFmt) in e.g. Philips EOTF-based format, a logarithmic representation of the luminances or color components, etc.
  • the luminance is an absolute optical quantity (and the other two color components, e.g. Cb and Cr, establish what are also universal color properties, namely a hue like e.g. Chartreuse, and a saturation).
  • the relative systems can be sufficiently unique, since then one has a percentage of some maximum, e.g. a display maximum, or the maximum of a composited total canvas presentation decided by the operating system.
  • the secondary color codes for actual communication of the asset to the operating system will represent a secondary brightness, along a different range of brightnesses, due to the brightness mapping (note that often advantageously the actual mapping processing may be applied to a brightness component per se, but one can also map on other representations, e.g. the RGB components, equivalently, so that the brightness mapping is achieved, in the whichever representation).
  • the important point is the common interfacing, allowing the coordinated use (typically re-grading) of any application's asset by the operating system.
  • the brightnesses are represented as lumas according to some elected EOTF, e.g.
  • the derivation of secondary brightness from the original brightness based on application of a brightness mapping function (TM1) to the original brightness may involve first converting e.g. absolute nit values to input (original) luma codes (e.g. PQ lumas) and then applying the function in the PQ domain to obtain e.g. output PQ lumas (in case output luminances are required, the PQ EOTF can be applied to those output lumas).
  • TM1 brightness mapping function
  • the color transformation control data (CTctrl) is associated with the asset (Ass1), which can happen in many manner, but typically the API will e.g. first communicate the color code array(s), or the data of the procedure to generate a set of pixel colors at a geometrical management unit of the OS, and thereafter the color transformation control data (CTctrl), or vice versa, it first sends the control data and then the pixel color data.
  • the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1) are for use by the operating system to understand which exactly of the many possible HDR colors of the asset the coding of the asset as received (e.g.
  • ML_V1 8000 nit
  • the applications may be supplied to the computer, via some software communication mechanism.
  • the techniques may be embodied as a method of communicating a visual asset having pixel colors to an operating system, comprising the steps of:
  • the techniques may be embodied as a method of operating a computer, comprising a step of running an operating system, wherein the operating system is arranged to manage color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Ass1) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness (ML_V), wherein the visual asset (Ass1) is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness;
  • TM1 first brightness mapping function
  • Fig. 5 we show generically (and schematically to the level of detail needed) a typical application scenario in which the current innovation embodiments would work.
  • client side system i.e. e.g. running on a personal computer or mobile phone
  • OS operating system
  • a mobile phone unless when in a screen casting application, will typically have one display only (second (HDR) display 522), but other client side systems may communicate with several displays.
  • HDR second
  • the web application (502) may be communicating directly with an operating system (504) without an intermediate native application (503), or there may only be a locally installed and running (specially written for some system) native app involved, at least for some of the visual assets being prepared for display, and no contacting with anything over the web regarding those visual assets.
  • the web application could be e.g. an internet banking website, remote gaming, a video on demand site of e.g. Netflix, etc.
  • the web application will run as software on typically a central processing unit (CPU) 501.
  • CPU central processing unit
  • the composition of a web presentation may be received as HTML code.
  • the web application will be to a large extent (for the user interaction) based on visual assets, such as e.g. text, images, or video.
  • Fig. 6 we have again generically (without wanting to be needlessly limiting) shown what the user would see on his one or more display screens, displaying a total viewable area or canvas (the screen buffer 601), containing two windows showing two such applications.
  • the web application may have needs to show its assets in high dynamic range (i.e. better visual quality, in particular a range of brighter pixel luminances than SDR, and correct/controlled pixel luminances).
  • it may show HDR still pictures showing more beautifully than with a plain SDR JPEG to let the customer browse through new available movies.
  • it may want to present commercial material which is already in some HDR format (or vice versa, an old not yet HDR commercial asset, i.e. still in Rec. 709).
  • a first web application may present its content in a first window 610.
  • the content of this window may be on the one hand a first plain text 611 (i.e. graphics colors for standard text, which is well-representable in SDR, e.g. the colors black for the text on an "ivory" background).
  • this webpage which gets composed in the first window for the viewer, may also present a HDR video 612 (e.g. a video streamed in real time from a file location on another server, coded in HLG, and assuming 1000 nit would be a good level for presenting its brightest video pixel colors).
  • This window could be side by side on a total canvas of the windowing server, which we will here call " screen buffer " (to discriminate from local canvases that windows may have).
  • Canvas is a nomenclature from the technical field, which we shall use in this text to denote some geometrical area, of a set of N horizontal by M vertical pixels (or similarly a general object like an oval), in which we can specify ("write” or “draw”) some final pixels, according to a pixel color representation (e.g. 3x8 bit R,G,B).
  • a pixel color representation e.g. 3x8 bit R,G,B
  • GUI Graphical User Interface
  • the operating system may deal with such issues (although it sometimes also expects applications to deal with at least part of the GUI widgets, e.g. remove buttons when the application is in full screen, and re-display them upon an action such as clicking or hovering at the bottom of the screen, or showing the widgets partially transparent, etc.).
  • GUI widgets e.g. remove buttons when the application is in full screen, and re-display them upon an action such as clicking or hovering at the bottom of the screen, or showing the widgets partially transparent, etc.
  • Such widget graphics may be solely in SDR, but that may change in the near future.
  • the two web applications may not know about each other (and then cannot coordinate colors or their brightnesses), but the local computer will (e.g. the windowing server application of the operating system).
  • the second window apart from its second window top bar 623 and second window scroll bar 624, may e.g. comprise a HDR 3D graphics 622 (say fireworks generated by some physical model, with very bright colors, illuminating a commercial), and in another area second text 621, which may now e.g. comprise very bright HDR text colors (e.g. defined on the Perceptual Quantizer scale, say non-linear R_PQ, G_PQ, B_PQ values).
  • the various parts of the screen may correspond to areas and their pixel sets, e.g. area tightly comprising all pixels related to the first window and nothing from the screen background (shown slightly larger to be visible).
  • a particular pastel yellow color of 750 nit gets defined and created (for ultimate display) at a certain pixel, just that this elementary action gets done.
  • a rendering pipeline which usually always consists of e.g. defining triangles, putting these triangles at the correct pixel coordinates, interpolating the colors within the triangle, etc. (and it doesn't matter whether one processor does this, even running on the CPU, or 20 parallel processing cores on the GPU, and where and how they cache certain data, etc.).
  • the web application may specify some intentions regarding colors to use (e.g. for unimportant text, important text, and text that is being hovered over by the cursor), e.g. by using Cascaded Style Sheets (CSS), but it may rely for the rendering of all the colors, in a canvas, on a native application 503. Since applications may desire a well-contemplated and consistent design, CSS is a manner to separate the content (i.e. e.g. what is said in a news article's text) from the definition and communication of its presentation, such as its color palette. For web browsing this native application would be e.g. Edge, or Firefox, etc. So it may call some standard functionality via a first API (API1).
  • API1 first API
  • the native application may do some preparing or processing on the colored text, but it may also rely on some functionality of the operating system by calling a second API (API2). But the operating system may, instead of doing some preparatory calculations on the CPU, and then simply write the resultant pixel colors in the VRAM of the GPU, also rely on the GPU to calculate the colors in the first place (e.g. calculate the fireworks by using some compute-shader). In any case, it may issue one or more function calls via third API (API3).
  • API3 third API
  • modern universal graphical APIs like WebGPU may have the web page already define calls for the GPU (fourth API API11).
  • An example of a command in WebGPU is copyExternalImageToTexture(source, destination, copySize), which copies the contents of a platform canvas or image to a destination texture.
  • the various APIs may also query e.g. what the preferred color format of a canvas is (colors may be compressed according to some format to save on bus communication, needed VRAM), etc.
  • the final buffer for colors of the screen to be shown, in the GPU 510, is the so-called ScanOut Buffer 511. From there the pixels as needed to be communicated to the display, e.g. over HDMI, are to be scanned, and formatted into the needed format (timing, signaling, etc.).
  • a first problem is that, although one may want to rely on other (layer) software for e.g. positioning and drawing a window, or capturing a mouse event, dynamic range conversion is far too complex, and quality-critical, to rely on other software components to do some probably fixed and simplistic luminance or luma mapping.
  • the creator of the web page, or the commercial content running on it, may have spent far too much time making beautiful HDR colors, only to see them mapped with a very coarse luminance mapping, potentially even clipping away relevant information.
  • the various applications might have defined various HDR colors (internally in their application space), but the windowing servers (e.g. the MS windows system) typically considered everything to be in the standard SDR color space sRGB . Because that is what they understood and had been using for decades, and HDR was complex, multi-variant, and ill-understood nor agreed. So you had HDR colors, but you needed to work with SDR colors, or that was what was going to come out to be able to see anything.
  • the windowing servers e.g. the MS windows system
  • one display may refresh at higher rate than the other, and then its part of the total canvas will be read out more frequently.
  • the innovation is about allowing the optimal coordination of the various parts the user ultimately gets to see, via the careful communications and attuned handling of the coordinated reference image representations (e.g. sRGB) and associated re-grading functions for defining HDR assets, and the coordinated use of it all when e.g. display adapting the final composition or part thereof.
  • the coordinated reference image representations e.g. sRGB
  • associated re-grading functions for defining HDR assets
  • a first software application 701 may be showing e.g. images of a motorcycle.
  • this motorcycle is a 2000x1000 pixel image, which may get re-scaled to the total output canvas by the OS.
  • Its asset may be defined primarily by an image Im1, which may be e.g. 3 pixel color component arrays (MPEG compressed, or non-compressed), e.g. an array for the lumas Y', and one for the blue chromas Cb and the red chromas Cr.
  • This asset is supplemented by associated metadata, e.g.
  • a second software application 702 may be showing e.g. an encyclopedic article about bats (either from an internal memory of the computer device, potentially composed on the fly and upon request of a user, or from internet, etc.).
  • third pixel set color specification procedure Tx3 may give a HDR color to ASCII-coded text characters, and another HDR color for the background, both of which will be converted to corresponding sRGB color codes (or the like in similar embodiments), yet those can be re-graded to the HDR colors (preferably both with the same inverse function; though in general they could have their separate transformation).
  • the brightness mapping unit 710 of the OS can, as shown symbolically, both apply upgrading (concave) or downgrading (convex) function to the input sRGB lumas, as the need may be. It will produce the correct tertiary range output lumas Y'o, e.g. typically in a re-graded image ImScal. Finally the GPU can communicate the total composited image to a display 750, via some image communication path 751, e.g. a HDMI cable, Wi-Fi screencasting, etc.
  • the geometrical management unit (711), e.g. a window compositing unit of the OS, is shown as the unit which does the overall asset management, i.e.
  • a first brightness mapping TM1 for the video image Im1 of the monster in the cave
  • WL1 here e.g. luma 64 out of 255 luma codes
  • the linear scale endpoint means 2000 nit; the horizontal input axis are just sRGB input lumas Y'_in_SDR, i.e. approximately the square root of the linear relative brightnesses).
  • the first reference white level WL1 can be used by the OS to e.g. construct a first, lower brightness part TMd_opt of the mapping curve (from input lumas of pixels in Im1, to output luminances).
  • the OS can determine (non-limiting) a "straight" line allocation (we have symbolically drawn a straight line, but since the input is square root, the shape should be approximately power 2), which maps the identified first reference white level WL1 to a first common white level WcomL1 of the composite canvas, and all input linear brightnesses of the Im1 correspondingly linearly proportional below this juncture point PJ1.
  • a "straight" line allocation (we have symbolically drawn a straight line, but since the input is square root, the shape should be approximately power 2), which maps the identified first reference white level WL1 to a first common white level WcomL1 of the composite canvas, and all input linear brightnesses of the Im1 correspondingly linearly proportional below this juncture point PJ1.
  • the brighter input lumas Y'_in_SDR will follow a shape-conforming optimal mapping function TM1_opt, which largely follows the shape guidance of the original communicated brightness mapping function TM1 of the first asset (see further elucidated in Fig. 9 how this can be achieved). Essentially, this means that primarily the flames should be well-boosted (note that the diagonal should be interpreted with a brighter output maximum luminance: if we allocate e.g. 100 nit to the input normalized maximum, mapping according to the diagonal already corresponds to a 20-fold brightening of all pixel luminances, at least if the input axis was linearly represented).
  • the creator of the TM1 function i.e.
  • the first software application and potentially any human behind it, has also created a steep slope where the monster is coded, ergo, this may mean that he wanted the monster to be rather contrasty in any re-grading, ergo this guiding shape should be followed at least as far as achievable (as we will also see in Fig. 9 , the shape of a function, i.e. basically how the output varies for increasing inputs, can elegantly be described by a set of distances of successive points on the diagonal to the locus of point of the function, e.g. PJ1). So the shape of TM1_opt, will at least be based on the communicated TM1 for asset Ass1 (Im1), and often also based on WL1.
  • the first reference white level WL1 can be communicated even though there is no object in the current image(s) that actually has this value, as it can still be used as important reference value by the OS, or the instructing software application.
  • the second asset is placed in the total canvas, after establishing suitable luma re-grading.
  • a third sRGB image Im3 we now assume without limitation that the text is not ASCII, but already communicated as a pixelated sRGB image comprising a rendering of that text), together with only a third reference white level WL3.
  • the OS sees that the text (e.g. yellow text with luma 175) is actually supposed to be brighter than the communicated third reference white level WL3. So the second software application (or the first SA for this third asset), wanted to show above averagely bright text, which can be characterized by a first contrast Cont1.
  • the OS can decide to respect this, and respect this compared to the tertiary brightness range of the total composition.
  • this processing may e.g. establish a characteristic brightness level CHRbriLev_COMP of the brightest regions of the composited canvas, which will contain the flames. E.g. areas or objects can be extracted, and an average output luminance can be determined. This may be the brightness the ultra-bright text of the third asset (i.e. Im3) has to compete with.
  • the OS can decide to derive one or more of a second contrast Cont2 of a to determine second common white level WcomL2 to the characteristic brightness level CHRbriLev_COMP, and a third contrast Cont3 of that second common white level WcomL2 to the first common white level WcomL1.
  • a characteristic brightness level CHRbriLev_COMP of the brightest regions of the composited canvas, which will contain the flames. E.g. areas or objects can be extracted, and an average output luminance can be determined. This may be the brightness the ultra-bright text of the third asset (i.e. Im3) has to compete with.
  • the OS can decide to derive one or more
  • WcomL2 can be lowered below CHRbriLev_COMP the smaller Cont1 is (and e.g. in a linear or non-linear proportion of CHRbriLev_COMP compared to WcomL1), etc.
  • other embodiments can ignore the maximum areas of the video, and merely raise the amount of Cont3 as a linear or non-linear function of how much Cont1 is above WL3.
  • Fig. 9 shows generically an example of an algorithm that the OS can use to map brightnesses to a different brightness range (e.g. 100 nit luminances that were to be reconstructed to 2000 nit luminances, will actually be re-graded to a composite range ending at maximum 1000 nit) using an essentially shape preserving tertiary mapping function TM1_opt.
  • the principle is elucidated with a function that is already composed of three parts, but the OS can segment the function in parts that behave essentially similarly in their mapping, e.g. relative brightening (partitioning algorithms are known, e.g.
  • the change of derivative may also be a good candidate for partitioning; note that the portioning is not necessary in all embodiments, but will be used if some partitions converge or diverge more strongly from the diagonal than others).
  • TM1 the input function
  • This initial distance d1 can be calculated by establishing a direction of projection (the OS can use a fixed direction, e.g. 80 degrees i.e. 10 degrees more slanted than vertical down-projection).
  • a final distance is determined, e.g. 80% of the initial distance for this bottom part of the curve (this ratio will depend, usually in a non-linear manner, corresponding with visual appearance, on the difference between the maximum luminance the original function TM1 was intended for, e.g. 2000 nit, and the current situation maximum, of the composite canvas; e.g.
  • the third function point Pf3 lies below the diagonal, meaning the output range should not be used in its entirety (because for this image, or these images, or this asset, 2000 nit is too much).
  • the OS could in principle deviate to an optimal mapping third point Pc3 which lies deeper than Pf3, but then the re-optimized curve is only partially shape preserving. In general, it will scale upwards, since for a lower output maximum luminance (1000 nit) one does not want to make the brightest pixels too dim.
  • These procedures together form the optimal brightness mapping curve TM1_opt for optimizing the lumas of the asset for the different maximum luminance situation of the tertiary brightness range of the total composite canvas.
  • this function will lie everywhere closer to the diagonal, but is still essentially shape preserving, meaning at least the variations of mapping over the input range, not of course the exact output values for any input (e.g. the middle part is still essentially a large contrast part, because usually there is some object of particular interest there, and the brightness of the brightest object pixels is still essentially kept under moderation).
  • the middle part is still essentially a large contrast part, because usually there is some object of particular interest there, and the brightness of the brightest object pixels is still essentially kept under moderation.
  • all assets that are to be juxtaposed will be mapped to a common tertiary brightness range by using the reference white level to map such level to one or more final white levels, with an equi-luminance ratio (i.e. 60% brightness of that white in the original asset becomes 60% of the chosen final white level, or at least close to that value) for the darker colors, and the (typically HDR effect) brighter colors are mapped by a final brightness mapping function, which distributes the remaining lumas in each asset above its white reference level over the remaining colors in the tertiary brightness range for output to the display, and in a manner which tries to follow the shape of the communicated brightness mapping function for the asset.
  • an equi-luminance ratio i.e. 60% brightness of that white in the original asset becomes 60% of the chosen final white level, or at least close to that value
  • the (typically HDR effect) brighter colors are mapped by a final brightness mapping function, which distributes the remaining lumas in each asset above its white reference level over the remaining colors
  • Fig. 10 gives some further detail on how the geometric graphical presentation of visual assets typically happens internally.
  • a human user 1000 typically interacts with some user app/application (1001). Although some apps could be talking with deeper levels more directly, typically they may be using a more generic graphical user interface language 1002 (a.k.a. shell), like e.g. KDE Plasma or GNOME (Gnu Network Object Model Environment).
  • a graphics protocol GRPROT i.e. a set of API calls to interact with the operating system 1010.
  • An example of a Linux graphics protocol is Wayland.
  • the operating system will typically contain (at least) three parts. Besides the kernel 1005, it may typically a window manager 1004 and a so-called display manager 1005, which managers can talk to each other (e.g. the window manager can call functionality of the display manager).
  • the window manager will take care of the user interaction (e.g. mouse focus), and the state of the windows (position, size, transparency and z-order, window shadow).
  • An example of a window manager for ChromeOS is Ash.
  • An example of a dynamic window manager for the X window system is Awesome.
  • the display server takes care of the low level drawing capabilities, e.g. it can draw a line or an area. So the above described luminance mapping processing may typically be performed by the display manager (or by some capability on request of that display manager, e.g.
  • the API calls of the protocol will according to the present innovation communicate via the basic pixel color array representation and one or more of the luminance mapping function (which can be formulated as a luma mapping function) and the reference white level, so this will typically be communicated over the graphics protocol GRPROT, but in complex operating system behavior it may also communicate in between modules of the OS, and even back to and back from another application etc.
  • the luminance mapping function which can be formulated as a luma mapping function
  • the reference white level so this will typically be communicated over the graphics protocol GRPROT, but in complex operating system behavior it may also communicate in between modules of the OS, and even back to and back from another application etc.
  • MacOS and iOS can use a display server like e.g. the Quartz display server (they can then use a somewhat different window manager).
  • Android can use SurfaceFlinger.
  • the operating system can talk with the hardware (the GPU 1006 and its connection to a display 1007, which may also be bidirectional in case properties of the display need to be polled like its maximum displayable luminance, for optimizing the tertiary range of brightnesses, i.e. of relative brightnesses or absolute luminances), via a GPU API, e.g. use the more generic APIs like e.g. Vulkan (which is an open standard cross-platform API for 3D graphics and computing), or OpenGL, etc.
  • Vulkan which is an open standard cross-platform API for 3D graphics and computing
  • OpenGL etc.
  • the protocol of communication with the GPU can use the present data formulation, but that is another concept that the communication by apps to, and coordination of luminance distributions of assets by, the OS.
  • the present innovative communication of HDR asset and its luminance re-grading desiderata metadata may also be incorporated into generic OS display server/window server abstraction languages.
  • the algorithmic components disclosed in this text may (entirely or in part) be realized in practice as hardware (e.g. parts of an application specific integrated circuit) or as software running on a special digital signal processor, or a generic processor, etc. At least some of the elements of the various embodiments may be running on a fixed or configurable CPU, GPU, Digital Signal Processor, FPGA, Neural Processing Unit, Application Specific Integrated Circuit, microcontroller, SoC, etc.
  • the images may be temporarily or for long term stored in various memories, in the vicinity of the processor(s) or remotely accessible e.g. over the internet.
  • the computer program product denotation should be understood to encompass any physical realization of a collection of commands enabling a generic or special purpose processor, after a series of loading steps (which may include intermediate conversion steps, such as translation to an intermediate language, and a final processor language) to enter the commands into the processor, and to execute any of the characteristic functions of an invention.
  • the computer program product may be realized as data on a carrier such as e.g. a disk, data present in a memory, data travelling via a network connection -wired or wireless-.
  • characteristic data required for the program may also be embodied as a computer program product.
  • Some of the technologies may be encompassed in signals, typically control signals for controlling one or more technical behaviors of e.g. a receiving apparatus, such as a television.
  • Some circuits may be reconfigurable, and temporarily configured for particular processing by software. Some parts of the apparatuses may be specifically adapted to receive, parse and/or understand innovative signals.
  • any reference sign between parentheses in the claim is not intended for limiting the claim.
  • the word “comprising” does not exclude the presence of elements or aspects not listed in a claim.
  • the word “portion” of a set of elements is not intended to exclude that portion may also cover the totality of the elements, because that may function equally in a same manner.
  • the word “a” or “an” preceding an element does not exclude the presence of a plurality of such elements, nor the presence of other elements.
  • “And/or” means that both options may be present together, or one of them may be present alone.
  • the word “e.g.” is typically used to indicate that we mean that something else is also belonging to the possibilities, e.g. a similar element, example, or teaching.
  • i.a means inter alia, or among others.
  • An element between ellipses will normally be used to indicate that something is optional, i.e. also possible as a variant of a more general concept, rather than necessary, e.g. (local) luminance boosting is intended to say, (primarily, as main level teaching) "luminance boosting" in general, which may be for all pixels the same, but may also be different, i.e. of the "local luminance boosting” variant, e.g. only applied to some locality of the image.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Computer Hardware Design (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Processing Of Color Television Signals (AREA)
  • Image Processing (AREA)

Abstract

To allow controlled display of various visual assets of different brightness dynamic range, a computer (700) or method of operation and corresponding software applications are arranged to run an operating system (504), wherein the operating system is arranged to manage color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Ass1) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness (ML_V), wherein the visual asset (Ass1) is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness;
wherein the visual asset in addition comprises color transformation control data (CTctrl) which comprises at least one of first brightness mapping function (TM1) or its inverse, and a reference white level (WL1);
wherein the operating system comprises a brightness mapping unit (710) arranged to receive the visual asset and to apply a tertiary brightness mapping function (FL_op) to the input luma code to obtain an output luma code (Y'o), wherein the output luma code specifies a tertiary brightness which lies in a tertiary brightness range which is different from the secondary brightness range, wherein the output luma codes a brightness of an output color, wherein the tertiary brightness mapping function (FL_op) is based on at least one of the first brightness mapping function (TM1) and the reference white level (WL1);
wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.

Description

    FIELD OF THE INVENTION
  • The invention relates to coordinating the brightnesses of pixels in various visual assets for display, in particular specifically for computer environments, in which various visual assets of different maximum luminance, or maximum brightness relative to a reference level, are combined in a total view. These assets may lie in different windows, and be generated or managed by different applications which run concurrently.
  • BACKGROUND OF THE INVENTION
  • For more than half a century, an image representation/coding technology which is now called Low Dynamic Range (LDR) or. Standard Dynamic Range (SDR) worked perfectly fine for creating, communicating and displaying electronic images such as videos (i.e. temporally successive sequences of images), for e.g. pre-recorded movies or live broadcasts, or still images, such as e.g. graphics, e.g. in games. Colorimetrically, i.e. regarding the specification of the pixel colors, it was based on the technology which already worked fine decades before for photographic materials and paintings: one merely needed to be able to define, and display, most of the colors projecting out of an axis of achromatic colors (a.k.a. greys) which spans from black at the bottom end to the brightest achromatic color giving the impression to the viewer of being white. An example of such a color gamut of all possible colors is the "Munsell tree", which distributes colors of various hues around a vertical axis of achromatic greys. The Munsell tree can be used to characterize the color of an object one finds in the world, e.g. a colored stone. The real world is different: from the darkest corner at night, till the middle of a supernova, colors (and their brightnesses) can be almost anything, which would be represented in an infinite cylinder rather than a limited gamut like e.g. a diamond-shaped gamut. If one represents a single relatively uniformly lit environment, a cut from that cylinder will be well-mappable to the limited (closed) gamut like e.g. the Munsell tree. But in general one may have differently lit environments side by side, and then a white color indoors may be a darker white color than e.g. a white car outdoors being strongly lit by the sun. Human eyes may sometimes see the former as some grey, but in general all those colors will look white (though differently bright whites). If one now wants to represent or display such an environment (ideally), more colors need attention than just the absolute brightest white of any representation of colors in a scene (which is an important characteristic of the representation, but merely indicates the brightest color one still wants to faithfully code and/or process).
  • Apart from theoretical colorimetry concerns, one needs technically simple systems, which not only are capable of relatively faithfully displaying the required colors, but also automatically capture the colors of a scene in the world into such a displayable representation, respectively one would like human operators to be able to change the colors to their liking. For television communication, which relies on an additive color creation mechanism at the display side, a triplet of red, green and blue color components needed to be communicated for each position on the display screen (pixel), since with a suitably proportioned triplet (e.g. 60%, 30%, 25%) one can make almost all colors, and in practice all needed colors (white being obtained by driving the three display channels to their maximum, with driving signal Rmax=Gmax=Bmax).
  • The earliest television standards (NTSC, PAL) communicated the color components as three voltage signals (which defined the amount of a color component between 0 and 700mV), where the time positions along the voltage signal corresponded by using a scan path with pixels on the screen.
  • The control signals generated at the creation side, directly instructed what the display should make as proportion (apart from their being an accidental fixed gamma pre-correction at the transmitter, because the physics of the cathode ray tube took approximately a square power of the input voltage, which would have made the dark colors much blacker than they were intended e.g. as seen by a camera) at the creation side). So a 60%, 30%, 25% color (which is a dark red) in the scene being captured, would look substantially similar on the display, since it would be re-generated as a 60%, 30%, 25% color (note that the absolute brightness didn't matter much, since the eye of the viewer would adapt to the white and the average brightness of the colors being displayed on the screen). One can call this "direct link driving", without further color processing (except for arbitrary and unnecessary processing a display maker might still perform to e.g. make his sky look more blue-ish). For reasons of backwards compatibility with the older black and white television broadcasts, instead of actually communicating a red, green and blue voltage signal, a brightness signal and two color difference signals called chroma were transmitted (in current nomenclature the blue chroma Cb, and the red Cr). The relationship between RGB and YCbCr is an easy one, namely they can be calculated into each other by using a simple fixed 3x3 matrix (the coefficients of which depend on the emission spectra of the three primaries, and are standardized, i.e. also to be emulated electronically internally by LCDs which actually may have different optical characteristics, so that from image communication point of view all SDR displays are alike).
  • These voltage signals were later for digital television (MPEG-based et al.) digitized, as prescribed in standard Rec. 709, and one defined the various amounts of e.g. the brightness component with an 8 bit code word, 0 coding for the darkest color (i.e. black), and 255 for white. Note that coded video signals need not be compressed in all situations (although they oftentimes are). In case they are we will use the wording compression (e.g. by MPEG-HEVC, AV1 etc.), not to be confused with the act of compressing colors in a smaller gamut respectively range. With brightness we mean the part of the color definition that will impact upon the colors (to be) displayed the visual property of being darker respectively brighter. In the light of the present technologies, it is important to correctly understand that there can be two kinds of brightnesses: relative brightness (as a percentage of something, which may be undefined until a choice is made, e.g. by the consumer buying a certain display, and setting its brightness setting to e.g. 120%, which will make e.g. the backlight emit a certain amount of light, and so also the white and colored pixels), and on the other hand absolute brightness. The latter can be characterized by the universal physical quantity luminance (which is measured technically in the unit nit, which is also candela per square meter). The luminance can be stated as an amount of photons coming out of a patch on an object, such as a pixel on the screen, towards the eye (and it is related to the lighting concept of illuminance, since such patch will receive a certain illuminance, and send some fraction of it towards the viewer).
  • Recently two unrelated technologies have emerged, which only came together because people argued that one might as well in one go improve the visual quality of images on all aspects, but those two technologies have quite different technical aspects.
  • On the one hand there was a strive towards wide gamut image technology. It can be shown that only colors can be made which lie within the triangle spanned by the red, green and blue primaries, and nothing outside. But one chose primaries (originally phosphors for the CRT, later color filters in the LCD, etc.) which lay relatively close to the spectral locus of all existing colors, so one could make sufficiently saturated colors (saturation specifies how far away a color lies from the achromatic colors, i.e. how much "color" there is). However, recently one wanted to be able to use novel displays with more saturated primaries (e.g. DCI_P3, or Rec. 2020), so that one also needed to be able to represent colors in such wider color spaces (a color space is the mathematical 3D space to represent colors (as geometric positions of coordinate numbers), the base of which being defined by the 3 primaries; for the technical discussion we may better use the word color gamut, which is the set of all colors that can be technically defined or displayed (i.e. a space may be e.g. a 3D coordinate system going to infinity whereas the gamut may be a cube of some size in that space); for the brightness aspect only, be will talk about brightness or luminance range (more commonly worded as "dynamic range", spanning from some minimum brightness to its maximum brightness). The present technologies will not primarily be about chromatic (i.e. color per se, such as more specifically its saturation) aspects, but rather about brightness aspects, so the chromatic aspects will only be mentioned to the extent needed for the relevant embodiments.
  • A more important new technology is High Dynamic Range (HDR). This should not be construed as "exactly this high" (since there can be many variants of HDR representations, with successively higher range maximum), but rather as "higher than the reference/legacy representation: SDR". Since there are new coding concepts needed, one may also discriminate HDR from SDR by aspects from the technical details of the representation, e.g. the video signal. One difference of absolute HDR systems is that they define a unique luminance for each image pixel (e.g. a pixel in an image object being a white dog in the sun may be 550 nit), where SDR signals only had relative brightness definitions (so the dog would happen look e.g. 75 nit (corresponding to 94%) on somebody's computer monitor which could maximally show 80 nit, but it would display at 234 nit on a 250 nit SDR display (yet the viewer would not typically see any difference in the look of the image, unless having those displays side by side). The reader should not confuse luminances as they exist (ultimately) at the front of any display, with luminances as are defined (i.e. establishable) on an image signal itself, i.e. even when that is stored but not displayed. Other differences are metadata that any flavor of HDR signal may have, but not the SDR signal.
  • Colorimetrically, HDR images can represent brighter colors than SDR images, so in particular brighter than white colors (glowing whites e.g.). Or in other words, the dynamic range will be larger. The SDR signal can represent a dynamic range of 1000: 1 (how much dynamic range is actual visible when displaying, will depend inter alia on the amount of surround light reflecting on the front of the display). So if one wants to represent e.g. 10,000: 1, one must resort to making a new HDR image format definition (we may in general use the word signal if the image representation is being or to be communicated rather than e.g. merely existing in the creating IC, and in general signaling will also have its own formatting and packaging, and may employ further techniques depending on the communication mechanism such as modulation).
  • Depending on the situation, the human eye can easily see (even if all on a screen in front corresponding to a small glare angle) 100,000: 1 (e.g. 10,000 nit maximum and 0.1 minimum, which is a good black for television home viewing, i.e. in a dim room which only support lighting of a relatively low level, such as in the evening). However, it is not necessary that all images as created by a creative go as high: (s)he may elect to make the brightest image pixel in an image or the video e.g. 1000 nit.
  • The luminance of SDR white for videos (a.k.a. the SDR White Point Luminance (WP) or maximum luminance (ML)), is standardized to be 100 nit (not to be confused with the reference luminance of white text in 1000 nit HDR images being 200 nit). I.e., a 1000 nit ML HDR image representation can make up to 10x brighter (glowing) object colors. What one can make with this are e.g. specular reflections on metals, such as a boundary of a metal window frame: in SDR the luminance has to end at 100 nit, making them visually on slightly brighter than the e.g. 70 nit light gray colors of the part of the window frame that does not specularly reflect. In HDR one can make those pixels that reflect the light source to the eye e.g. 900 nit, making them glow nicely giving a naturalistic look to the image as if it were a real scene. The same can be done with fire balls, light bulbs, etc. This regards the definition of images; how a display which can only display whites as bright as 650 nit (the display maximum luminance ML_D) is to actually display the images is an entirely different matter, namely one of display adaptation a.k.a. display tuning, not of image (de)coding. Also the relationship with how a camera captures HDR scene colors may be tight or lose: we will in general assume that HDR colors have already been defined in the HDR image when talking about such technologies as coding, communication, dynamic range conversion and the like. In fact, the original camera-captured colors or specifically their luminances may have been changed into different values by e.g. a human color grader (who defines the ultimate look of an image, i.e. which color triplet values each pixel color of the image(s) should have), or some automatic algorithm. So already for the fact that the chromatic gamut size may stretch with less than a factor 2, whereas the brightness range e.g. luminance range may stretch by a factor 100, one expects different technical rationales and solutions for the two improvement technologies.
  • A HDR image may be associated with a metadatum called mastering display white point luminance (MDWPL), a.k.a. ML_V. This value, which is typically communicated in metadata of the signal, and is a characterizer of the HDR video images (rather than of a specific display, as it is an element of a virtual display associated specifically with the video, being some ideal intended display for which the video pixel colors have been optimized to be conforming). This is an electable parameter of the video, which can be contemplated as similar to the election of the painting canvas aspect ratio by a painter: first the painter chooses an appropriate AR, e.g. 4:1 for painting a landscape, or 1:1 when he wants to make a still life, and thereafter he starts to optimally position all his objects in that elected painting canvas. In an HDR image the creator will then, after having established that the MDWPL is e.g. 5000 nit, make his secondary elections that in a specific scene this lamp shade should be at 700 nit, the flames in the hearth distributed around 500 nit, etc.
  • The primary visual aspect of HDR images is (since the blacks are more tricky) the additional bright colors (so one can define ranges with only a MDWPL value, if one assumes the bottom luminance to be fixed to e.g. 0.1 nit). However, HDR image creation can also involve deeper blacks, up to as deep as e.g. 0.0001 nit (although that is mostly relevant for dark viewing environments, such as in cinema theatres).
  • The other objects, e.g. the objects which merely reflect the scene light, will be coordinated to be e.g. at least 40x darker in a 5000 nit MDWPL graded video, and e.g. at least 20x darker in a 2000 nit video, etc. So the distribution of all image pixel luminances will typically depend on the MDWPL value (not making most of the pixels very bright). Relative brightness systems can code brighter pixels compared to some reference relative brightness or luminance level. E.g., one may use the level 100% to indicate the classical Lambertian reflection base lighting white of legacy LDR image representations, and denote brighter whites or colors with higher percentages (e.g. 1000% white is 10x brighter than the normal LDR white; and if one were to map the normal LDR white to 100 nit, one would typically map the 1000% white to 1000 nit).
  • It may be advantageous for technical systems to not encode the brightness in its native manner, i.e. e.g. a float number, or some N bit representation which varies linearly with the brightness values, e.g. if 255 codes 100 nit, then 128 codes 50 nit instead of e.g. 25 nit. The digital coding of the brightness, involves a technical quantity called luma (Y). We will give luminances the letter L, and (relative) brightnesses the letter B. Note that technically, e.g. for ease of definition of some operations, one can always normalize even 5000 nit WPDPL range luminances to the normalized range [0,1], but that doesn't detract from the fact that these normalized luminances still represent absolute luminances on a range ending at 5000 nit (in contrast to relative brightnesses that never had any clear associated absolute luminance value, and can only be converted to luminances ad hoc, typically with some arbitrary value).
  • For SDR signals the luma coding used a so-called Opto-electronic Transfer Function (OETF), between the optical brightnesses and the electronic typically 8 bit luma codes which (approximately) by definition was: Y _ float = sqrt B _ relative .
  • If B_relative is a float number ranging from 0 to 1.0, so will Y_float.
  • Subsequently that signal value Y_float is quantized, because we want 8 bit digital representations, ergo, the Y_dig value that is communicated to receivers over e.g. airways DVB (or internet-supplied video on demand, or blu-ray disk, etc.) has a value between 0 and 255 (i.e. power(2;8)-1).
  • One can show the gamut of all SDR colors (or similarly one can show gamuts of HDR colors, which would if using the same RGB primaries defining the chromatic gamut have the same base, but after renormalization stretch vertically to a larger absolute gamut solid) as in Fig. 1B . Larger Cb and Cr values will lead to a (psychovisually more relevant color characterizer) larger saturation (sat), which moves outwards from the unsaturated or colorless colors vertical axis showing the achromatic colors in the middle, towards the maximally saturated colors on the circle (which is a transformation of the usual color triangle spanned by the RGB primaries as vertices). Hues h (i.e. the color category, yellows, versus greens, versus blues) will be angles along the circle. The vertical axis represents normalized luminances (in a linear gamut representation), or normalized lumas (in a non-linear representation (coding) of those luminances, e.g via a psychovisually uniformized OETF or its inverse the EOTF). Since after normalization (i.e. division by the respective MDWPL values, e.g. 2000 for an HDR image of a particular video and 100 for an SDR image), the common representation will become easy, and one can define luminance (or luma) mapping functions on normalized axes as shown in Fig. 1D (i.e. which function F_comp re-distributes the various image object pixel values as needed, so that in the actual luminance representation e.g. a dark object looks the same, i.e. has the same luminance, but a different normalized luminance since in one situation a pixel normalized luminance will get multiplied by 100 and in the other situation by 2000, so to get the same end luminance the latter pixel should have a normalized luminance of 1/20th of the former). When down-grading to a smaller range of luminances, one will typically get (in the normalized to 1.0 representation, i.e. mapping a range of normalized input luminances L_in between zero and one to normalized output luminances L_out) a convex function which everywhere lies above the diagonal diag, but the exact shape of that function F_comp, e.g. how fact it has to rise at the blacks, will depend typically not only on the two MDWPL values, but (to have the most perfect re-grading technology version) also on the scene contents of various (video or still) images, e.g. whether the scene is a dark cave, and there is action happening in a shadowy area, which must still be reasonably visible even on 100 nit luminance ranges (hence the strong boost of the blacks for such a scenario, compared to a daylight scene which may employ a near linear function almost overlapping with the diagonal).
  • In the representation of Fig. 1B one can show the mapping from a HDR color (C_H) to an LDR color (C_L) of a pixel as a vertical shift (assuming that both colors should have the same proper color, i.e. hue and saturation, which usually is the desired technical requirement, i.e. on the circular ground plane they will project to the same point). Ye means the color yellow, and its complementary color on the opposite side of the achromatic axis of luminances (or lumas) is blue (B), and W signifies white (the brightest color in the gamut a.k.a. the white point of the gamut, with the darkest colors, the blacks being at the bottom).
  • So the receiving side (in the old days, or today) will know it has an SDR video, if it gets this format. The maximum white (of SDR) will be by definition the brightest color that SDR can define. So if one now wants to make brighter image colors (e.g of real luminous lamps), that should be done with a different codec (as one can show the math of the Rec. 709 OETF allows only a coding of up to 1000:1 and no more).
  • So one defined new frameworks with different code allocation functions (EOTFs, or OETFs). What is of interest here is primarily the definition of the luma codes.
  • For reasons beyond what is needed for the present discussion, most HDR codecs start by defining an Electro-optical transfer function instead of its inverse, the OETF. Then one can at least basically define brighter (and darker) colors. That as such is not enough for a professional HDR coding system, since because it is different from SDR, and there are even various flavors, one wants more (new compared to SDR coding) technical information relating to the HDR images, which will be metadata.
  • The property of those HDR EOTFs is that they are much steeper, to encode a much larger range of needed to be coded HDR luminances, and a significant part of that range coding specifically darker colors (relatively darker, since although one may be coding absolute luminances with e.g. the Perceptual Quantizer (PQ) EOTF (standardized in SMPTE 2084), one applies the function after normalization). In fact if one were to use exact power functions as EOTFs for coding HDR luminances as HDR lumas, one would have a power of 4, or even 7. When a receiver gets a video image signal defined by such an EOTF (e.g. Perceptual Quantizer) it will know it gets a HDR video. It will need the EOTF to be able to decode the pixel lumas in the plane of lumas spanning the image (i.e. having a width of e.g. 4000 pixels and a height of 2000), which will simply be binary numbers. Typically HDR images will also have a larger word length, e.g. 10 bit. However, one should not confuse the non-linear coding one can at will design by optimizing a non-linear EOTF shape with linear codings and the amount of bits needed for them. If one needs to drive, with a linear (bit-represented) code, e.g. a DMD pixel, indeed to reach e.g. 10000: 1 modulation darkest to brightest, one needs to take the log2 to obtain the number of bits. There one would need to have at least 14 bits (which may for technical reasons get rounded upwards to 16 bits), since power(2;14)= 16384 > 10000. But being able to smartly design the shape of the EOTF, and knowing that the visual system sees not all luminance differences equally, the present applicant has shown that (surprisingly) quite reasonable HDR television signals can be communicated with only 8 bit per pixel color component (of course if technically achievable in a system, 10 bits may be better and more preferable). So the receiving side may in both situations get as input a coded pixel color (luma and Cb, Cr; or in some systems by matrixing equivalent non-linear R'G'B' component values) which lie between 0 and 255, or 0 and 1023, but it will know the kind of signal it is getting (hence what ought to be displayed) from the metadata, such as the metadata (e.g. MPEG VUI metadata) co-communicated EOTF (e.g. a value 16 meaning PQ; 18 means another OETF was used to create the lumas, namely the Hybrid LogGamma OETF, ergo the inverse of that function should be used to decode the luma plane), in many HDR codings the MDWPL value (e.g. 2000 nit), and in more advanced HDR codings further metadata (some may e.g. co-encode luminance -or luma- mapping functions to apply for mapping image luminances from a primary luminance dynamic range to a secondary luminance dynamic range, such as one function FL_enc per image).
  • We detail the typical needs of an already more sophisticated HDR image handling chain with the aid of simple elucidation Fig. 1 (for a typical nice HDR scene image, of a monster being fought in a cave with a flame thrower, the master grading (Mstr_HDR) of which is shown spatially in Fig. 1A, and the range of occurring pixel luminances on the left of Fig. 1C). The master grading or master graded image is where the image creator can make his image look as impressive (e.g. realistic) as desired. E.g., in a Christmas movie he can make a baker's shop window look somewhat illuminated by making the yellow walls somewhat brighter than paper white, e.g. 150 nit (and real colorful yellow instead of pale yellow), and the light bulbs can be made 900 nit (which will give a really lit Christmas-like look to the image, instead of a dull one in which all lights are clipped white, and not much more bright than the rest of the image objects, such as the green of the Christmas tree).
  • So the basic thing one must be able to do is encode (and typically also decode and display) brighter image objects than in a typical SDR image.
  • Let's look at it colorimetrically now. SDR (and its coding and signaling) was designed to be able to communicate any Lambertian reflecting color (i.e. a typical object, like your blue jeans pants, which absorbs some of the infalling light, e.g. the red and green wavelengths, to emit only blue light to the viewer or capturing camera) under good uniform lighting (of the scene where the action is camera-captured). Just like we would do on a painting: if we don't add paint we get the full brightness reflecting back from the white painting canvas, and if we add a thick layer of strongly absorbing paint we will see a black stroke or dot. We can represent all colors brighter than blackest black and darker than white in a so-called color gamut of representable colors, as in Fig. 1B (the "tent"). As a bottom plane, we have a circle of all representable chromaticities (note that one can have long discussions that in a typical RGB system this should be a triangle, but those details are beyond the needs of the present teachings). A chromaticity is composed of a certain (rotation angle) hue h (e.g. bluish-green e.g. "teal"), and a saturation sat, which is the amount of pure color mixed in a grey, e.g. the distance from the vertical axis in the middle which represents all achromatic colors from black at the bottom becoming increasingly bright till we arrive at white. Chromatic colors, e.g. a half-saturated purple, can also have a brightness, the same color being somewhat darker or brighter. However, the brightest color in an additive system can only be (colorless) white, since it is made by setting all color channels to maximum R=G=B=255, ergo, there is no unbalance which would make the color clearly red (there is still a little bit of bluishness respectively yellowishness in the elected white point chromaticity, but that is also an unnecessary further discussion, we will assume D65 daylight white). We can define those SDR colors by setting MDWPL a (relative) 100% for white W (n.b., in SDR white does not actually have a luminance, since legacy SDR does not have a luminance associated with the image, but we can pretend it to be X nit, e.g. typically 100 nit, which is good average representative value of the various legacy SDR tv's).
  • Now we want to represent brighter than Lambertian colors, e.g. the self-luminous flame object (flm) of the flame thrower of the soldier (sol) fighting the monster (mon) in this dark cave.
  • Let's say we define a (video maximum luminance ML_V) 5000 nit master HDR grading (master means the starting image -most important in this case best quality grading- which we will optimally grade first, to define the look of this HDR scene image, and from which we can derive secondary gradings a.k.a. graded images as needed). We will for simplicity talk about what happens to (universal) luminances, then we can for now leave the debate about the corresponding luma codes out of the discussion, and indeed PQ can code between 1/10,000 nit and 10,000 nit, so there is no problem communicating those graded pixel luminances as a e.g. 10 bit per component YCbCr pixelized HDR image, if coding according to that PQ EOTF (of course, the mappings can also be represented, and e.g. implemented in the processing IC units, as an equivalent luma mapping).
  • The two dotted horizontal lines represent the limitations of the SDR codable image, when associating 100 nit with the 100% of SDR white.
  • Although in a cave, the monster will be strongly illuminated by the light of the flames, so we will give it an (average) luminance of 300 nit (with some spread, due to the square power law of light dimming, skin texture, etc.).
  • The soldier may be 20 nit, since that is a nicely slightly dark value, still giving some good basic visibility.
  • A vehicle (veh) may be hidden in some shadowy corner, and therefore in a archetypical good impact HDR scene of a cave e.g. have a luminance of 0.01 nit. The flames one may want to make impressively bright (though not too exaggerated). On an available 5000 nit HDR range, we could elect 2500 nit, around which we could still gradually make some darker and brighter parts, but all nicely colorful (yellow and maybe some oranges).
  • What would now happen in a typical SDR representation, e.g. a straight from camera SDR image capturing?
  • The camera operator would open his iris so that the soldier comes out at "20 nit", or in fact more precisely 20%. Since the flames are much brighter (note: we didn't actually show the real world scene luminances, since master HDR video Mstr_HDR is already an optimal grading to have best impact in a typical living room viewing scenario, but also in the real world the flames would be quite brighter than the soldier, and certainly the vehicle), they would all clip to maximum white. So we would see a bright area, without any details, and also not yellow, since yellow must have a lower luminance (of course the cinematographer may optimize things so that there is still somewhat of a flame visible even in LDR, but then that is firstly never as impactful extra bright, and secondly at the detriment of the other objects which must become darker).
  • The same would also happen if we built a SDR (max. 100 nit) TV which would map equi-luminance, i.e. it would accurately represent all luminances of the master HDR grading it can represent, but clip all brighter object to 100 nit white.
  • So the usual paradigm in the LDR era was to relatively map, i.e. the brightest brightness (here luminance) of the received image to the maximum capability of the display. So as this maps 5000 nit by division by 50 on 100 nit, the flames would still be okay since the are spread as yellows and oranges around 50 nit (which is a brightness representable for a yellow, since as we see in Fig. 1B the gamut tent for yellows goes down in luminance only a little bit when going towards the most saturated yellows, in contrast to blues (B) on the other side of the slice for this hue angle B-Ye, which blues can only be made in relatively dark versions). However this would be at the detriment of everything else becoming quite dark, e.g. the soldier 20/50 nit which is pure black (and this is typically a problem that we see in SDR renderings of such kinds of movie scene).
  • So, if having established a good HDR maximum luminance (i.e. ML_V) for the master grading, and a good EOTF e.g. PQ for coding it, we can in principle start communicating HDR images to receivers, e.g. consumer television displays, computers, cinema projectors, etc.
  • But that is only the most basic system of HDR.
  • The problem is that, unless the receiving side has a display which can display pixels at least as bright as 5000 nit, there is still a question of how to display those pixels.
  • Some (DR adaptation) luminance down-mapping must be performed in the TV, to make darker pixels which are displayable. E.g. if the display has a (end-user) display maximum luminance ML_D of 1500 nit, one could somehow try to calculate 1200 nit yellow pixels for the flame (potentially with errors, like some discoloration, e.g. changing the oranges into yellows).
  • This luminance down-mapping is not really an easy task, especially to do very accurately instead of sufficiently well, and therefore various technologies have been invented (also for the not necessarily similar task of luminance up-mapping, to create an output image of larger dynamic range and in particular maximum luminance than the input image).
  • Typically one wants a mapping function (generically, i.e. used for simplicity of elucidation) of a convex shape in a normalized luminance (or brightness) plot, as shown in Fig. 1D. Both input and output luminances are defined here on a range normalized to a maximum equaling one, but one must mind that on the input axis this one corresponds to e.g. 5000 nit, and on the output axis e.g. 200 nit (which to and for can be easily implemented by division respectfully multiplication). In such a normalized representation the darkest colors will typically be too dark for the grading with the lower dynamic range of the two images (here for down-conversion shown on the vertical output axis, of normalized output luminances L_out, the horizontal axis showing all possible normalized input luminances L_in). Ergo, to have a satisfactory output image corresponding to the input image, we must relatively boost those darkest luminances, e.g. by multiplying by 3x, which is the slope of this luminance compression function F_comp for its darkest end. But one cannot boost forever if one wants no colors to be clipped to maximum output, ergo, the curve must get an increasingly lower slope for brighter input luminances, e.g. it may typically map input 1.0 to output 1.0. In any case the luminance compression function F_comp for down-grading will lie above the 45 degree diagonal (diag) typically.
  • Care must still be taken to do this correctly. E.g., some people like to apply three such compressive functions to the three red, green and blue color channels separately. Whilst this is a nice and easy guarantee that all colors will fit in the output gamut (an RGB cube, which in chromaticity-luminance (L) view becomes the tent of Fig. 1B) especially with higher non-linearities it can lead to significant color errors. A e.g. reddish orange hue is determined by the percentage of red and green, e.g. 30% green and 70% red. If the 30% now gets doubled by the mapping function, but the red stays in the feeble-sloped part of the mapping function almost unchanged, we will have a 60/70, i.e. 50/50 i.e. a yellow instead of an orange. This can be particularly annoying if it depends on (in contrast to the SDR paradigm) non-uniform scene lighting, e.g. an sports car entering the shadows suddenly turning yellow.
  • Ergo, whilst the general desired shape for the brightening of the colors may still be the function F_comp (e.g. determined by the video creator, when grading a secondary image corresponding to his master HDR image already optimally graded), one wants a more savvy down-mapping. As shown in Fig. 1B, for many scenarios one may desire a re-grading which merely changes the brightness of the normalized luminance component (L), but now the innate type of color, i.e. its chromaticity (hue and saturation). If both SDR and HDR are represented with the same red, green and blue color primaries, they will have a similarly shaped gamut tent, only one being higher than the other in absolute luminance representation. If one scales both gamuts with their respective MDWPL values (e.g. MDWPL1= 100 nit, and MDWPL2= 5000 nit), both gamuts will exactly overlap. The desired mapping from a HDR color C_H to a corresponding output SDR color CL (or vice versa) will simply be a vertical shifting, whilst the projection to the chromaticity plane circle stays the same.
  • Although the details of such approaches are also beyond the need of the present application, we have thought examples of such color mapping mechanism before, where the three color components are processed coordinately, although in a separate luminance and chroma processing path, e.g. in WO2017157977 .
  • If it is now possible to down-grade with one (or more) luminance mapping functions (the shape of which may be optimized by the creator of the video(s)), in case one uses invertible functions one can design a more advanced HDR codec.
  • Instead of just making some final secondary grading from the master image, e.g. in a television, one can make a lower dynamic range image version for communication, communication image Im_comm. We have elected in the example this image to be defined with its communication image maximum luminance ML_C equal to 200 nit. The original 5000 nit image can then be reconstructed (a.k.a. decoded) as a reconstructed image Rec_HDR (i.e. with the same reconstructed image maximum luminance ML_REC) by receivers, if they receive in metadata the decoding luminance mapping function FL_dec, which is typically substantially the inverse of the coding luminance mapping function FL_enc, which was used by the encoder to map all pixel luminances of the master HDR image into corresponding lower pixel luminances of the communication image Im_comm. So the proxy image for communicating actually an image a higher dynamic range (DR_H, e.g. spanning from 0.001 nit to 5000 nit) is an image of a different, lower dynamic range (DR_L).
  • Interestingly, one can even elect the communication image to be a 100 nit LDR (i.e. SDR) image, which is immediately ready (without further color processing) to be displayed on legacy LDR images (which is a great advantage, because legacy displays don't have HDR knowledge on board). How does that work? The legacy TV doesn't recognize the MDWPL metadatum (cos that didn't exist in the SDR video standard, so the TV is also not arranged to go look for it somewhere in the signal, e.g. in a Supplemental Enhancement Information message, which is MPEG's mechanism to introduce all kinds of pre-agreed new technical information). It is also not going to look for the function. It just looks at the YCbCr e.g. 1920x1080 pixel color array, and displays those colors as usual, i.e. according to the SDR Rec. 709 interpretation. And the creator has chosen in this particular codec embodiment his FL_enc function so that all colors, even the flame, map to reasonable colors on the limited SDR range. Note that, in contrast to a simple multiplicative change corresponding to the opening or shutting of a camera iris in an SDR production (which typically leads to clipping to at least one of white and/or black), now a very complicated optimal function shape can be elected, as long as it is invertible (e.g. we have taught systems with first a coarse pre-grading and then a fine-grading). E.g. one can move the luminance (respectively relative brightness) of the car to a level which is just barely visible in SDR, e.g. 1% deep black, whilst moving the flame to e.g. 90% (as long as everything stays invertible). That may seem extremely daunting if not impossible at first sight, but many field tests with all kinds of video material and usage scenarios have shown that it is possible in practice, as long as one does it correctly (following e.g. the principles of WO2017157977 ).
  • How do we now know that this is actually a HDR video signal, even if it contains an LDR-usable pixel color image, or in fact that any HDR-capable receiver can reconstruct it to HDR: because there are also the functions FL_dec in metadata, typically one per image. And hence that signal codes what is also colorimetrically, i.e. according to our above discussion and definition, a (5000 nit) HDR image.
  • Although already more complex than the basic system which communicates only a PQ-HDR image, this per SDR proxy coding is still not the best future-proof system, as it still leaves the receiving side to guess how to down-map the colors if it has e.g. a 1500 nit, or even a 550 nit, tv.
  • Therefore we added a further technical insight, and developed so-called display tuning technology (a.k.a. display adaptation): the image can be tuned for any possible connected tv, i.e. any ML_D, because one can double the function of the coding function FL_enc as some guidance function for the up-mapping from 100 nit Im_comm not to a 5000 nit reconstructed image, but to e.g. a 1500 nit image. The concave function, which is substantially the inverse of F_comp (note, for display tuning there is no requirement of exact inversion as there is for reconstruction), will now have to be scaled to be somewhat less steep (i.e. from the reference decoding function FL_dec a display adapted luminance mapping function FL_DA will be calculated), since we expand to only 1500 nit instead of 5000 nit. I.e. an image of tertiary dynamic range (DR_T) can be calculated, e.g. optimized for a particular display in that the maximum luminance of that tertiary dynamic range is typically the same as the maximum displayable luminance of a particular display.
  • Techniques for this are described in WO2017108906 (we can transform a function of any shape into a similarly-shaped function which lies closer to the 45 degree diagonal, by an amount which depends on the ratio between the maximum luminances of the input image and the desired output image, versus the ratio of the maximum luminances of the input image and a reference image which would here be the reconstructed image, by e.g. using that ratio to obtain closer points on for all points on the diagonal orthogonally projecting a line segment from the respective diagonal point till it meets a point on the input function, which closer points together define the tuned output function, for calculating the to be displayed image Im_disp luminances from the Im_comm luminances).
  • Not only did we get more kinds of displays even for basic movie or television video content (LCD tv, mobile phone, home cinema projector, professional movie theatre digital projector), and more video different sources and communication media (satellite, streaming over the internet, e.g. OTT, streaming over 5G), but also did we get more production manners of video.
  • Fig. 2 shows -in general, without desiring to be limiting- a few typical creations of video where the present teachings may be usefully deployed.
  • In a studio environment, e.g. for the news or a comedy, there may still be a tightly controlled shooting environment (although HDR allows relaxation of this, and shooting in real environments). There will be controlled lighting (202), e.g. a battery of base lights on the ceiling, and various spot lights. There will be a number of bulky relatively stationary television cameras (201). Variations on this often real-time broadcast will be e.g. a sports show like soccer, which will have various types of cameras like near-the-goal cameras for a local view, overview cameras, drones, etc.
  • There will be some production environment 203, in which the various feeds from the cameras can be selected to become the final feed, and various (typically simple, but potentially more complex) grading decisions can be taken. In the past this often happened in e.g. a production truck, which had many displays and various operators, but with internet-based workflows, where the raw feeds can travel via some network, the final composition may happen at the premises of the broadcaster. Finally, when simplifying the production for this elucidation, some coding and formatting for broadcast distribution to end (or intermediate, such as local cable stations) customers will happen in formatter 204. This will typically do the conversion to e.g. PQ YCbCr from the luminances as graded as explained with Fig. 1, for e.g. an intermediate dynamic range format, calculate and format all the needed metadata, convert to some broadcasting format like DVB or ATSC, packetize in chunks for distribution, etc. (the etc. indicating there may be tables added for signaling available content, sub-titling, encryption, but at least some of that will be of lesser interest to understand the details of the present technical innovations).
  • In the example the video (a television broadcast in the example) is communicated via a television satellite 250 to a satellite dish 260 and a satellite signal capable set-top-box 261. Finally it will be displayed on an end-user display 263.
  • This display may be showing this first video, but it may also show other video feed, potentially even at the same time, e.g. in Picture-in-Picture windows (or some data of the first HDR video program may come via some distribution mechanism and other data via another).
  • A second production is typically an off-line production. We can think of a Hollywood movie, but it can also be a show of somebody having a race through a jungle. Such a production may be shot with other optimal cameras, e.g. steadicam 211 and drone 210. We again assume that the camera feeds (which may be raw, or already converted to some HDR production format like HLG) are stored somewhere on network 212, for later processing. In such a production we may in the last months of production have some human color grader use grading equipment 213 to determine the optimal luminances (or relative brightnesses in case of HLG production and coding) of the master grading. Then the video may be uploaded to some internet-based video service 251. For professional video distribution this may be e.g. Netflix.
  • A third example is consumer video production. Here the user will have e.g. when making a vlog a ring lighter 221, and will capture via a mobile phone 220, but (s)he may also be capturing in some exterior location without supplementary lighting. She/he will typically also upload to the internet, but now maybe to YouTube, or TikTok, etc.
  • In case of reception via the internet, the display 263 will be connected via a modem, or router 262 or the like (more complicated setups like in-house Wi-Fi and the like are not shown in this mere elucidation).
  • Another user may be viewing the video content on a portable display (271), such as a laptop (or similarly other users may use a non-portable desktop PC), or a mobile phone etc. The may access the content over a wireless connection (270), such as Wi-Fi, 5G, etc.
  • So it can be seen that today, various kinds of video, in various technical codings, can be generated and communicated in various manners, and our coding and processing systems have been designed to handle substantially all those variants.
  • Fig. 3 shows an example of an absolute (nit-level-defined) dynamic range conversion circuit 300 for a (HDR) image or video decoder shown in a video processing circuit chain in Fig. 3C. (The encoder would typically work similarly but with inverted functions typically, i.e. the function to be applied being the function of the other side mirrored over the diagonal). It is based on coding a primary image (e.g. a master HDR grading) with a primary luminance dynamic range (DR_Prim) as another (so-called proxy) image with a different secondary range of pixel luminances (DR_Sec). If the encoder and all its supply-able decoders have pre-agreed or know that the proxy image has a maximum luminance of 100 nit, this need not be communicated as an SDR_WPL metadatum. If the proxy image is e.g. a 200 nit maximum image, this will be indicated by filling its proxy white point luminance P_WPL with the value 200, or similarly for 80 nit etc. The maximum of the primary image (HDR_WPL= 1000), to be reconstructed by the dynamic range conversion circuit, will normally be co-communicated as metadata of the received input image, or video signal, i.e. together with the input pixel color triplets (Y_in, Cb_in, Cr_in). The various pixel lumas will typically come in as a luma image plane, i.e. the sequential pixels will have first luma Y11, second Y21, etc. (typically these will be scanned, and the dynamic range conversion circuit will convert pixel by pixel to output pixel color triplets (Y_out, Cb_out, Cr_out). We will primarily focus on the brightness dimension of the pixel colors in this elucidation. Various dynamic range conversion circuits may internally work differently, to achieve basically the same thing: a correctly reconstructed output luminance L_out for all image pixels (the actual details don't matter for this innovation, and the embodiments will focus on teaching only aspects as far as needed).
  • The mapping of luminances from the secondary dynamic range to the primary dynamic range may happen on the luminances themselves, but also on any luma representation (i.e. according to any EOTF, or OETF), provided it is done correctly, e.g. not separately on non-linear R'G'B' components. The internal luma representation need not even be the one of the input (i.e. of Y_in), or for that manner of whatever output the dynamic range conversion circuitry or its encompassing decoder may deliver (e.g. a format luma Y_sigfm for a particular communication format or communication system, "communicating" including storage to a memory, e.g. inside a PC, a hard disk, an optical storage medium, etc.).
  • We have optionally (dotted) shown a luma conversion circuit 301, which turns the input lumas Y_in into perceptionally uniformized lumas Y_pc.
  • Applicant standardized in ETSI 103433 a useful equation to convert luminances in any range to such a perceptual luma representation: Y _ pc = log _ 10 1 + RHO WPL _ inrep 1 power Ln _ in ; 1 / 2.4 / log _ 10 RHO WPL _ inrep
  • In which the function RHO is defined as RHO WPL _ inrep = 1.32 power WPL _ inrep / 10000 ; 1 / 2.4
  • The value WPL_inrep is the maximum luminance of the range that needs to be converted to psychovisually uniformized lumas, so for the 100 nit SDR image this value would be 100, and for the to be reconstructed output image (or the originally coded image at the creation side) the value would be 1000.
  • Ln_in are the luminances along that whichever range which need to be converted, after normalization by dividing by its respective maximum luminance, i.e. within range [0,1],
  • Once we have an input and an output range normalized to 1.0, we can apply a luminance mapping function actually in the luma domain, as shown inside the luma mapping circuit 302, which does the actual luma mapping for each incoming pixel.
  • In fact, this mapping function had been specifically chosen by the encoder of the image (at least for yielding good quality reconstructability, and maybe also a reduced amount of needed bits when MPEG compressing, but sometimes also fulfilling further criteria like e.g. the SDR proxy image being of correct luminance distribution for the particular scene -a dark cave, or a daytime explosion- on a legacy SDR display, etc.). So this function F_dec (or its inverse) will be extracted from metadata of the input image signal or representation, and supplied to the dynamic range conversion circuit for doing the actual per pixel luma mapping. In this example the function F_dec directly specifies the needed mapping in the perceptual luma domain, but other variants are of course possible, as the various conversions can also be applied on the functions. Furthermore, although for simplicity of explanation, and to guarantee the teaching is understood, we teach here a pure decoder dynamic range conversion, but other dynamic range conversions may use other functions, e.g. a function derived from F_dec, etc. The details of all of that are not needed for understanding the present innovative contribution to the technology.
  • In general one will not only change the luminances, but there will be a corresponding change in the chromas Cb and Cr. That can also be done in various manners, from strictly inversely decoding, to implementing additional features like a saturation boost, since Cb and Cr code the saturation of the pixels. Thereto another function is typically communicated in metadata (recoloring specification function FCOL), which determines the chromatic recoloring behavior, i.e. the mapping of Cb and Cr (note that Cb and Cr will typically be changed by the same multiplicative amount, since the ratio of Cr/Cb determines the hue, and generally one does not want to have hue changes when decoding, i.e. the lower and higher dynamic range image will in general have object pixels of different brightness, and oftentimes at least some of the pixels will have different saturation, but ideally the hue of the pixels in both image versions will be the same). This color function will typically specify a multiplier which has a value dependent on a brightness code Y (e.g. the Y_pc, or other codes in other variants). A multiplier establishment circuit 305 will yield the correct multiplier m for the brightness situation of the pixel being processed. A multiplier 306 will multiply both Cb_in and Cr_in by this same multiplier, to obtain the corresponding output chromas Cb_out= m*Cb_in and Cr_out=m*Cr_in. So the multiplier realizes the correct chroma processing, therefore the whole color processing of any dynamic range conversion being correctly configurable in the dynamic range conversion circuit.
  • Furthermore, there may typically be (at least in a decoder) a formatting circuit 310, so that the output color triplet (Y_out, Cb_out, Cr_out) can be converted to whatever needed output format (e.g. an RGB format, or a communication YCbCr format, Y_sigfm, Cb_sigfm, Cr_sigfm). E.g. if the circuit outputs to a version of a communication channel 379 which is an HDMI cable, such cables typically use PQ-based YCbCr pixel color coding, ergo, the lumas will again be converted from the perceptual domain to the PQ domain by the formatting circuit.
  • It is important that the reader well understands what is a (de)coding, and how an absolute HDR image, or its pixel colors, is different from a legacy SDR image. There may be a connection to a display tuning circuit 380, which calculates ultimate pixel colors and luminances to be displayed at the screen of some display, e.g. a 450 nit tv which some consumer has at home.
  • However, in absolute HDR, one can establish pixel luminances already in the decoding step, at least for the output image (here the 1000 nit image).
  • We have shown this in Fig. 3B, for some typical HDR image being an indoors/outdoors image (the geometry and comprised image objects of which are shown in Fig. 3A). Note that, whereas in the real world the outdoors luminances may typically be 100 times brighter than the indoors luminances, in an actual master graded HDR image it may be better to make them e.g. 10x brighter, since the viewer will be viewing all together on a screen, in a fixed viewing angle, even typically in a dimly illuminated room in the evening, and not in the real world.
  • Nevertheless, we find that when we look at the luminances corresponding to the lumas, e.g. the HDR luminances L_out, we typically see a large histogram (of counts N(L_out) of each occurring luminance in an output image of this homely scene). This spans considerably above some lower dynamic range lobe of luminances, and above the low dynamic range 100 nit level, because the sunny outdoors images have their own histogram lobe. Note that the luminance representation is drawn non-linearly, e.g. logarithmically. We can also trace what the encoder would do at the encoding side, when making the 100 nit proxy image (and its histogram of proxy luminance counts N(L_in)). A convex function, as shown in Fig. 1, or inside luma mapper 302, is used which squeezes in the brighter luminances, due to the limitations of the smaller luminance dynamic range. There is still some difference between the brightness of indoors and outdoors, and still a considerable range for the upper lobe of the outdoors objects, so that one can still make the different colors needs to color the various objects, such as the various greens in the tree. However, there must also be some sacrifices. Firstly the indoors objects will display (assuming for the moment an 100 or 200 nit display would faithfully display those luminances as coded, and not e.g. do some arbitrary beautification processing which brightens them) darker, darker than ideally desired, i.e. up to the indoors threshold luminance T_in of the HDR image. Secondly, the span of the upper lobe is squeezed, which may give the outdoors objects less contrast. Thirdly, since bright colors in the tent-shaped gamut as shown in Fig. 1 cannot have large saturation, the outdoors colors may also be somewhat pastellized, i.e. of lowered saturation. But of course if the grader at the creation side has control over all the functions (F_enc, FCOL), hey may balance those features, so that some have a lesser deviation at the detriment of others. E.g. if the outdoors shows a plain blue sky, the grader may opt for making it brighter, yet less blue. If there was a beautiful sunset, he may want to retain all its colors, and make everything dimmer instead, in particular if there are no important dark corners in the indoors part of the image, which would then have their contents badly visible, especially when watching tv with all the lights on (note that there are also techniques for handling illumination differences and the visibility of the blacks, but that is too much information for this patent application's elucidation).
  • The middle graph shows what the lumas would look like for the proxy luminances, and that may typically give a more uniform histogram, with e.g. approximately the same span for the indoors and outdoors image object luminances. The lumas are however only relevant to the extent of coding the luminances, or in case some calculations are actually performed in the luma domain (which has advantages for the size of the word length on the processing circuitry). Note that whereas the absolute formalism can allocate luminances on the input side between zero and 100 nit, one can also treat the SDR luminances as relative brightnesses (which is what a legacy display would do, when discarding all the HDR knowledge and communicated metadata, and looking merely at the 0-255 luma and chroma codes).
  • Fig. 4 elucidates how the computer world looked at HDR, in particular for the calculation of HDR scenes, e.g. by ray-tracing (i.e. the equivalent of actual camera capturing). This technology, which can be used in e.g. gaming, in which a different view on an environment has to be re-calculated each time a player moves, potentially with a specular reflection appearing that wasn't in view when the player was positioned one game meter to the left, was not primarily geared for communication, e.g. from a broadcaster to a receiver, any receiver, with any of various kinds of displays with different maximum brightness characteristics. These calculations and representations originated for a typical application (tightly managed) inside a single PC, and were certainly not developed with a view on many future quite different applications. When one actually calculates a HDR image, by defining e.g. a very bright light source and (internally) physically modeling how it illuminates a pixel of a mirror, one can define just any HDR luminance. But one of the problems is already one can calculate it, but not necessarily display it. E.g., one would typically use a floating point representation (e.g. 16 bit half-float for representing the luminance, in a logarithmic format, i.e. able to represent almost infinite luminances, basically ridiculously large from the point of view of actually using for display). One would then consider all luminances lower than 1 (which can also be stated as the relative 100%) as the normal SDR luminances. It may not be the best way to treat, e.g. display, an HDR float image, but some systems or components, or software would indeed clip everything above 1 to 1.0, and just use the lower luminances. The outdoors objects, like the house, when being computationally generated, may have luminances a multiplication factor higher than 1.0, e.g. 5x or 20x (if one were to associate 1.0 with 100 nit, which was not necessarily done in the computer view, one could say these would be e.g. 500 nit). One could make the sun e.g. as bright as 1 million nit, which would need severe down-mapping to make it visible as a light yellow sphere, or, it would typically clip to white (which would not be an issue, since also in the real world the sun is so bright that it will clip to white in camera capturings). However, there may be issues with other objects in the down-mapping.
  • One difference one already sees with e.g. the PQ representation is that this representation is closed (one assumes every luminance one is ever going to need in practice may reasonably fall below its maximum being 10,000 nit), and therefore one can do e.g. luminance mappings based on this end-point, whereas the computer log format is -pragmatically- open, in the sense that it can go almost infinitely above 1. This may involve some (undesirable) clipping. One could say that even for a 16 bit logarithmic format there is some end-point, but given that luminance or brightness will be extremely high, that is much less interesting to map, than mapping the values around 1.0 (compressing e.g. linearly from 1 million nit would make all indoors pixels pitch black). So open-ended representations need a somewhat different luminance mapping (called tone mapping in this sub-area), but on a more general level one can come to some common ground approach, at least in the sense that both sub-fields of technology have been able to create LDR images (and similarly, in the modern area, one can also create (common) HDR images for both).
  • Computer games, which were until recently played on SDR displays mostly anyway, despite making beautiful lifelike images, needed to down-map those to yield SDR images. One could argue that would not be so different from the down-mapping of natural, camera-captured HDR images, which was not typical until recently anyway. But the tone mappers were sometimes difficult to fathom, and somewhat ad hoc. Nevertheless, HDR gaming did become possible, even with some introductory pains, and people loved the look. But having things sorted out for, like HDR television, another tightly controlled application, doesn't mean yet that all HDR problems are solved in the computer world.
  • Television versus computing environment
  • Although there are still many flavors (and technical visions, e.g. the absolute PQ-based coding versus the relative HLG-based coding), as shown above at least for the simple video communication (e.g. for broadcast television services) the situation is relatively simple and akin to the "direct display control line" situation of the analog PAL era (which used the display paradigm: the brightest code in the image -white- gets displayed as the brightest thing on the display, and perceived by the viewer as the whitest white). It is generalized beyond the exact direct display (of square root mapping followed by inverse square power mapping of SDR systems), as there may be e.g. complex display adaptation involved, but it is still similar in the sense that one video takes total area of view of a display, and will in general be optimized for one display of an end-consumer, to be thereafter directly displayed as sole asset on the end-user display.
  • Computer systems may be more complex, as there may be various unrelated visual assets, and the viewer may also be using different displays (e.g. an old SDR one, and a new one, and show at least some of the content on one of the displays, but he may also change the display for that content by dragging e.g. its window to the other display). Furthermore, whereas a video processor/processing is usually provided by one manufacturer, and usually follows well-standardized principles, a computer may be running applications from just about anybody, and those manufacturers (e.g. of software and its look and feel) may have different visions about higher luminance representation and/or display in more than just minor details.
  • Also, if this issue would have been handled in the previous century, one could still argue that the physics of the at the time universal Cathode Ray Tube monitors would put some limit on the wildless of higher brightness colors that various manufacturers/suppliers could be using, but we now also may want to cater for very dissimilar types of displays in very dissimilar viewing scenarios (e.g. one may want to display the same content on a small LCD-based mobile phone watched on the train, a projector in a darkened home cinema in a consumer's attic, LED panels in a supermarket, etc.). It seems that also the underlying circuits and software may sometimes be multiplying instead of converging: in addition to the classical operating systems Linux, Windows, Android, one may now have proprietary operating systems (e.g. Tizen) that, even when derived from a basic operating system may have some differential behavior regarding high brightness colors at least in one sense or condition.
  • A computer may also be seen as a "kit of parts", and that may be good for its general usability for a myriad of tasks, but that doesn't necessarily mean that these parts would work together in a stable well-defined manner.
  • The problems described herein, and the solutions offered by the embodiments will be similar just as well for native applications (which are specifically written for a specific computer platform, and run on that computer platform) and for the currently popular web applications, which are served from a remote server, i.e. have a number of services from that remote server, yet may rely for some computations (e.g. execute downloaded JavaScript) on the local client computer, and may need to do so for certain steps of the program. To the customer it does not seem to make much difference whether his spreadsheet is installed locally, or running on the cloud (except maybe for the subscription fee). He may not even realize that if he is typing an email in Gmail running in a browser, that under the hood he is actually making use of internet technologies, as he is writing html lines. The difference between a local file browser and an internet file browser is becoming less. But although the user may not care or want to see what exactly is going on to produce a well-working and visually pleasing result, having various parties work (and decide) on assets makes for big question who is in charge of the colorimetry. When any service running over the internet, those systems may want to make use of ever more advanced features, which may need to take recourse to details of the client's computing device, under the hood. E.g., a server-controlled internet game, may use a farm of processors to calculate the behavior of virtual actors, yet want to benefit from the hardware acceleration of calculations on the client's GPU, e.g. for doing the final shading (with shading in the computer sense we mean the calculations needed to come to the correct colors including brightnesses of an object, e.g. using a simulation of illumination of some object with some physical texture like a tapestry, or interpolation of colors of vertices of a triangle, which pixel colors typically end up as red, green and blue color components in a color buffer; the name rendering can also be used, but means the higher level whole process of coming to an image of an object from a viewpoint, for any generation of a NxM matrix of pixels in a so-called canvas, which is a memory in which to put finally rendered pixels, taking into account also such aspects as visibility and occlusion). Whereas we do not want to imply any limitations, we will elucidate some technical details of our innovations with some more challenging web application embodiments.
  • There are particular problems when several visual assets, of different brightness dynamic range, need to be displayed in a coordinated manner, such as e.g. in a window system on a computer (and possibly more challenging on several screens, of possibly different brightness capability). This is already a difficult technical problem per se, and the fact that in practice many different components hence producers/companies are involved (various software or middleware applications, different operating systems controlled by companies like Microsoft, Apple or Google, different Graphics Processing Unit vendors, etc.) does not necessarily make things easier. In this text with visual asset we do not necessarily mean a displayable thing in an area that has been prepared previously, and e.g. stored in a memory, but also visual assets that can be generated on the fly, e.g. a text with HDR text colors generated as it is being typed, into some window, and to end up in some canvas which collects all the visual assets in the various windows, to ultimately get displayed after traveling through the whole processing pipe which happens to be in place in the computer being configured with that particular set of applications presenting their visual assets.
  • To be clear, what we mean by graphics processing unit is a circuit (or maybe in some systems circuits plural, if there are two or more separate GPUs being supplied with the visual assets to be displayed in totality) which contain the final buffering for such pixels that should be shown on a display, i.e. be scanned out to at least one display (and not processors which do not have this scanout buffer and circuitry for communicating the video signals out to display(s); this GPU is sometimes also called Display Processing Unit DPU).
  • So there is a need for a universal well-coordinatable approach of handling HDR (and possibly some SDR) visual assets on computers, which may have several applications running in parallel, none of them controlling the entire visible screen, and which may work (in the sense of at least sending their preferred pixel colors to the GPU scanout buffer) through various layers of software, middleware, APIs etc. (i.e. which may communicate directly to an operating system, or via other processes, and typically via calls which can contain and communicate configurable data).
  • SUMMARY OF THE INVENTION
  • The indicated problems are handled by a computer (700) arranged to run an operating system (504), wherein the operating system is arranged to manage the color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Ass1) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness, wherein the visual asset is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness;
    • wherein the visual asset in addition comprises color transformation control data (CTctrl) which comprises at least one of first brightness mapping function (TM1) or its inverse and a reference white level (WL1);
    • wherein the operating system comprises a brightness mapping unit (710) arranged to receive the visual asset and to apply a tertiary brightness mapping function (FL_op) to the input luma code to obtain an output luma code, wherein the output luma code specifies a tertiary brightness which lies in a tertiary brightness range which is different from the secondary brightness range, wherein the output luma codes a brightness of an output color, wherein the tertiary brightness mapping function (FL_op) is based on at least one of, or both of, the first brightness mapping function (TM1) and the reference white level (WL1);
    • wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format (ImFinFmt).
  • The new technical effect hitherto unmanageable is that this approach to in particular the operating system construction and its configurable communication with applications and their visual color asset desiderata, by means of the new color transformation control data (CTctrl), makes it possible that the operating system creates better coordinated colors for the one or more visual assets, for ultimate display. So the final displayed look of all the visual assets together will be more controllable towards a better looking total output picture. Unit in terms of software code may typically mean a sub-process, which will run on an electronic circuit (it will have access to memory for temporarily storing the data, and may have access to reconfigurable processing, such as different algorithms for color transformation, and different manners of deciding the geometric positioning of the various transformed pixels on a canvas, such as a windowing system, the details of the latter being of lesser relevance for understanding the color coordination technology of the present technical system). The brightness mapping unit (710) will receive the pixel color codes of a visual asset (from the geometrical management unit 711) to be brightness mapped to a different brightness range, and apply e.g. a function-based transformation to at least the pixel's input luma to obtain the output luma as needed for a particular brightness range which the operating system decides to use in its final presentation of the aggregate one or more visual asset (e.g. on a background). This suitable (final) brightness range may be determined based on various parameters, e.g. for which kind of display the composited canvas is generated, user preferences, the kind of total presentation (e.g. a user-interface-focused presentation), etc., but those details are not the core aspects of the present innovation. The original ranges of the assets may be determined by the various applications based on very different criteria. E.g., a game which expects lots of bright explosions to be displayed on the end-user display, may make its graphics (e.g. a notation of successful shots) equally super-bright.
  • A set of lumas respectively colors may for many situations simply means one or more arrays. There will be an array of e.g. 200x100 luma code values (e.g. for a graphics to put on the top-left of the screen), and typically two more arrays for the chromas. In other situations the set of lumas (respectively colors) will be specified by drawing instructions (e.g. draw a line from beginning position (x1,y1) to end position (x2,y2) with luma 422 out of 1023).
  • An important technical property is that the original asset colors will be communicated according to a (possibly a few selectable alternatives, but ideally a single) common format for a communication version of a pixelated image of an asset (example elucidated with Im1), or a procedural description of generation of an asset in a color space (example of Tx3), which the operating system can prescribe to all the communicating applications.
  • A color code representing an input luma code means that, whichever the color specification/model used in any detailed embodiment, the coding of the color is specifying colors in the respective gamut (e.g. an SDR gamut, or some HDR gamut). So it may be that the color model directly comprises the input luma code, as in the model which is popular for video: Y'CbCr, in which Y' is the luma code of a pixel. However, all these colors (i.e. the e.g. SDR gamut) can also be represented in another additive color model. E.g., in the computer world the color code R',G',B' is a popular coding. The pixel luma (and therefore in case of an absolute system its luminance) is then still uniquely represented (coded) since it follows from a universal fixed colorimetric equation: Y'=a*R'+b*G'+c*B', in which a,b, and c are fixed constants, depending on the primaries of the gamut.
  • A suitable embodiment of a common format will be based on an SDR color gamut, e.g. a normalized Rec. 709 primaries gamut in which the brightest white luma code may represent 100 nit (e.g. the 8 bit luma code 255 dictates that pixels having these values, are to be displayed at 100 nit on any display, if not transformed into a tertiary brightness). Note that it is not obligatory that the (advantageously e.g. SDR) color representation for communication (i.e. of the Im1 which codes the spatial and color structure of the first asset Ass1, or a corresponding set defined as vector graphics) to the OS has an actual maximum luminance as a nit number. It is enough that the function Tm 1 is able to allocate (during the reconstruction) an original maximum luminance (ML_V) to the largest possibly occurring code of the pixels in Im 1 (or equivalently a percentage of that largest HDR white if the asset only goes as bright as e.g. 40% grey). In other words it is sufficient if one can establish the dynamic range (preferably in absolute luminances, or alternatively relative brightnesses) in which the asset resides, with its HDR white if the asset actually has white pixels, or compared to its HDR white if it is darker (which can be achieved by giving it lower sRGB values in an Im 1 which forms a proxy for a e.g. 1000 nit video: even if there is no pixel in the asset which actually reaches 100% sRGB, i.e. to be shown as 1000 nit; it suffices that TM1 instructs, via its shape, that pixels of 100% sRGB would be reconstructed to 1000 nit, rather than in another TM1 shape to e.g. 2000 nit or 750 nit). In such a case the maximum brightness of the communicated image may be taken as fixed to 100% (the normal sRGB interpretation). Note that this 100%, after down-grading, will correspond to some potentially quite bright HDR maximum luminance, and the actual "SDR white" (i.e. a dimmer white that one would see in the scene as e.g. an averagely lit piece of paper, casu quo display on a display as a typical SDR-brightness range white, e.g. 200 nit on a display than can go as high as 600 nit) may be coded in the sRGB communicated image e.g. at luma 60%, or 30% (and that level can be communicated as the reference white level WL1). Even though both the original and communicated representation may be using the same amount of bits per color component, e.g. 10 or 12 (or the original may be 3x12 bit and the communicated image Im1 3x8 bit), the difference in luminance range will be apparent from the squeezing together of at least some of the object colors, e.g. typically the brighter colors (this can be verified for relative communication images by allocating an actual maximum luminance to the maximum brightness, i.e. the relative brightness of 100% luma code, and looking at the histogram and inter- and intra-object contrasts, e.g. the ratios of the pixel brightnesses of two brighter pixels will be smaller in Im1 than in the original image). Several technical communication mechanisms may fulfil this property, but one can e.g. formulate TM1 as a function which maps the largest possible communicated pixel luma, i.e. input luma code = 255* 100% to the level of 1.0 of the reconstructed HDR range, and then associate a maximum luminance ML_V with that maximum HDR output luminance or luma code, which ML_V will in such an embodiment get communicated as part of the definition of the first brightness mapping function TM1. This function could in some embodiments directly map input luma codes to normalized output luminances, and one need then only multiply 1.0 (or any value below) by ML_V to obtain the absolute output luminance of the pixel of the reconstructed output color. In case the brightness mapping unit embodiment directly produces relative brightnesses or especially when it produces absolute luminances, the output luma code (Y'o) may be coded as a native, linear coding of said brightness respectively luminance. In case the communication to the OS communicates the (typically down-grading) original function TM1 determined by the software application, the OS will invert that received function (possibly scaled for display adaptation) before doing its brightness range conversion. In some embodiments a flag may indicate whether the direct or already pre-inverted form of the brightness mapping function is put in the metadata (or the system may work in a pre-agreed manner).
  • It may alternatively be more pragmatic to calculate in luma code domain also for the output, e.g. PQ lumas. Since the PQ definition already has a universally recognized maximum of 10,000 nit, if one then communicates a relative function TM1 which maps from e.g. sRGB normalized lumas to PQ lumas, then if 100% SDR luma maps to e.g. 75% PQ luma, we know the brightest pixels of the communicated image is ideally to be displayed as 1000 nit, because 0.75 in PQ means 1000 nit (unless the display or receiving side apparatus still needs to display optimize for a display of lower maximum display luminance, in which case it will use the color transformation control data CTctrl, and specifically TM1 to guide how this down-grading should ideally happen).
  • The original colors and their lumas and the luminances (as an amount of nit a.k.a. Cd/m2) or relative brightnesses they code (relative to e.g. the 100% level of SDR white), i.e. of the visual asset as the creating/communicating application ideally wants to see it displayed, may be considerably different than the coded colors as communicated: i.e. the normally decoded colors of the communicated coded proxy colors (using the normal sRGB definition instead of the brightness mapping function TM1 for the decoding) will usually be in a much smaller secondary brightness range, which ends at e.g. 100 nit instead of the original e.g. 20,000 nit. Note that the secondary range as communicated may also be larger than the original range of the visual asset, in particular it may end at a larger maximum luminance than the original maximum luminance of the asset pixels, but usually a smaller brightness range ending at a lowered maximum luminance will perform sufficiently well for communicating assets to the operating system.
  • Because the operating system also receives the color transformation control data (CTctrl), it can perform suitable tertiary transformations in line with what the original asset's colors were, i.e. are supposed to be displayed as. E.g., the operating system may chose the tertiary brightness range to be identical to the primary brightness range, and invert the brightness mapping function (TM1), and use inverted first brightness mapping function (ITM1) on the input luma codes to obtain output luma codes which are a reconstruction of the original lumas of the asset of the application.
  • But in general things won't necessarily be so simple. If different applications (e.g. a game and a website, or an encyclopaedia) use assets of considerably different brightness range, and in particular maximum luminance or maximum relative brightness, the operating system may want to coordinate the output luma of the geometric composition of various assets together in a total campus, into a tertiary brightness range which may end lower than a primary brightness range of at least one of the assets of at least one of the applications. The operating system may also take into account what display will be served by the GPU with the composite images of its scanout buffer (ScOBff), and if that display does not have a high displayable brightness range, or the viewer wants to see everything bright near the upper end of the displayable brightness range, the OS may take this into account when firstly electing a tertiary brightness range, and secondly deriving optimal tertiary brightness mapping functions (FL_op) for mapping the various luma code values of the various assets to that common range. In general the mapping any luma code gets, which can equivalently be described as a multiplication by a multiplier which depends on the value of that luma code (a multiplier larger than 1 indicating a -relative if performed in a normalized to 1.0 representation of the lumas or absolute- brightness boost, and a multiplier smaller than one performing a brightness dimming), will depend on where in the input range the luma code falls. E.g., as one may want to squeeze most or all of the possible input lumas into the tertiary/output range, the amount of boost a luma say halfway gets may depend on how much the darker lumas get brightened in their mapping. So in general each pixel luma will get an optimal mapping so that the range of input brightnesses (e.g. luminances in some embodiments) is well-represented in the output range, but the many shapes of brightness mapping function that any application or the operating system may chose is also a detail we need not dive into, since the new technical system construction must be able to function which substantially each desired function.
  • It will be shown that some embodiments of the operating system's common mapping may depend only on the communicated at least one brightness mapping function, whilst others may work solely on the communicated at least one reference white level (WL1), whereas other more sophisticated tertiary mappings may design their tertiary mapping function (or algorithm, e.g. taking into account the spatial nature of an asset, such as geometrically non-uniform shading) on both of those communicated color transformation control data elements.
  • Whereas the brightness mapping function essentially communicates how many luminances respectively lumas of the original asset's colors are squeezed into the typically smaller brightness range of the pixel color array (Im1) that gets communicated, one may know where the brightest street lamp falls in that lower brightness range (namely near 1.0, or 255 in an 8 bit coding of the brightnesses respectively luminances), but one doesn't know yet where the reference level of the uniformly lit (under the average base lighting of the scene, usually the bigger area of the geometrical frame of the images) Lambertian diffusive white object falls. That reference white level (WL1) of the asset might (depending on how much brighter the brightest HDR objects in the original asset representation are) e.g. fall at 128 for a not so impressive brightness dynamic range (also depending on which EOTF is used for the lumas, e.g. PQ being pre-agreed between operating system and applications, or HLG), but it may also fall on luma value 55. So that 100% Lambertian reference white level (or e.g. 90% of that level if one desires), may also be communicated, and be used to the benefit by the OS when determining its optimal tertiary mapping function(s).
  • So the first mapping (with TM1) will map the original pixel colors and their lumas, originally lying in the first luminance dynamic range which is determined by whatever the asset was created to be (e.g. a games designer may create a blue laser beam which is as bright as 8000 nit). The secondary brightnesses may be coded as lumas which either code absolute nit secondary color lumas, but which may pragmatically end at a lower maximum luminance, e.g. 100 nit for SDR (reversible) common asset communication, or relative/percentual brightnesses. This first mapping establishes the relationship between the original asset colors, and the common interface colors (in the pixel color component arrays) which actually get communicated to the OS. So conversely, for the OS only getting the interface colors, it establishes what the original colors were, and were supposed to be in a displaying. So the maximum brightness of the secondary colors will typically be lower than that of the original colors of the various assets (transforming HDR assets into SDR assets, but in a smart invertible manner by using typically invertible or largely invertible first mapping function(s) TM1). But as regards the tertiary mapping the OS can basically do what it wants (except for it should normally try to fulfil the desired look of the original colors, by at least taking into account the original first mapping function and or reference white level, and typically using functions which keep the color look of most of the colors, such as their differences, still reasonably similar as far as the OS-side, e.g. output-side desiderata enable (which can be elegantly realized e.g. by giving the tertiary mapping function a shape which is similar to the shape of the first brightness mapping, e.g. a weakened down version of that function). E.g. it may lift primarily the darkest colors, and shift the whole secondary brightness range (or even when considering from the original brightness range) upwards to basically much brighter colors for display.
  • By geometrical management unit (711) is a unit which manages the geometrical aspects of the assets, and their basic ingestion. So e.g. it will determine the position and possibly scale of assets in the total canvas (but not solely as a set of parameters, such as a top-left (x,y) coordinate pair, but generically also with its content filling, e.g. an image (i.e. the which asset to go where)), potentially in windows and the like. In this innovation, it will also take care of the control of luma processing on demand (to the brightness mapping unit functionality), so they become already of the correct color, and only the geometric aspects of pixel placement are in order. E.g. this unit may determine a window size and position for a window in the total canvas showing a video, and it may then scale (by geometric interpolation) the various correctly mapped pixel colors received from the brightness mapping unit, so that the video fits the window. Brightness mapping unit means any unit that can do brightness processing for one of more pixels of an asset, so that ultimately these pixels will have the desired brightness, either relative to some intermediate or maximum value, or absolute in nits. We show just a simple version for understanding where a single brightness mapping function will be applied merely on the value of the input luma, irrespective where it resides in the asset, but more complex scenarios could involve a central mapping function for mapping the pixels in say a circle in the middle of the asset, and a surrounding mapping function for the surrounding pixels. In that case the application will, for this asset, communicate two brightness mapping function, and instructions where to apply which function (e.g. with a bitmap, where 1 means first function and 0 means the second function should be applied by the operating system, or more precisely, its tertiary function should be based on that second function). Any geometric information needed may be communicated by the application in a geometric data section (Geo1).
  • Finally, once the assets have all been suitably mixed, the problem is less difficult. It may determine an output pixelated image format (ImFinFmt) for writing the geometric composition in the total canvas of the one or more assets into one or more portions of the scanout buffer, e.g. first memory portion 720 and second memory portion 721. We will assume -without wanting to be limiting- a common format which is also useful for further communication by the GPU to a display, such as a perceptual quantizer luma based format (such a 10 bit format may represent pixel luminances up to 10,000 nit, and even if the original asset had some higher pixel luminances, this may be satisfactory; for relative brightness communications one can still use this coding, by assuming, or explicitly communicating to a display by a further metadatum, that some value is the 100% value, e.g. 100 or 200 nit).
  • A software application is a computer program designed to carry out a specific task other than one relating to the operation of the computer itself, i.e. other than the basic control of the computer hardware. It may implement various functionalities for the user, e.g. online shopping, presentation of visual media, etc. In the present discussion we need not formulate differences between applications for e.g. classical personal computers, or computer-orchestrated professional systems, and apps for portable apparatuses (the shorthand app can mean all of those).
  • It may be advantageous in practice to have most or all software applications (being directed to) communicate to the operating system with color codes of the various pixels of their one or more visual assets represented in an sRGB color representation. This is a well-understood reference color system, but for SDR, but now with the additional color transformation control data CTctrl it can be used to also communicate a myriad of different kinds of HDR assets.
  • Advantageously the operating system will work in a manner enabling software applications to specify and communicate the original brightness of any pixel of any asset specifying a luminance of a pixel to be displayed as an amount of nits. This means that such an embodiment will communicate, even if the image array of Im1 itself codes only relative (0-100%) lumas or normalized brightnesses, or only relative non-linear R'G'B' values are communicated, the totality with the color transformation control data allows the receiving operating system to establish for each pixel (i.e. its reconstruction, or some derived re-graded image of pixels) an absolute luminance. This can be performed e.g. by specifying the brightness mapping function TM1 in a format which established or allows to establish absolute nit outputs, such as Perceptual Quantizer EOTF luma output domain values for the function (or equivalently in other embodiments one could add additional data to the color transformation control data CTctrl enabling e.g. an absolute scaling of the normalized to 1.0 brightnesses, such as a common multiplier, typically a maximum luminance value for the image or video ML_V).
  • Corresponding to the color-coordinating operation of the operating system, there will be software applications (e.g. web application 502, or native application 503) arranged to create at least one visual asset (Ass1) comprising pixels wherein a pixel has an original color code which specifies an original brightness and to communicate such visual asset to an operating system (504), wherein the original brightness lies in a first brightness range which has a first maximum brightness (ML_V);
    • wherein the original color code is transformed into a secondary color code for communication to the operating system, wherein the secondary color code represents a second luma code, which codes a secondary brightness which lies in a secondary brightness range which is different from the first brightness range, wherein the secondary brightness is derived from the original brightness based on application of a brightness mapping function (TM1) to the original brightness;
    • characterized in that the software application communicates color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Ass1), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1).
  • The application may elect what original HDR format it will use for its asset's colors, but PQ luma-based colors will be a useful manner. E.g., the original asset pixel colors may be 0-2000 nit absolute luminance colors, represented as equivalent PQ Y'CbCr or PQ R'G'B' values. Actually, it doesn't matter if the color representation for unified communication is e.g. sRGB, since then the receiving operating system need only understand the original colors per se from the proxy sRGB color, and it may internally generate its output colors (of the output intermediate pixelated image format ImFinFmt) in e.g. Philips EOTF-based format, a logarithmic representation of the luminances or color components, etc. Especially absolute systems will be understandable because the luminance is an absolute optical quantity (and the other two color components, e.g. Cb and Cr, establish what are also universal color properties, namely a hue like e.g. Chartreuse, and a saturation). But also the relative systems can be sufficiently unique, since then one has a percentage of some maximum, e.g. a display maximum, or the maximum of a composited total canvas presentation decided by the operating system.
  • Just like the original color codes (whatever their actual codification) will represent original brightnesses, the secondary color codes for actual communication of the asset to the operating system will represent a secondary brightness, along a different range of brightnesses, due to the brightness mapping (note that often advantageously the actual mapping processing may be applied to a brightness component per se, but one can also map on other representations, e.g. the RGB components, equivalently, so that the brightness mapping is achieved, in the whichever representation). The important point is the common interfacing, allowing the coordinated use (typically re-grading) of any application's asset by the operating system. In case the brightnesses are represented as lumas according to some elected EOTF, e.g. PQ, the derivation of secondary brightness from the original brightness based on application of a brightness mapping function (TM1) to the original brightness may involve first converting e.g. absolute nit values to input (original) luma codes (e.g. PQ lumas) and then applying the function in the PQ domain to obtain e.g. output PQ lumas (in case output luminances are required, the PQ EOTF can be applied to those output lumas).
  • The color transformation control data (CTctrl) is associated with the asset (Ass1), which can happen in many manner, but typically the API will e.g. first communicate the color code array(s), or the data of the procedure to generate a set of pixel colors at a geometrical management unit of the OS, and thereafter the color transformation control data (CTctrl), or vice versa, it first sends the control data and then the pixel color data. The first brightness mapping function (TM1) or its inverse, and a reference white level (WL1) are for use by the operating system to understand which exactly of the many possible HDR colors of the asset the coding of the asset as received (e.g. image array(s) Im1) originally represented, and therefore also how it should ideally present (re-grade) such asset colors in any of the many possible final composited canvases it may want to generate, and send to the GPU for ultimate display. As a first application may send an asset of a very different maximum brightness (e.g. ML_V1= 8000 nit) than a second application (e.g. ML_V2= 1000 nit), the operating system can understand this, and then also in any embodiment of its brightness mapping unit 710 decide how to coordinate those two assets in the final canvas (e.g. it may dim the first asset, or boost the second one somewhat, etc.).
  • The applications (or even the operating system) may be supplied to the computer, via some software communication mechanism.
  • The techniques may be embodied as a method of communicating a visual asset having pixel colors to an operating system, comprising the steps of:
    • creating at least one visual asset (Ass1) comprising pixels wherein a pixel has an original color code which specifies an original brightness, wherein the original brightness lies in a first brightness range which has a first maximum brightness (ML_V);
    • transforming the original color code into a secondary color code for communication to the operating system, wherein the secondary color code represents a second luma code, which codes a secondary brightness which lies in a secondary brightness range which is different from the first brightness range, wherein the secondary brightness is derived from the original brightness based on application of a brightness mapping function (TM1) to the original brightness;
    • communicating the secondary color code and color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Ass1), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1).
  • The techniques may be embodied as a method of operating a computer, comprising a step of running an operating system, wherein the operating system is arranged to manage color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Ass1) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness (ML_V), wherein the visual asset (Ass1) is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness;
    • wherein the visual asset in addition comprises color transformation control data (CTctrl) which comprises at least one of first brightness mapping function (TM1) or its inverse, and a reference white level (WL1);
    • wherein the operating system comprises a brightness mapping unit (710) arranged to receive the visual asset and to apply a tertiary brightness mapping function (FL_op) to the input luma code to obtain an output luma code (Y'o), wherein the output luma code specifies a tertiary brightness which lies in a tertiary brightness range which is different from the secondary brightness range, wherein the output luma codes a brightness of an output color, wherein the tertiary brightness mapping function (FL_op) is based on at least one of the first brightness mapping function (TM1) and the reference white level (WL1);
    • wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
    BRIEF DESCRIPTION OF THE DRAWINGS
  • These and other aspects of the method and apparatus according to the invention will be apparent from and elucidated with reference to the implementations and embodiments described hereinafter, and with reference to the accompanying drawings, which serve merely as non-limiting specific illustrations exemplifying the more general concepts, and in which dashes are used to indicate that a component is optional, non-dashed components not necessarily being essential. Dashes can also be used for indicating that elements, which are explained to be essential, but hidden in the interior of an object, or for intangible things such as e.g. selections of objects/regions.
  • In the drawings:
    • Fig. 1 schematically explains various luminance dynamic range re-gradings, i.e. mappings between input luminances (typically specified to desire by a creator, human or machine, of the video or image content) and corresponding output luminances (of various image object pixels), of a number of steps or desirable images that can occur in a HDR image handling chain;
    • Fig. 2 schematically introduces (non-limiting) some typical examples of use scenarios of HDR video or image communication from origin (e.g. production) to usage (typically a home consumer);
    • Fig. 3 schematically illustrates how the brightness conversion and the corresponding conversion of the pixel chromas may work and how the pixel brightness histogram distributions of a lower and higher brightness, here specifically luminance, dynamic range corresponding image may look;
    • Fig. 4 shows schematically how computer representations of HDR images would represent a HDR range of image object pixel luminances;
    • Fig. 5 shows on a high level the concepts of the present application, how various applications may be dealing with the display of pixel colors via and/or in competition with all kinds of other software processes, which may not result in good or stable display behavior, certainly when several different displays are connected or connectable;
    • Fig. 6 shows an archetypical example of a user scenario, the user dealing with and looking at several HDR-capable web-applications in different windows occupying different areas of a same screen, may benefit from the current new technical approach and elements;
    • Fig. 7 schematically shows an embodiment of a computer according to the present innovations running a few applications which coordinate their assets with the innovative operating system according to the innovative technical specification format for the asset colors of arbitrary higher brightness dynamic range;
    • Fig. 8 schematically illustrates one embodiment of an algorithm by which the OS can beneficially use the new color transform control data to come to a total canvas for showing together several of the visual assets in a well-coordinated manner, respectful of both the original intended color and the appearance of the totality;
    • Fig. 9 schematically illustrates how the OS can derive secondary brightness mapping functions from any brightness mapping function it receives in the color transform control data; and
    • Fig. 10 discusses (without intending to be limited) so more details on how the present coordination framework concepts can be mapped internally in an OS.
    DETAILED DESCRIPTION OF THE DRAWINGS
  • In Fig. 5 we show generically (and schematically to the level of detail needed) a typical application scenario in which the current innovation embodiments would work. As this is currently becoming ever more popular, we show -from the client side system (500), i.e. e.g. running on a personal computer or mobile phone, how a web-app may work via a native app and finally to the operating system (OS), which may coordinate the ultimate communication with the GPU and management of ultimate display. A mobile phone, unless when in a screen casting application, will typically have one display only (second (HDR) display 522), but other client side systems may communicate with several displays. We have shown some options in dotted, to indicate optionality, e.g. the web application (502) may be communicating directly with an operating system (504) without an intermediate native application (503), or there may only be a locally installed and running (specially written for some system) native app involved, at least for some of the visual assets being prepared for display, and no contacting with anything over the web regarding those visual assets.
  • The web application could be e.g. an internet banking website, remote gaming, a video on demand site of e.g. Netflix, etc. The web application will run as software on typically a central processing unit (CPU) 501. E.g., the composition of a web presentation may be received as HTML code. The web application will be to a large extent (for the user interaction) based on visual assets, such as e.g. text, images, or video.
  • In Fig. 6 we have again generically (without wanting to be needlessly limiting) shown what the user would see on his one or more display screens, displaying a total viewable area or canvas (the screen buffer 601), containing two windows showing two such applications. The web application may have needs to show its assets in high dynamic range (i.e. better visual quality, in particular a range of brighter pixel luminances than SDR, and correct/controlled pixel luminances). E.g., it may show HDR still pictures showing more beautifully than with a plain SDR JPEG to let the customer browse through new available movies. Or it may want to present commercial material which is already in some HDR format (or vice versa, an old not yet HDR commercial asset, i.e. still in Rec. 709).
  • E.g., a first web application may present its content in a first window 610. The content of this window, may be on the one hand a first plain text 611 (i.e. graphics colors for standard text, which is well-representable in SDR, e.g. the colors black for the text on an "ivory" background). On the other hand, this webpage, which gets composed in the first window for the viewer, may also present a HDR video 612 (e.g. a video streamed in real time from a file location on another server, coded in HLG, and assuming 1000 nit would be a good level for presenting its brightest video pixel colors). This window could be side by side on a total canvas of the windowing server, which we will here call "screen buffer" (to discriminate from local canvases that windows may have). Canvas is a nomenclature from the technical field, which we shall use in this text to denote some geometrical area, of a set of N horizontal by M vertical pixels (or similarly a general object like an oval), in which we can specify ("write" or "draw") some final pixels, according to a pixel color representation (e.g. 3x8 bit R,G,B). Even if there was only one such window on the screen buffer, it could already partially project on a first display 520, and partially on a second display 522, e.g. when dragging that window across screens. If the screen buffer were to represent the pixels in one kind of color representation, e.g. SDR, then one of the challenging questions would already be how the second display, which is a HDR display, which expects PQ-defined YCbCr pixel color representations, and can display pixel colors as bright as its second display maximum luminance ML_D2 capability being 1000 nit (compared to first display maximum luminance ML_D1 capability being 100 nit) would display its part of the window. The viewer may find it at least weird or annoying if a white text background suddenly jumps from 100 nit to 1000 nit. Note that there will also be Graphical User Interface (GUI) widgets, such as first window top bar 613, and first window scroll bar 614. The operating system (OS, 504) may deal with such issues (although it sometimes also expects applications to deal with at least part of the GUI widgets, e.g. remove buttons when the application is in full screen, and re-display them upon an action such as clicking or hovering at the bottom of the screen, or showing the widgets partially transparent, etc.). Usually, at least for the moment, such widget graphics may be solely in SDR, but that may change in the near future.
  • In the second window 620 we show another web application. The two web applications may not know about each other (and then cannot coordinate colors or their brightnesses), but the local computer will (e.g. the windowing server application of the operating system). The second window, apart from its second window top bar 623 and second window scroll bar 624, may e.g. comprise a HDR 3D graphics 622 (say fireworks generated by some physical model, with very bright colors, illuminating a commercial), and in another area second text 621, which may now e.g. comprise very bright HDR text colors (e.g. defined on the Perceptual Quantizer scale, say non-linear R_PQ, G_PQ, B_PQ values). The various parts of the screen (windows, or parts of windows, background), may correspond to areas and their pixel sets, e.g. area tightly comprising all pixels related to the first window and nothing from the screen background (shown slightly larger to be visible).
  • Applications nowadays call functionality of other software (i.e. calculations or other actions that the other software processes can do) or hardware via Application Programming Interface Specifications (API). This hides the very different deep details of e.g. another software, or a hardware component like a GPU (like its specific machine code instructions that must be done). This works because there is universality in the tasks. E.g., the end user, or any upper layer application, doesn't care about how exactly a GPU gets the drawing of a line around a scroll bar done, just that it gets done. Also for HDR pixel colors, although there may be many flavors, at least the way in which e.g. a particular pastel yellow color of 750 nit gets defined and created (for ultimate display) at a certain pixel, just that this elementary action gets done. The same is true for a rendering pipeline, which usually always consists of e.g. defining triangles, putting these triangles at the correct pixel coordinates, interpolating the colors within the triangle, etc. (and it doesn't matter whether one processor does this, even running on the CPU, or 20 parallel processing cores on the GPU, and where and how they cache certain data, etc.).
  • At least in the SDR era this didn't matter, because SDR was just what it was, a simple universal manner to represent all colors between and projected around black and white. And a certain blue was a certain blue, and therefore GUI widget blue could be easily mixed (composed) with movie or still image blue, etc.
  • So e.g. the web application may specify some intentions regarding colors to use (e.g. for unimportant text, important text, and text that is being hovered over by the cursor), e.g. by using Cascaded Style Sheets (CSS), but it may rely for the rendering of all the colors, in a canvas, on a native application 503. Since applications may desire a well-contemplated and consistent design, CSS is a manner to separate the content (i.e. e.g. what is said in a news article's text) from the definition and communication of its presentation, such as its color palette. For web browsing this native application would be e.g. Edge, or Firefox, etc. So it may call some standard functionality via a first API (API1). The native application may do some preparing or processing on the colored text, but it may also rely on some functionality of the operating system by calling a second API (API2). But the operating system may, instead of doing some preparatory calculations on the CPU, and then simply write the resultant pixel colors in the VRAM of the GPU, also rely on the GPU to calculate the colors in the first place (e.g. calculate the fireworks by using some compute-shader). In any case, it may issue one or more function calls via third API (API3). In fact, modern universal graphical APIs like WebGPU, may have the web page already define calls for the GPU (fourth API API11). An example of a command in WebGPU is copyExternalImageToTexture(source, destination, copySize), which copies the contents of a platform canvas or image to a destination texture. The various APIs may also query e.g. what the preferred color format of a canvas is (colors may be compressed according to some format to save on bus communication, needed VRAM), etc. The final buffer for colors of the screen to be shown, in the GPU 510, is the so-called ScanOut Buffer 511. From there the pixels as needed to be communicated to the display, e.g. over HDMI, are to be scanned, and formatted into the needed format (timing, signaling, etc.).
  • A first problem is that, although one may want to rely on other (layer) software for e.g. positioning and drawing a window, or capturing a mouse event, dynamic range conversion is far too complex, and quality-critical, to rely on other software components to do some probably fixed and simplistic luminance or luma mapping. The creator of the web page, or the commercial content running on it, may have spent far too much time making beautiful HDR colors, only to see them mapped with a very coarse luminance mapping, potentially even clipping away relevant information. If it was a simple color operation, such as the correction of a slight bluish tint, which would also not involve too severe differential errors or too different a look if not done appropriately, or at all, one could rely on a universal mechanism for doing that task, which could then be done anywhere in the chain. But color-critical HDR imagery has too many detail aspects to just handle it ad hoc.
  • Indeed, until recently, as a second problem the various applications might have defined various HDR colors (internally in their application space), but the windowing servers (e.g. the MS windows system) typically considered everything to be in the standard SDR color space sRGB. Because that is what they understood and had been using for decades, and HDR was complex, multi-variant, and ill-understood nor agreed. So you had HDR colors, but you needed to work with SDR colors, or that was what was going to come out to be able to see anything.
  • That was good for compositing any kind of SDR visual asset together. But an application having generated e.g. very colorful 750 nit fireworks, would need to luma map those to dull SDR firework colors, and then the windowing server would compose everything together on the screen buffer. Even if somewhere in between there would be some re-mapping to another dynamic range, that would be very ad hoc.
  • So at best any application would need to convert its beautiful HDR visual content to sRGB visual assets, and also there would be no coordination with other application's assets. Furthermore, worst case there could be a cascade of various luma mappings in the various layers of software acting on top of each other, e.g. some of those involving SDR-to-HDR stretching, and all in an uncoordinated manner, ergo, one would have no idea of what came out displayed at the end, to the detriment of visual quality and pleasure of the viewer (e.g. (parts of) windows could be too dark, of insufficient contrast, if some pixels colors lay outside expected boundaries of some software some parts of the visual asset may have disappeared when displaying, etc.). Recently some windowing systems have introduced that applications can communicate to the window composition process canvases that can contain 16 bit floating point (so-called "half float") color components (and consequently luminances of those colors), however this only allows to basically communicate a large amount of colors (per se), but doesn't guarantee anything how these will be treated down the line, e.g. which luma mapping will be involved when the operating system makes the final composition of all visual assets, to correctly make the screen buffer, and ultimately the scanout buffer of the GPU. Note that how the screen buffer (especially if several displays are involved in an extended view) is exactly managed in one or more memory parts, with one or more API calls etc., is not critical to the present innovation embodiments. E.g., one display may refresh at higher rate than the other, and then its part of the total canvas will be read out more frequently. The innovation is about allowing the optimal coordination of the various parts the user ultimately gets to see, via the careful communications and attuned handling of the coordinated reference image representations (e.g. sRGB) and associated re-grading functions for defining HDR assets, and the coordinated use of it all when e.g. display adapting the final composition or part thereof.
  • In Fig. 7 , which shows a computer 700 (which could also be e.g. a mobile phone) connected to a (external or internal) display, a first software application 701 may be showing e.g. images of a motorcycle. Say this motorcycle is a 2000x1000 pixel image, which may get re-scaled to the total output canvas by the OS. Its asset may be defined primarily by an image Im1, which may be e.g. 3 pixel color component arrays (MPEG compressed, or non-compressed), e.g. an array for the lumas Y', and one for the blue chromas Cb and the red chromas Cr. This asset is supplemented by associated metadata, e.g. in this first example at least a first brightness mapping function TM1. The ellipses around the first reference white level WL1 in the schematic Figure indicate that this data may be present in the communication to the OS (e.g. API call), or not. The may be geometric information Geo1, which can e.g. indicate preferred size, position, whether the application would like to be on top at least partly (or transparency information), etc. Especially if this application communicates multiple assets, the geometric information may convey the relative positioning of these assets. A second software application 702 may be showing e.g. an encyclopedic article about bats (either from an internal memory of the computer device, potentially composed on the fly and upon request of a user, or from internet, etc.). It may communicate to operating system 504, again according to the innovative stable mechanism of the present application, two assets. Second asset Ass2 may e.g. be the image (potentially a computer-generated graphic instead of a photo) of the bat. In this case this graphic is communicated again as an (second) image Im2 (again one or more color component arrays coding the respective magnitude of the color component for any pixel position in the array(s)). In addition the second color transform control data now has both a second brightness mapping function TM2 and a second reference white level WL2. The third asset (the text below the bat graphic) may be procedural in coding. E.g. third pixel set color specification procedure Tx3 may give a HDR color to ASCII-coded text characters, and another HDR color for the background, both of which will be converted to corresponding sRGB color codes (or the like in similar embodiments), yet those can be re-graded to the HDR colors (preferably both with the same inverse function; though in general they could have their separate transformation). In this example (non-limiting), we have shown that we could also merely give a third reference white level WL3. Based on this level, the operating system can then judge how bright or dark the creating software desired the colors to be, and take this into account when positioning the colors in the tertiary brightness range (e.g. it may boost colors to the level of an explosion, yet still, compared to that new reference level, keep dark colors of the text or background sufficiently dark).
  • The brightness mapping unit 710 of the OS can, as shown symbolically, both apply upgrading (concave) or downgrading (convex) function to the input sRGB lumas, as the need may be. It will produce the correct tertiary range output lumas Y'o, e.g. typically in a re-graded image ImScal. Finally the GPU can communicate the total composited image to a display 750, via some image communication path 751, e.g. a HDMI cable, Wi-Fi screencasting, etc. The geometrical management unit (711), e.g. a window compositing unit of the OS, is shown as the unit which does the overall asset management, i.e. receives or retrieves all necessary information, does the brightness mapping via call to brightness mapping unit 710, and ultimately sends the composited image to one or more scanout buffers or portions of buffers (first scanout buffer portion 720 and second scanout buffer portion 721) of the GPU 510.
  • Fig. 8 elucidates an embodiment of how the OS can use the new color transformation control data, namely a first brightness mapping TM1 for the video image Im1 (of the monster in the cave), and a first reference white level WL1 (here e.g. luma 64 out of 255 luma codes), and simply a third reference white level (WL3=100) for some graphics text, in deciding its final (original color-attentive) tertiary brightness range composition of the total assets canvas.
  • First regarding the pixel normalized luminances (on the output vertical axis, where the OS has chosen it needs a 2000 nit tertiary range maximum (ML_COMP) for the composition, ergo the linear scale endpoint means 2000 nit; the horizontal input axis are just sRGB input lumas Y'_in_SDR, i.e. approximately the square root of the linear relative brightnesses). The first reference white level WL1, can be used by the OS to e.g. construct a first, lower brightness part TMd_opt of the mapping curve (from input lumas of pixels in Im1, to output luminances). E.g. the OS can determine (non-limiting) a "straight" line allocation (we have symbolically drawn a straight line, but since the input is square root, the shape should be approximately power 2), which maps the identified first reference white level WL1 to a first common white level WcomL1 of the composite canvas, and all input linear brightnesses of the Im1 correspondingly linearly proportional below this juncture point PJ1. This becomes elegantly configurable: the OS can establish e.g. 250 nit to be a good first common white level. The brighter input lumas Y'_in_SDR will follow a shape-conforming optimal mapping function TM1_opt, which largely follows the shape guidance of the original communicated brightness mapping function TM1 of the first asset (see further elucidated in Fig. 9 how this can be achieved). Essentially, this means that primarily the flames should be well-boosted (note that the diagonal should be interpreted with a brighter output maximum luminance: if we allocate e.g. 100 nit to the input normalized maximum, mapping according to the diagonal already corresponds to a 20-fold brightening of all pixel luminances, at least if the input axis was linearly represented). We however also see that the creator of the TM1 function, i.e. the first software application, and potentially any human behind it, has also created a steep slope where the monster is coded, ergo, this may mean that he wanted the monster to be rather contrasty in any re-grading, ergo this guiding shape should be followed at least as far as achievable (as we will also see in Fig. 9, the shape of a function, i.e. basically how the output varies for increasing inputs, can elegantly be described by a set of distances of successive points on the diagonal to the locus of point of the function, e.g. PJ1). So the shape of TM1_opt, will at least be based on the communicated TM1 for asset Ass1 (Im1), and often also based on WL1.
  • This constitutes already a good final coloring (re-grading) of the first asset, on the tertiary range image (i.e. ImFinFmt, or its internal intermediate precursor of different colorimetric definition), for the GPU, and ultimate display. Note that the first reference white level WL1 can be communicated even though there is no object in the current image(s) that actually has this value, as it can still be used as important reference value by the OS, or the instructing software application.
  • Now the second asset is placed in the total canvas, after establishing suitable luma re-grading. We show an example what can be done if the text is encoded in a third sRGB image Im3 (we now assume without limitation that the text is not ASCII, but already communicated as a pixelated sRGB image comprising a rendering of that text), together with only a third reference white level WL3.
    The OS sees that the text (e.g. yellow text with luma 175) is actually supposed to be brighter than the communicated third reference white level WL3. So the second software application (or the first SA for this third asset), wanted to show above averagely bright text, which can be characterized by a first contrast Cont1. The OS can decide to respect this, and respect this compared to the tertiary brightness range of the total composition.
  • In order to do this processing, it may e.g. establish a characteristic brightness level CHRbriLev_COMP of the brightest regions of the composited canvas, which will contain the flames. E.g. areas or objects can be extracted, and an average output luminance can be determined. This may be the brightness the ultra-bright text of the third asset (i.e. Im3) has to compete with. Ergo, the OS can decide to derive one or more of a second contrast Cont2 of a to determine second common white level WcomL2 to the characteristic brightness level CHRbriLev_COMP, and a third contrast Cont3 of that second common white level WcomL2 to the first common white level WcomL1. E.g. WcomL2 can be lowered below CHRbriLev_COMP the smaller Cont1 is (and e.g. in a linear or non-linear proportion of CHRbriLev_COMP compared to WcomL1), etc. Alternatively, other embodiments can ignore the maximum areas of the video, and merely raise the amount of Cont3 as a linear or non-linear function of how much Cont1 is above WL3.
  • Fig. 9 shows generically an example of an algorithm that the OS can use to map brightnesses to a different brightness range (e.g. 100 nit luminances that were to be reconstructed to 2000 nit luminances, will actually be re-graded to a composite range ending at maximum 1000 nit) using an essentially shape preserving tertiary mapping function TM1_opt. The principle is elucidated with a function that is already composed of three parts, but the OS can segment the function in parts that behave essentially similarly in their mapping, e.g. relative brightening (partitioning algorithms are known, e.g. based on extent of deviation from a common joining line, but the change of derivative may also be a good candidate for partitioning; note that the portioning is not necessary in all embodiments, but will be used if some partitions converge or diverge more strongly from the diagonal than others).
  • E.g. for all the darkest lumas, up to first function point Pf1, we can take at least one distance of the input function (TM1) to be shape-adapted, to the diagonal. This initial distance d1 can be calculated by establishing a direction of projection (the OS can use a fixed direction, e.g. 80 degrees i.e. 10 degrees more slanted than vertical down-projection). A final distance is determined, e.g. 80% of the initial distance for this bottom part of the curve (this ratio will depend, usually in a non-linear manner, corresponding with visual appearance, on the difference between the maximum luminance the original function TM1 was intended for, e.g. 2000 nit, and the current situation maximum, of the composite canvas; e.g. 2000/1000 is 2x more, but visually a factor 2 is not a large amount, so one need not deviate the distance d2 by a factor 2, but can keep it close to the initial distance). This corresponds to optimal mapping first point Pc1. All intermediate points can be established correspondingly, meaning, if it is a line then scaled, and if the curve has non-linear shape, the corresponding differing respective distances can all be similarly scaled to 80%. For the middle part up to optimal mapping second point Pc2 respectively second function point Pf2, the OS could elect not to scale to 80%, but e.g. to 90%. The third function point Pf3 lies below the diagonal, meaning the output range should not be used in its entirety (because for this image, or these images, or this asset, 2000 nit is too much). The OS could in principle deviate to an optimal mapping third point Pc3 which lies deeper than Pf3, but then the re-optimized curve is only partially shape preserving. In general, it will scale upwards, since for a lower output maximum luminance (1000 nit) one does not want to make the brightest pixels too dim. These procedures together form the optimal brightness mapping curve TM1_opt for optimizing the lumas of the asset for the different maximum luminance situation of the tertiary brightness range of the total composite canvas. As can be seen, often this function will lie everywhere closer to the diagonal, but is still essentially shape preserving, meaning at least the variations of mapping over the input range, not of course the exact output values for any input (e.g. the middle part is still essentially a large contrast part, because usually there is some object of particular interest there, and the brightness of the brightest object pixels is still essentially kept under moderation). Note that any variants of operating system mapping behavior are not quintessential to the technical contribution of the present application, and merely introduced to elucidate some possible uses of the innovative computer systems or its components. So one example of application of the two color transform control data elements, is that all assets that are to be juxtaposed will be mapped to a common tertiary brightness range by using the reference white level to map such level to one or more final white levels, with an equi-luminance ratio (i.e. 60% brightness of that white in the original asset becomes 60% of the chosen final white level, or at least close to that value) for the darker colors, and the (typically HDR effect) brighter colors are mapped by a final brightness mapping function, which distributes the remaining lumas in each asset above its white reference level over the remaining colors in the tertiary brightness range for output to the display, and in a manner which tries to follow the shape of the communicated brightness mapping function for the asset. If the asset's maximum luminance would fall above the maximum luminance of the tertiary range chosen by the OS, the function can be scaled so that the maximum of such a brighter asset maps to the final maximum brightness of the tertiary range. But other manners of processing guided by the color transform control data are also possible.
  • Fig. 10 gives some further detail on how the geometric graphical presentation of visual assets typically happens internally. A human user 1000 typically interacts with some user app/application (1001). Although some apps could be talking with deeper levels more directly, typically they may be using a more generic graphical user interface language 1002 (a.k.a. shell), like e.g. KDE Plasma or GNOME (Gnu Network Object Model Environment). This will contain a graphics protocol GRPROT, i.e. a set of API calls to interact with the operating system 1010. An example of a Linux graphics protocol is Wayland.
  • The operating system will typically contain (at least) three parts. Besides the kernel 1005, it may typically a window manager 1004 and a so-called display manager 1005, which managers can talk to each other (e.g. the window manager can call functionality of the display manager). The window manager will take care of the user interaction (e.g. mouse focus), and the state of the windows (position, size, transparency and z-order, window shadow). An example of a window manager for ChromeOS is Ash. An example of a dynamic window manager for the X window system is Awesome. The display server takes care of the low level drawing capabilities, e.g. it can draw a line or an area. So the above described luminance mapping processing may typically be performed by the display manager (or by some capability on request of that display manager, e.g. in a calculation circuit of a GPU). The API calls of the protocol will according to the present innovation communicate via the basic pixel color array representation and one or more of the luminance mapping function (which can be formulated as a luma mapping function) and the reference white level, so this will typically be communicated over the graphics protocol GRPROT, but in complex operating system behavior it may also communicate in between modules of the OS, and even back to and back from another application etc.
  • MacOS and iOS can use a display server like e.g. the Quartz display server (they can then use a somewhat different window manager). Android can use SurfaceFlinger.
  • The operating system can talk with the hardware (the GPU 1006 and its connection to a display 1007, which may also be bidirectional in case properties of the display need to be polled like its maximum displayable luminance, for optimizing the tertiary range of brightnesses, i.e. of relative brightnesses or absolute luminances), via a GPU API, e.g. use the more generic APIs like e.g. Vulkan (which is an open standard cross-platform API for 3D graphics and computing), or OpenGL, etc. For request of specific luminance versions of parts of assets also the protocol of communication with the GPU can use the present data formulation, but that is another concept that the communication by apps to, and coordination of luminance distributions of assets by, the OS.
  • The present innovative communication of HDR asset and its luminance re-grading desiderata metadata may also be incorporated into generic OS display server/window server abstraction languages.
  • The algorithmic components disclosed in this text may (entirely or in part) be realized in practice as hardware (e.g. parts of an application specific integrated circuit) or as software running on a special digital signal processor, or a generic processor, etc. At least some of the elements of the various embodiments may be running on a fixed or configurable CPU, GPU, Digital Signal Processor, FPGA, Neural Processing Unit, Application Specific Integrated Circuit, microcontroller, SoC, etc. The images may be temporarily or for long term stored in various memories, in the vicinity of the processor(s) or remotely accessible e.g. over the internet.
  • It should be understandable to the skilled person from our presentation which components may be optional improvements and can be realized in combination with other components, and how (optional) steps of methods correspond to respective means of apparatuses, and vice versa. Some combinations will be taught by splitting the general teachings to partial teachings regarding one or more of the parts. The word "apparatus" in this application is used in its broadest sense, namely a group of means allowing the realization of a particular objective, and can hence e.g. be (a small circuit part of) an IC, or a dedicated appliance (such as an appliance with a display), or part of a networked system, etc. "Arrangement" is also intended to be used in the broadest sense, so it may comprise inter alia a single apparatus, a part of an apparatus, a collection of (parts of) cooperating apparatuses, etc.
  • The computer program product denotation should be understood to encompass any physical realization of a collection of commands enabling a generic or special purpose processor, after a series of loading steps (which may include intermediate conversion steps, such as translation to an intermediate language, and a final processor language) to enter the commands into the processor, and to execute any of the characteristic functions of an invention. In particular, the computer program product may be realized as data on a carrier such as e.g. a disk, data present in a memory, data travelling via a network connection -wired or wireless-. Apart from program code, characteristic data required for the program may also be embodied as a computer program product. Some of the technologies may be encompassed in signals, typically control signals for controlling one or more technical behaviors of e.g. a receiving apparatus, such as a television. Some circuits may be reconfigurable, and temporarily configured for particular processing by software. Some parts of the apparatuses may be specifically adapted to receive, parse and/or understand innovative signals.
  • Some of the steps required for the operation of the method may be already present in the functionality of the processor instead of described in the computer program product, such as data input and output steps.
  • It should be noted that the above-mentioned embodiments illustrate rather than limit the invention. Where the skilled person can easily realize a mapping of the presented examples to other regions of the claims, we have for conciseness not mentioned all these options in-depth. Apart from combinations of elements of the invention as combined in the claims, other combinations of the elements are possible. Any combination of elements can in practice be realized in a single dedicated element, or split elements.
  • Any reference sign between parentheses in the claim is not intended for limiting the claim. The word "comprising" does not exclude the presence of elements or aspects not listed in a claim. In several situations the word "portion" of a set of elements is not intended to exclude that portion may also cover the totality of the elements, because that may function equally in a same manner. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements, nor the presence of other elements. "And/or" means that both options may be present together, or one of them may be present alone. The word "e.g." is typically used to indicate that we mean that something else is also belonging to the possibilities, e.g. a similar element, example, or teaching. "i.a." means inter alia, or among others. An element between ellipses will normally be used to indicate that something is optional, i.e. also possible as a variant of a more general concept, rather than necessary, e.g. (local) luminance boosting is intended to say, (primarily, as main level teaching) "luminance boosting" in general, which may be for all pixels the same, but may also be different, i.e. of the "local luminance boosting" variant, e.g. only applied to some locality of the image.

Claims (15)

  1. A computer (700) arranged to run an operating system (504), wherein the operating system is arranged to manage color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Ass1) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness (ML_V), wherein the visual asset (Ass1) is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness;
    wherein the visual asset in addition comprises color transformation control data (CTctrl) which comprises at least one of first brightness mapping function (TM1) or its inverse, and a reference white level (WL1);
    wherein the operating system comprises a brightness mapping unit (710) arranged to receive the visual asset and to apply a tertiary brightness mapping function (FL_op) to the input luma code to obtain an output luma code (Y'o), wherein the output luma code specifies a tertiary brightness which lies in a tertiary brightness range which is different from the secondary brightness range, wherein the output luma codes a brightness of an output color, wherein the tertiary brightness mapping function (FL_op) is based on at least one of the first brightness mapping function (TM1) and the reference white level (WL1);
    wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
  2. The computer as claimed in claim 1, wherein the visual asset is received from a software application by color codes represented in an sRGB color representation.
  3. The computer as claimed in claim 1 or 2, wherein the original brightness specifies a luminance of a pixel to be displayed as an amount of nits.
  4. An operating system (504) for being supplied to and for operation on a computer, wherein the operating system is arranged to manage color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Ass1) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness (ML_V), wherein the visual asset (Ass1) is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness;
    wherein the visual asset in addition comprises color transformation control data (CTctrl) which comprises at least one of first brightness mapping function (TM1) or its inverse, and a reference white level (WL1);
    wherein the operating system comprises a brightness mapping unit (710) arranged to receive the visual asset and to apply a tertiary brightness mapping function (FL_op) to the input luma code to obtain an output luma code (Y'o), wherein the output luma code specifies a tertiary brightness which lies in a tertiary brightness range which is different from the secondary brightness range, wherein the output luma codes a brightness of an output color, wherein the tertiary brightness mapping function (FL_op) is based on at least one of the first brightness mapping function (TM1) and the reference white level (WL1);
    wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
  5. The operating system as claimed in claim 4, wherein the visual asset is received from a software application by color codes represented in an sRGB color representation.
  6. The operating system as claimed in claim 4, wherein the original brightness specifies a luminance of a pixel to be displayed as an amount of nits.
  7. A tangible data source, comprising code enabling execution on a computer of the operating system as claimed in claim 4.
  8. A method of operating a computer, comprising a step of running an operating system, wherein the operating system is arranged to manage color coordination of visual assets received from one or more software applications (701, 702), wherein the operating system comprises a geometrical management unit (711) arranged to receive at least one visual asset (Ass1) comprising a pixel color having an original brightness lying in a first brightness range which has a first maximum brightness (ML_V), wherein the visual asset (Ass1) is received comprising a set of pixel color codes, wherein a color code represents an input luma code that defines a secondary brightness, which is derived from the original brightness by applying a first brightness mapping function (TM1) to the original brightness, which yields the secondary brightness as output of the function, wherein the secondary brightness lies in a second brightness range which has a second maximum brightness which is different from the first maximum brightness;
    wherein the visual asset in addition comprises color transformation control data (CTctrl) which comprises at least one of first brightness mapping function (TM1) or its inverse, and a reference white level (WL1);
    wherein the operating system comprises a brightness mapping unit (710) arranged to receive the visual asset and to apply a tertiary brightness mapping function (FL_op) to the input luma code to obtain an output luma code (Y'o), wherein the output luma code specifies a tertiary brightness which lies in a tertiary brightness range which is different from the secondary brightness range, wherein the output luma codes a brightness of an output color, wherein the tertiary brightness mapping function (FL_op) is based on at least one of the first brightness mapping function (TM1) and the reference white level (WL1);
    wherein the geometrical management unit (711) is arranged to establish an output pixelated image format (ImFinFmt) and to write the output color in a portion of a scanout buffer (511) of a graphics processing unit (510) according to the output pixelated image format.
  9. The method of operating a computer as claimed in claim 8, wherein the visual asset is received from a software application by color codes represented in an sRGB color representation.
  10. The method of operating a computer as claimed in claim 8, wherein the original brightness specifies a luminance of a pixel to be displayed as an amount of nits.
  11. A software application (502, or 503) arranged to create at least one visual asset (Ass1) comprising pixels wherein a pixel has an original color code which specifies an original brightness and to communicate such visual asset to an operating system (504), wherein the original brightness lies in a first brightness range which has a first maximum brightness (ML_V);
    wherein the original color code is transformed into a secondary color code for communication to the operating system, wherein the secondary color code represents a second luma code, which codes a secondary brightness which lies in a secondary brightness range which is different from the first brightness range, wherein the secondary brightness is derived from the original brightness based on application of a brightness mapping function (TM1) to the original brightness;
    characterized in that the software application communicates color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Ass1), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1).
  12. The software application (502, or 503) as claimed in claim 11, arranged to communicate the asset to the operating system in an sRGB representation.
  13. A method of communicating a visual asset having pixel colors to an operating system, comprising the steps of:
    creating at least one visual asset (Ass1) comprising pixels wherein a pixel has an original color code which specifies an original brightness, wherein the original brightness lies in a first brightness range which has a first maximum brightness (ML_V);
    transforming the original color code into a secondary color code for communication to the operating system, wherein the secondary color code represents a second luma code, which codes a secondary brightness which lies in a secondary brightness range which is different from the first brightness range, wherein the secondary brightness is derived from the original brightness based on application of a brightness mapping function (TM1) to the original brightness;
    communicating the secondary color code and color transformation control data (CTctrl) to the operating system, which color transformation control data (CTctrl) is associated with the asset (Ass1), wherein the color transformation control data (CTctrl) comprises at least one or both of the first brightness mapping function (TM1) or its inverse, and a reference white level (WL1).
  14. The method of communicating a visual asset to an operating system, wherein the secondary color code is an sRGB color code.
  15. A computer program product comprising code enabling a computer to execute the software application (502, or 503) as claimed in claim 11.
EP24157639.6A 2024-02-14 2024-02-14 Brightness range adaptation for computers Pending EP4604114A1 (en)

Priority Applications (2)

Application Number Priority Date Filing Date Title
EP24157639.6A EP4604114A1 (en) 2024-02-14 2024-02-14 Brightness range adaptation for computers
PCT/EP2025/053542 WO2025172273A1 (en) 2024-02-14 2025-02-11 Brightness range adaptation for computers

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
EP24157639.6A EP4604114A1 (en) 2024-02-14 2024-02-14 Brightness range adaptation for computers

Publications (1)

Publication Number Publication Date
EP4604114A1 true EP4604114A1 (en) 2025-08-20

Family

ID=89940847

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24157639.6A Pending EP4604114A1 (en) 2024-02-14 2024-02-14 Brightness range adaptation for computers

Country Status (2)

Country Link
EP (1) EP4604114A1 (en)
WO (1) WO2025172273A1 (en)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2017108906A1 (en) 2015-12-21 2017-06-29 Koninklijke Philips N.V. Optimizing high dynamic range images for particular displays
WO2017157977A1 (en) 2016-03-18 2017-09-21 Koninklijke Philips N.V. Encoding and decoding hdr videos
US20170347113A1 (en) * 2015-01-29 2017-11-30 Koninklijke Philips N.V. Local dynamic range adjustment color processing
US20210152801A1 (en) * 2018-07-05 2021-05-20 Huawei Technologies Co., Ltd. Video Signal Processing Method and Apparatus

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105009567B (en) * 2013-02-21 2018-06-08 杜比实验室特许公司 For synthesizing the system and method for the appearance of superposed graph mapping
WO2017089146A1 (en) * 2015-11-24 2017-06-01 Koninklijke Philips N.V. Handling multiple hdr image sources

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170347113A1 (en) * 2015-01-29 2017-11-30 Koninklijke Philips N.V. Local dynamic range adjustment color processing
WO2017108906A1 (en) 2015-12-21 2017-06-29 Koninklijke Philips N.V. Optimizing high dynamic range images for particular displays
WO2017157977A1 (en) 2016-03-18 2017-09-21 Koninklijke Philips N.V. Encoding and decoding hdr videos
US20210152801A1 (en) * 2018-07-05 2021-05-20 Huawei Technologies Co., Ltd. Video Signal Processing Method and Apparatus

Also Published As

Publication number Publication date
WO2025172273A1 (en) 2025-08-21

Similar Documents

Publication Publication Date Title
JP7343629B2 (en) Method and apparatus for encoding HDR images
US11521537B2 (en) Optimized decoded high dynamic range image saturation
US10878776B2 (en) Optimizing high dynamic range images for particular displays
EP3381179B1 (en) Handling multiple hdr image sources
CN107111980A (en) Optimize high dynamic range images for particular display
JP2024517241A (en) Display-optimized HDR video contrast adaptation
EP4604114A1 (en) Brightness range adaptation for computers
EP4568235A1 (en) Hdr range adaptation luminance processing on computers
US20240273692A1 (en) Display-Optimized Ambient Light HDR Video Adapation
EP4657423A1 (en) Visual asset state-dependent coordinated luminance processing
EP4607917A1 (en) Improved encoding and decoding for images
EP4636683A1 (en) Improved luma and chroma mapping for images
EP4523416B1 (en) Hdr video reconstruction by converted tone mapping
EP4567783A1 (en) Image display improvement in brightened viewing environments
KR102279842B1 (en) Methods and apparatuses for encoding hdr images
WO2025176561A1 (en) Luminance mapping for images
JP2024519606A (en) Display-optimized HDR video contrast adaptation

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION HAS BEEN PUBLISHED

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR