EP4736416A1 - Improvements to intra block copy - Google Patents

Improvements to intra block copy

Info

Publication number
EP4736416A1
EP4736416A1 EP24732322.3A EP24732322A EP4736416A1 EP 4736416 A1 EP4736416 A1 EP 4736416A1 EP 24732322 A EP24732322 A EP 24732322A EP 4736416 A1 EP4736416 A1 EP 4736416A1
Authority
EP
European Patent Office
Prior art keywords
candidate
block
offset
diagonal
ibc
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24732322.3A
Other languages
German (de)
French (fr)
Inventor
Fabrice Le Leannec
Karam NASER
Antoine Robert
Tangi POIRIER
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
InterDigital CE Patent Holdings SAS
Original Assignee
InterDigital CE Patent Holdings SAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by InterDigital CE Patent Holdings SAS filed Critical InterDigital CE Patent Holdings SAS
Publication of EP4736416A1 publication Critical patent/EP4736416A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/593Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial prediction techniques
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/103Selection of coding mode or of prediction mode
    • H04N19/105Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/176Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/503Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
    • H04N19/51Motion estimation or motion compensation
    • H04N19/56Motion estimation with initialisation of the vector search, e.g. estimating a good candidate to initiate a search

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Mobile Radio Communication Systems (AREA)

Abstract

Disclosed herein are systems, methods and instrumentalities associated with video content coding (e.g., encoding and decoding). A video coding device (e.g., a video encoder or a video decoder) as described herein may be configured to determine a base block vector (BV) predictor associated with a current coding unit (CU) and further determine a block vector difference (BVD) associated with the base BV predictor, wherein the BVD may be associated with a diagonal direction and a magnitude corresponding to the diagonal direction. The video coding device may be further configured to determine a BV associated with the current CU based on the based BV predictor and the BVD, and code (e.g., encode or decode) the current CU based at least on the determined BV.

Description

IMPROVEMENTS TO INTRA BLOCK COPY
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of European Patent Application No. 23306059.9, filed June 29, 2023, the disclosure of which is incorporated herein by reference in its entirety.
BACKGROUND
[0002] Intra block copy (IBC) may be used as a tool in video coding (e.g., encoding and/or decoding). IBC may be improved to improve the efficiency of video coding.
SUMMARY
[0003] Disclosed herein are systems, methods and instrumentalities associated with video content coding (e.g., encoding and decoding). A video coding device (e.g., a video encoder or a video decoder) as described herein may include a processor configured to determine a base block vector (BV) predictor associated with a current coding unit (CU) and further determine a block vector difference (BVD) associated with the base BV predictor, wherein the BVD may be associated with a diagonal direction and a magnitude corresponding to the diagonal direction. The video coding device may be further configured to determine a BV associated with the current CU based on the base BV predictor and the BVD, and code (e.g., encode or decode) the current CU based at least on the determined BV.
[0004] In examples, the diagonal direction described herein may be non-vertical and non-horizontal. For instance, the diagonal direction may be at an angle from the vertical and horizontal directions and the angle may be a multiple of TT/2, TT/4 or TT/8.
[0005] In examples, the video coding device may be configured to code the current CU in an intra block copy (IBC) mode, such as, e.g., an IBC merge with block vector differences (IBC-MBVD) mode.
[0006] A video decoding device as described herein may be configured to obtain a base BV associated with a current video block coded in a block vector-based coding mode. The video decoding device may be further configured to determine a plurality of candidate BV offsets associated with the base BV, wherein the plurality of candidate BV offsets may include a diagonal BV offset (e.g., a BV offset pointing at a non-zero angle from the horizontal or vertical direction). The video decoding device may determine a refined BV based on the base BV and the plurality of candidate BV offsets, wherein the determination may be made at least partially based on the diagonal BV offset. The video decoding device may then decode the current video block based on the refined BV.
[0007] In examples, the diagonal BV offset may point at a 45-degree angle or a 22.5-degree from a horizontal direction or a vertical direction. In examples, the plurality of candidate BV offsets may further include a horizontal BV offset or a vertical BV offset, and the video decoding device may be further configured to determine a first set of candidate magnitudes associated with the diagonal BV offset and a second set of candidate magnitudes associated with the horizontal or vertical BV offset, wherein the first set of candidate magnitudes may include fewer candidates than the second set of candidate magnitudes. In examples, the number of candidates included in the first set of candidate magnitudes may be dependent on an angle of the diagonal BV offset.
[0008] In examples, the current video block may include camera-captured video content and/or be decoded using an intra block copy (IBC) prediction mode such as an IBC merge with block vector differences (IBC-MBVD) prediction mode.
[0009] A video encoding device as described herein may be configured to obtain a base BV associated with a current video block coded in a block vector-based coding mode. The video encoding device may determine a plurality of candidate BV offsets associated with the base BV, wherein the plurality of candidate BV offsets may include a diagonal BV offset (e.g., a BV offset pointing at a non-zero angle such as a 45-degree or a 22.5-degree angle from the horizontal or vertical direction). The video encoding device may determine a refined BV based on the base BV and the plurality of candidate BV offsets, wherein the determination may be made at least partially based on the diagonal BV offset. The video encoding device may then encode the current video block based on the refined BV.
BRIEF DESCRIPTION OF THE DRAWINGS
[0010] FIG. 1 A is a system diagram illustrating an example communications system in which one or more disclosed embodiments can be implemented.
[0011] FIG. 1 B is a system diagram illustrating an example wireless transmit/receive unit (WTRU) that can be used within the communications system illustrated in FIG. 1A according to an embodiment.
[0012] FIG. 1 C is a system diagram illustrating an example radio access network (RAN) and an example core network (CN) that can be used within the communications system illustrated in FIG. 1 A according to an embodiment.
[0013] FIG. 1 D is a system diagram illustrating a further example RAN and a further example CN that can be used within the communications system illustrated in FIG. 1 A according to an embodiment.
[0014] FIG. 2 illustrates an example video encoder.
[0015] FIG. 3 illustrates an example video decoder.
[0016] FIG. 4 illustrates an example of a system in which various aspects and examples can be implemented.
[0017] FIG. 5 illustrates examples of positions of spatial and temporal motion vector predictors that may be used in a merge mode. [0018] FIG. 6 illustrates an example of constructing a list of merge motion vector predictor candidates.
[0019] FIG. 7 illustrates another example of constructing a list of merge motion vector predictor candidates.
[0020] FIG. 8 illustrates examples of motion representation categories.
[0021] FIG. 9 illustrates an example of non-sub-block merge candidate list construction.
[0022] FIG. 10 illustrates example of allowed motion vector differences (MVd or MVD) in a merge with motion vector differences (MMVD) mode.
[0023] FIG. 11 illustrates examples of additional block vector difference (BVD) directions (e.g., along kxrr/8 diagonal angles, where the positions marked in black such as 1702 may be used in an anchor).
[0024] FIG. 12 illustrates an example of a processing order for a current coding tree unit (CTU) and examples of available reference samples (e.g., in the current CTU and a left CTU).
[0025] FIG. 13 illustrates examples of padding candidates that may be used for the replacement of zerovectors in an IBC candidate list.
[0026] FIG. 14 illustrates an example of an extended reference region for IBC.
[0027] FIG. 15 illustrates examples of improvements to a merge with block vector differences (MBVD) mode.
[0028] FIG. 16 illustrates an example of a decoding process for a CU in an IBC MBVD mode (also referred to as an IBC-MBVD mode).
[0029] FIG. 17 illustrates examples of additional BVD directions (e.g., along kxrr/8 diagonal angles) that may be used in an MBVD mode.
[0030] FIG. 18 illustrates an example of an extended set of BVD candidates.
[0031] FIG. 19 illustrates another example of an extended set of BVD candidates.
DETAILED DESCRIPTION
[0032] A more detailed understanding can be had from the following description, given by way of example in conjunction with the accompanying drawings.
[0033] FIG. 1 A is a diagram illustrating an example communications system 100 in which one or more disclosed embodiments can be implemented. The communications system 100 can be a multiple access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communications system 100 can enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communications systems 100 can employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block-filtered OFDM, filter bank multicarrier (FBMC), and the like.
[0034] As shown in FIG. 1 A, the communications system 100 can include wireless transmit/receive units (WTRUs) 102a, 102b, 102c, 102d, a RAN 104/113, a ON 106/115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, though it will be appreciated that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and/or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to operate and/or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d, any of which can be referred to as a "station” and/or a "STA”, can be configured to transmit and/or receive wireless signals and can include a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular telephone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fl device, an Internet of Things (loT) device, a watch or other wearable, a head-mounted display (HMD), a vehicle, a drone, a medical device and applications (e.g., remote surgery), an industrial device and applications (e.g., a robot and/or other wireless devices operating in an industrial and/or an automated processing chain contexts), a consumer electronics device, a device operating on commercial and/or industrial wireless networks, and the like. Any of the WTRUs 102a, 102b, 102c and 102d can be interchangeably referred to as a UE.
[0035] The communications systems 100 can also include a base station 114a and/or a base station 114b. Each of the base stations 114a, 114b can be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as the CN 106/115, the Internet 110, and/or the other networks 112. By way of example, the base stations 114a, 114b can be a base transceiver station (BTS), a Node-B, an eNode B, a Home Node B, a Home eNode B, a gNB, a NR NodeB, a site controller, an access point (AP), a wireless router, and the like. While the base stations 114a, 114b are each depicted as a single element, it will be appreciated that the base stations 114a, 114b can include any number of interconnected base stations and/or network elements.
[0036] The base station 114a can be part of the RAN 104/113, which can also include other base stations and/or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and/or the base station 114b can be configured to transmit and/or receive wireless signals on one or more carrier frequencies, which can be referred to as a cell (not shown). These frequencies can be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell can provide coverage for a wireless service to a specific geographical area that can be relatively fixed or that can change over time. The cell can further be divided into cell sectors. For example, the cell associated with the base station 114a can be divided into three sectors. Thus, in one embodiment, the base station 114a can include three transceivers, i.e., one for each sector of the cell. In an embodiment, the base station 114a can employ multiple-input multiple output (MIMO) technology and can utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and/or receive signals in desired spatial directions.
[0037] The base stations 114a, 114b can communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 can be established using any suitable radio access technology (RAT).
[0038] More specifically, as noted above, the communications system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and the like. For example, the base station 114a in the RAN 104/113 and the WTRUs 102a, 102b, 102c can implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can establish the air interface 115/116/117 using wideband CDMA (WCDMA). WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and/or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and/or High-Speed UL Packet Access (HSUPA).
[0039] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can establish the air interface 116 using Long Term Evolution (LTE) and/or LTE-Advanced (LTE-A) and/or LTE-Advanced Pro (LTE-A Pro).
[0040] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as NR Radio Access, which can establish the air interface 116 using New Radio (NR).
[0041] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c can implement LTE radio access and NR radio access together, for instance using dual connectivity (DC) principles. Thus, the air interface utilized by WTRUs 102a, 102b, 102c can be characterized by multiple types of radio access technologies and/or transmissions sent to/from multiple types of base stations (e.g., a eNB and a gNB).
[0042] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), and the like.
[0043] The base station 114b in FIG. 1 A can be a wireless router, Home Node B, Home eNode B, or access point, for example, and can utilize any suitable RAT for facilitating wireless connectivity in a localized area, such as a place of business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a roadway, and the like. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR etc.) to establish a picocell or femtocell. As shown in FIG. 1A, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b can not be required to access the Internet 110 via the CN 106/115.
[0044] The RAN 104/113 can be in communication with the CN 106/115, which can be any type of network configured to provide voice, data, applications, and/or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data can have varying quality of service (QoS) requirements, such as differing throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, and the like. The CN 106/115 can provide call control, billing services, mobile location-based services, pre-paid calling, Internet connectivity, video distribution, etc., and/or perform high-level security functions, such as user authentication. Although not shown in FIG. 1A, it will be appreciated that the RAN 104/113 and/or the CN 106/115 can be in direct or indirect communication with other RANs that employ the same RAT as the RAN 104/113 or a different RAT. For example, in addition to being connected to the RAN 104/113, which can be utilizing a NR radio technology, the CN 106/115 can also be in communication with another RAN (not shown) employing a GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0045] The CN 106/115 can also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and/or the other networks 112. The PSTN 108 can include circuit-switched telephone networks that provide plain old telephone service (POTS). The Internet 110 can include a global system of interconnected computer networks and devices that use common communication protocols, such as the transmission control protocol (TCP), user datagram protocol (UDP) and/or the internet protocol (IP) in the TCP/IP internet protocol suite. The networks 112 can include wired and/or wireless communications networks owned and/or operated by other service providers. For example, the networks 112 can include another CN connected to one or more RANs, which can employ the same RAT as the RAN 104/113 or a different RAT.
[0046] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 can include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d can include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU 102c shown in FIG. 1A can be configured to communicate with the base station 114a, which can employ a cellular-based radio technology, and with the base station 114b, which can employ an IEEE 802 radio technology.
[0047] FIG. 1 B is a system diagram illustrating an example WTRU 102. As shown in FIG. 1 B, the WTRU 102 can include a processor 118, a transceiver 120, a transmit/receive element 122, a speaker/microphone 124, a keypad 126, a display/touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and/or other peripherals 138, among others. It will be appreciated that the WTRU 102 can include any sub-combination of the foregoing elements while remaining consistent with an embodiment.
[0048] The processor 118 can be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs) circuits, any other type of integrated circuit (IC), a state machine, and the like. The processor 118 can perform signal coding, data processing, power control, input/output processing, and/or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to the transceiver 120, which can be coupled to the transmit/receive element 122. While FIG. 1 B depicts the processor 118 and the transceiver 120 as separate components, it will be appreciated that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0049] The transmit/receive element 122 can be configured to transmit signals to, or receive signals from, a base station (e.g., the base station 114a) over the air interface 116. For example, in one embodiment, the transmit/receive element 122 can be an antenna configured to transmit and/or receive RF signals. In an embodiment, the transmit/receive element 122 can be an emitter/detector configured to transmit and/or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit/receive element 122 can be configured to transmit and/or receive both RF and light signals. It will be appreciated that the transmit/receive element 122 can be configured to transmit and/or receive any combination of wireless signals. [0050] Although the transmit/receive element 122 is depicted in FIG. 1 B as a single element, the WTRU 102 can include any number of transmit/receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit/receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0051] The transceiver 120 can be configured to modulate the signals that are to be transmitted by the transmit/receive element 122 and to demodulate the signals that are received by the transmit/receive element 122. As noted above, the WTRU 102 can have multi-mode capabilities. Thus, the transceiver 120 can include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11 , for example.
[0052] The processor 118 of the WTRU 102 can be coupled to, and can receive user input data from, the speaker/microphone 124, the keypad 126, and/or the display/touchpad 128 (e.g., a liquid crystal display (LCD) display unit or organic light-emitting diode (OLED) display unit). The processor 118 can also output user data to the speaker/microphone 124, the keypad 126, and/or the display/touchpad 128. In addition, the processor 118 can access information from, and store data in, any type of suitable memory, such as the non-removable memory 130 and/or the removable memory 132. The non-removable memory 130 can include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 can include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 can access information from, and store data in, memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0053] The processor 118 can receive power from the power source 134, and can be configured to distribute and/or control the power to the other components in the WTRU 102. The power source 134 can be any suitable device for powering the WTRU 102. For example, the power source 134 can include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, and the like.
[0054] The processor 118 can also be coupled to the GPS chipset 136, which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or in lieu of, the information from the GPS chipset 136, the WTRU 102 can receive location information over the air interface 116 from a base station (e.g., base stations 114a, 114b) and/or determine its location based on the timing of the signals being received from two or more nearby base stations. It will be appreciated that the WTRU 102 can acquire location information by way of any suitable locationdetermination method while remaining consistent with an embodiment. [0055] The processor 118 can further be coupled to other peripherals 138, which can include one or more software and/or hardware modules that provide additional features, functionality and/or wired or wireless connectivity. For example, the peripherals 138 can include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photographs and/or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands free headset, a Bluetooth® module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a Virtual Reality and/or Augmented Reality (VR/AR) device, an activity tracker, and the like. The peripherals 138 can include one or more sensors, the sensors can be one or more of a gyroscope, an accelerometer, a hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor; a geolocation sensor; an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and/or a humidity sensor.
[0056] The WTRU 102 can include a full duplex radio for which transmission and reception of some or all of the signals (e.g., associated with particular subframes for both the UL (e.g., for transmission) and downlink (e.g., for reception) can be concurrent and/or simultaneous. The full duplex radio can include an interference management unit to reduce and or substantially eliminate self-interference via either hardware (e.g., a choke) or signal processing via a processor (e.g., a separate processor (not shown) or via processor 118). In an embodiment, the WTRU 102 can include a half-duplex radio for which transmission and reception of some or all of the signals (e.g., associated with particular subframes for either the UL (e.g., for transmission) or the downlink (e.g., for reception)).
[0057] FIG. 1 C is a system diagram illustrating the RAN 104 and the CN 106 according to an embodiment. As noted above, the RAN 104 can employ an E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 104 can also be in communication with the CN 106.
[0058] The RAN 104 can include eNode-Bs 160a, 160b, 160c, though it will be appreciated that the RAN 104 can include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 160a, 160b, 160c can each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the eNode-Bs 160a, 160b, 160c can implement MIMO technology. Thus, the eNode-B 160a, for example, can use multiple antennas to transmit wireless signals to, and/or receive wireless signals from, the WTRU 102a.
[0059] Each of the eNode-Bs 160a, 160b, 160c can be associated with a particular cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and/or DL, and the like. As shown in FIG. 1 C, the eNode-Bs 160a, 160b, 160c can communicate with one another over an X2 interface. [0060] The CN 106 shown in FIG. 1 C can include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. While each of the foregoing elements are depicted as part of the CN 106, it will be appreciated that any of these elements can be owned and/or operated by an entity other than the CN operator.
[0061] The MME 162 can be connected to each of the eNode-Bs 162a, 162b, 162c in the RAN 104 via an S1 interface and can serve as a control node. For example, the MME 162 can be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation/deactivation, selecting a particular serving gateway during an initial attach of the WTRUs 102a, 102b, 102c, and the like. The MME 162 can provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies, such as GSM and/or WCDMA.
[0062] The SGW 164 can be connected to each of the eNode Bs 160a, 160b, 160c in the RAN 104 via the S1 interface. The SGW 164 can generally route and forward user data packets to/from the WTRUs 102a, 102b, 102c. The SGW 164 can perform other functions, such as anchoring user planes during inter- eNode B handovers, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, managing and storing contexts of the WTRUs 102a, 102b, 102c, and the like.
[0063] The SGW 164 can be connected to the PGW 166, which can provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0064] The CN 106 can facilitate communications with other networks. For example, the CN 106 can provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communications between the WTRUs 102a, 102b, 102c and traditional land-line communications devices. For example, the CN 106 can include, or can communicate with, an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. In addition, the CN 106 can provide the WTRUs 102a, 102b, 102c with access to the other networks 112, which can include other wired and/or wireless networks that are owned and/or operated by other service providers.
[0065] Although the WTRU is described in FIGS. 1 A-1 D as a wireless terminal, it is contemplated that in certain representative embodiments that such a terminal can use (e.g., temporarily or permanently) wired communication interfaces with the communication network.
[0066] In representative embodiments, the other network 112 can be a WLAN.
[0067] A WLAN in Infrastructure Basic Service Set (BSS) mode can have an Access Point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have an access or an interface to a Distribution System (DS) or another type of wired/wireless network that carries traffic in to and/or out of the BSS. Traffic to STAs that originates from outside the BSS can arrive through the AP and can be delivered to the STAs. Traffic originating from STAs to destinations outside the BSS can be sent to the AP to be delivered to respective destinations. Traffic between STAs within the BSS can be sent through the AP, for example, where the source STA can send traffic to the AP and the AP can deliver the traffic to the destination STA. The traffic between STAs within a BSS can be considered and/or referred to as peer-to- peer traffic. The peer-to-peer traffic can be sent between (e.g., directly between) the source and destination STAs with a direct link setup (DLS). In certain representative embodiments, the DLS can use an 802.11e DLS or an 802.11z tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode cannot have an AP, and the STAs (e.g., all of the STAs) within or using the IBSS can communicate directly with each other. The IBSS mode of communication can sometimes be referred to herein as an "ad-hoc” mode of communication.
[0068] When using the 802.11ac infrastructure mode of operation or a similar mode of operations, the AP can transmit a beacon on a fixed channel, such as a primary channel. The primary channel can be a fixed width (e.g., 20 MHz wide bandwidth) or a dynamically set width via signaling. The primary channel can be the operating channel of the BSS and can be used by the STAs to establish a connection with the AP. In certain representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA/CA) can be implemented, for example in in 802.11 systems. For CSMA/CA, the STAs (e.g., every STA), including the AP, can sense the primary channel. If the primary channel is sensed/detected and/or determined to be busy by a particular STA, the particular STA can back off. One STA (e.g., only one station) can transmit at any given time in a given BSS.
[0069] High Throughput (HT) STAs can use a 40 MHz wide channel for communication, for example, via a combination of the primary 20 MHz channel with an adjacent or nonadjacent 20 MHz channel to form a 40 MHz wide channel.
[0070] Very High Throughput (VHT) STAs can support 20MHz, 40 MHz, 80 MHz, and/or 160 MHz wide channels. The 40 MHz, and/or 80 MHz, channels can be formed by combining contiguous 20 MHz channels. A 160 MHz channel can be formed by combining 8 contiguous 20 MHz channels, or by combining two non-contiguous 80 MHz channels, which can be referred to as an 80+80 configuration. For the 80+80 configuration, the data, after channel encoding, can be passed through a segment parser that can divide the data into two streams. Inverse Fast Fourier Transform (IFFT) processing, and time domain processing, can be done on each stream separately. The streams can be mapped on to the two 80 MHz channels, and the data can be transmitted by a transmitting STA. At the receiver of the receiving STA, the above described operation for the 80+80 configuration can be reversed, and the combined data can be sent to the Medium Access Control (MAC). [0071] Sub 1 GHz modes of operation are supported by 802.11 af and 802.11 ah. The channel operating bandwidths, and carriers, are reduced in 802.11 af and 802.11 ah relative to those used in 802.11 n, and 802.11 ac. 802.11 af supports 5 MHz, 10 MHz and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, and 802.11 ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non- TVWS spectrum. According to a representative embodiment, 802.11 ah can support Meter Type Control/Machine-Type Communications, such as MTC devices in a macro coverage area. MTC devices can have certain capabilities, for example, limited capabilities including support for (e.g., only support for) certain and/or limited bandwidths. The MTC devices can include a battery with a battery life above a threshold (e.g., to maintain a very long battery life).
[0072] WLAN systems, which can support multiple channels, and channel bandwidths, such as 802.11 n, 802.11 ac, 802.11 af, and 802.11 ah, include a channel which can be designated as the primary channel.
The primary channel can have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and/or limited by a STA, from among all STAs in operating in a BSS, which supports the smallest bandwidth operating mode. In the example of 802.11 ah, the primary channel can be 1 MHz wide for STAs (e.g., MTC type devices) that support (e.g., only support) a 1 MHz mode, even if the AP, and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and/or other channel bandwidth operating modes. Carrier sensing and/or Network Allocation Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, for example, due to a STA (which supports only a 1 MHz operating mode), transmitting to the AP, the entire available frequency bands can be considered busy even though a majority of the frequency bands remains idle and can be available.
[0073] In the United States, the available frequency bands, which can be used by 802.11 ah, are from 902 MHz to 928 MHz. In Korea, the available frequency bands are from 917.5 MHz to 923.5 MHz. In Japan, the available frequency bands are from 916.5 MHz to 927.5 MHz. The total bandwidth available for 802.11 ah is 6 MHz to 26 MHz depending on the country code.
[0074] FIG. 1 D is a system diagram illustrating the RAN 113 and the CN 115 according to an embodiment. As noted above, the RAN 113 can employ an NR radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 113 can also be in communication with the CN 115.
[0075] The RAN 113 can include gNBs 180a, 180b, 180c, though it will be appreciated that the RAN 113 can include any number of gNBs while remaining consistent with an embodiment. The gNBs 180a, 180b, 180c can each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, 180c can implement MIMO technology. For example, gNBs 180a, 108b can utilize beamforming to transmit signals to and/or receive signals from the gNBs 180a, 180b, 180c. Thus, the gNB 180a, for example, can use multiple antennas to transmit wireless signals to, and/or receive wireless signals from, the WTRU 102a. In an embodiment, the gNBs 180a, 180b, 180c can implement carrier aggregation technology. For example, the gNB 180a can transmit multiple component carriers to the WTRU 102a (not shown). A subset of these component carriers can be on unlicensed spectrum while the remaining component carriers can be on licensed spectrum. In an embodiment, the gNBs 180a, 180b, 180c can implement Coordinated Multi-Point (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNB 180a and gNB 180b (and/or gNB 180c).
[0076] The WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using transmissions associated with a scalable numerology. For example, the OFDM symbol spacing and/or OFDM subcarrier spacing can vary for different transmissions, different cells, and/or different portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using subframe or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing varying number of OFDM symbols and/or lasting varying lengths of absolute time).
[0077] The gNBs 180a, 180b, 180c can be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and/or a non-standalone configuration. In the standalone configuration, WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c without also accessing other RANs (e.g., such as eNode-Bs 160a, 160b, 160c). In the standalone configuration, WTRUs 102a, 102b, 102c can utilize one or more of gNBs 180a, 180b, 180c as a mobility anchor point. In the standalone configuration, WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using signals in an unlicensed band. In a non-standalone configuration WTRUs 102a, 102b, 102c can communicate with/connect to gNBs 180a, 180b, 180c while also communicating with/connecting to another RAN such as eNode-Bs 160a, 160b, 160c. For example, WTRUs 102a, 102b, 102c can implement DC principles to communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c substantially simultaneously. In the non-standalone configuration, eNode-Bs 160a, 160b, 160c can serve as a mobility anchor for WTRUs 102a, 102b, 102c and gNBs 180a, 180b, 180c can provide additional coverage and/or throughput for servicing WTRUs 102a, 102b, 102c.
[0078] Each of the gNBs 180a, 180b, 180c can be associated with a particular cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and/or DL, support of network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards User Plane Function (UPF) 184a, 184b, routing of control plane information towards Access and Mobility Management Function (AMF) 182a, 182b and the like. As shown in FIG. 1 D, the gNBs 180a, 180b, 180c can communicate with one another over an Xn interface.
[0079] The CN 115 shown in FIG. 1 D can include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements are depicted as part of the CN 115, it will be appreciated that any of these elements can be owned and/or operated by an entity other than the CN operator.
[0080] The AMF 182a, 182b can be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and can serve as a control node. For example, the AMF 182a, 182b can be responsible for authenticating users of the WTRUs 102a, 102b, 102c, support for network slicing (e.g., handling of different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, management of the registration area, termination of NAS signaling, mobility management, and the like. Network slicing can be used by the AMF 182a, 182b in order to customize CN support for WTRUs 102a, 102b, 102c based on the types of services being utilized WTRUs 102a, 102b, 102c. For example, different network slices can be established for different use cases such as services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, services for machine type communication (MTC) access, and/or the like. The AMF 162 can provide a control plane function for switching between the RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and/or non-3GPP access technologies such as WiFi.
[0081] The SMF 183a, 183b can be connected to an AMF 182a, 182b in the CN 115 via an N11 interface. The SMF 183a, 183b can also be connected to a UPF 184a, 184b in the CN 115 via an N4 interface. The SMF 183a, 183b can select and control the UPF 184a, 184b and configure the routing of traffic through the UPF 184a, 184b. The SMF 183a, 183b can perform other functions, such as managing and allocating UE IP address, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notifications, and the like. A PDU session type can be IP-based, non-IP based, Ethernetbased, and the like.
[0082] The UPF 184a, 184b can be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface, which can provide the WTRUs 102a, 102b, 102c with access to packet- switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPF 184, 184b can perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, and the like. [0083] The CN 115 can facilitate communications with other networks. For example, the CN 115 can include, or can communicate with, an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that serves as an interface between the CN 115 and the PSTN 108. In addition, the CN 115 can provide the WTRUs 102a, 102b, 102c with access to the other networks 112, which can include other wired and/or wireless networks that are owned and/or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c can be connected to a local Data Network (DN) 185a, 185b through the UPF 184a, 184b via the N3 interface to the UPF 184a, 184b and an N6 interface between the UPF 184a, 184b and the DN 185a, 185b.
[0084] In view of Figures 1A-1 D, and the corresponding description of Figures 1A-1 D, one or more, or all, of the functions described herein with regard to one or more of: WTRU 102a-d, Base Station 114a-b, eNode-B 160a-c, MME 162, SGW 164, PGW 166, gNB 180a-c, AMF 182a-b, UPF 184a-b, SMF 183a-b, DN 185a-b, and/or any other device(s) described herein, can be performed by one or more emulation devices (not shown). The emulation devices can be one or more devices configured to emulate one or more, or all, of the functions described herein. For example, the emulation devices can be used to test other devices and/or to simulate network and/or WTRU functions.
[0085] The emulation devices can be designed to implement one or more tests of other devices in a lab environment and/or in an operator network environment. For example, the one or more emulation devices can perform the one or more, or all, functions while being fully or partially implemented and/or deployed as part of a wired and/or wireless communication network in order to test other devices within the communication network. The one or more emulation devices can perform the one or more, or all, functions while being temporarily implemented/deployed as part of a wired and/or wireless communication network. The emulation device can be directly coupled to another device for purposes of testing and/or can performing testing using over-the-air wireless communications.
[0086] The one or more emulation devices can perform the one or more, including all, functions while not being implemented/deployed as part of a wired and/or wireless communication network. For example, the emulation devices can be utilized in a testing scenario in a testing laboratory and/or a non-deployed (e.g., testing) wired and/or wireless communication network in order to implement testing of one or more components. The one or more emulation devices can be test equipment. Direct RF coupling and/or wireless communications via RF circuitry (e.g., which can include one or more antennas) can be used by the emulation devices to transmit and/or receive data.
[0087] This application describes a variety of aspects, including tools, features, examples, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that can sound limiting. However, this is for purposes of clarity in description, and does not limit the application or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well.
[0088] The aspects described and contemplated in this application can be implemented in many different forms. The figures provided herein can provide some examples, but other examples are contemplated. The discussion of the figures does not limit the breadth of the implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable medium (e.g., storage medium) comprising (e.g., having stored thereon) instructions for encoding or decoding video data according to any of the methods described, and/or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described. When referred to herein, a bitstream can refer to transmitted data, but can also refer to data that is stored, generated, and/or accessed without being transmitted (e.g., non-transitory data).
[0089] In the present application, the terms "reconstructed” and "decoded” can be used interchangeably, the terms "pixel” and "sample” can be used interchangeably, the terms "image,” "picture” and "frame” can be used interchangeably.
[0090] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and/or use of specific steps and/or actions can be modified or combined. Additionally, terms such as "first”, "second”, etc. can be used in various examples to modify an element, component, step, operation, etc., such as, for example, a "first decoding” and a "second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and can occur, for example, before, during, or in an overlapping time period with the second decoding.
[0091] Various methods and other aspects described in this application can be used to modify modules, for example, decoding modules, of a video encoder 200 and decoder 300 as shown in FIG. 2 and FIG. 3. Moreover, the subject matter disclosed herein can be applied, for example, to any type, format or version of video coding, whether described in a standard or a recommendation, whether pre-existing or future- developed, and extensions of any such standards and recommendations. Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination. [0092] Various numeric values are used in examples described in the present application. These and other specific values are for purposes of describing examples and the aspects described are not limited to these specific values.
[0093] FIG. 2 is a diagram showing an example video encoder. Variations of example encoder 200 are contemplated, but the encoder 200 is described below for purposes of clarity without describing all expected variations.
[0094] Before being encoded, the video sequence can go through pre-encoding processing 201 , for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of one of the color components). Metadata can be associated with the pre-processing, and attached to the bitstream.
[0095] In the encoder 200, a picture is encoded by the encoder elements as described below. The picture to be encoded is partitioned 202 and processed in units of, for example, coding units (CUs). Each unit is encoded using, for example, either an intra or inter mode. When a unit is encoded in an intra mode, it performs intra prediction 260. In an inter mode, motion estimation 275 and compensation 270 are performed. The encoder decides 205 which one of the intra mode or inter mode to use for encoding the unit, and indicates the intra/inter decision by, for example, a prediction mode flag. Prediction residuals are calculated, for example, by subtracting 210 the predicted block from the original image block.
[0096] The prediction residuals are then transformed 225 and quantized 230. The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded 245 to output a bitstream. The encoder can skip the transform and apply quantization directly to the non-transformed residual signal. The encoder can bypass both transform and quantization, i.e., the residual is coded directly without the application of the transform or quantization processes.
[0097] The encoder decodes an encoded block to provide a reference for further predictions. The quantized transform coefficients are de-quantized 240 and inverse transformed 250 to decode prediction residuals. Combining 255 the decoded prediction residuals and the predicted block, an image block is reconstructed. In-loop filters 265 are applied to the reconstructed picture to perform, for example, deblocking/SAO (Sample Adaptive Offset) filtering to reduce encoding artifacts. The filtered image is stored at a reference picture buffer (280).
[0098] FIG. 3 is a diagram showing an example of a video decoder. In example decoder 300, a bitstream is decoded by the decoder elements as described below. Video decoder 300 generally performs a decoding pass reciprocal to the encoding pass as described in FIG. 2. The encoder 200 also generally performs video decoding as part of encoding video data. [0099] In particular, the input of the decoder includes a video bitstream, which can be generated by video encoder 200. The bitstream is first entropy decoded 330 to obtain transform coefficients, motion vectors, and other coded information. The picture partition information indicates how the picture is partitioned. The decoder can therefore divide 335 the picture according to the decoded picture partitioning information. The transform coefficients are de-quantized 340 and inverse transformed 350 to decode the prediction residuals. Combining 355 the decoded prediction residuals and the predicted block, an image block is reconstructed. The predicted block can be obtained 370 from intra prediction 360 or motion- compensated prediction (i.e., inter prediction) 375. In-loop filters 365 are applied to the reconstructed image. The filtered image is stored at a reference picture buffer 380.
[0100] The decoded picture can further go through post-decoding processing 385, for example, an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping performing the inverse of the remapping process performed in the pre-encoding processing 201. The postdecoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream. In an example, the decoded images (e.g., after application of the in-loop filters 365 and/or after post-decoding processing 385, if post-decoding processing is used) can be sent to a display device for rendering to a user.
[0101] FIG. 4 is a diagram showing an example of a system in which various aspects and examples described herein can be implemented. System 400 can be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this document. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and/or discrete components. For example, in at least one example, the processing and encoder/decoder elements of system 400 are distributed across multiple ICs and/or discrete components. In various examples, the system 400 is communicatively coupled to one or more other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and/or output ports. In various examples, the system 400 is configured to implement one or more of the aspects described in this document.
[0102] The system 400 includes at least one processor 410 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this document. Processor 410 can include embedded memory, input output interface, and various other circuitries as known in the art. The system 400 includes at least one memory 420 (e.g., a volatile memory device, and/or a non-volatile memory device). System 400 includes a storage device 440, which can include non-volatile memory and/or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drive, and/or optical disk drive. The storage device 440 can include an internal storage device, an attached storage device (including detachable and non-detachable storage devices), and/or a network accessible storage device, as non-limiting examples.
[0103] System 400 includes an encoder/decoder module 430 configured, for example, to process data to provide an encoded video or decoded video, and the encoder/decoder module 430 can include its own processor and memory. The encoder/decoder module 430 represents module(s) that can be included in a device to perform the encoding and/or decoding functions. As is known, a device can include one or both of the encoding and decoding modules. Additionally, encoder/decoder module 430 can be implemented as a separate element of system 400 or can be incorporated within processor 410 as a combination of hardware and software as known to those skilled in the art.
[0104] Program code to be loaded onto processor 410 or encoder/decoder 430 to perform the various aspects described in this document can be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. In accordance with various examples, one or more of processor 410, memory 420, storage device 440, and encoder/decoder module 430 can store one or more of various items during the performance of the processes described in this document. Such stored items can include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0105] In some examples, memory inside of the processor 410 and/or the encoder/decoder module 430 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other examples, however, a memory external to the processing device (for example, the processing device can be either the processor 410 or the encoder/decoder module 430) is used for one or more of these functions. The external memory can be the memory 420 and/or the storage device 440, for example, a dynamic volatile memory and/or a non-volatile flash memory. In several examples, an external non-volatile flash memory is used to store the operating system of, for example, a television. In at least one example, a fast external dynamic volatile memory such as a RAM is used as working memory for video encoding and decoding operations.
[0106] The input to the elements of system 400 can be provided through various input devices as indicated in block 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (ill) a Universal Serial Bus (USB) input terminal, and/or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG. 4, include composite video.
[0107] In various examples, the input devices of block 445 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (I) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) downconverting the selected signal, (ill) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain examples, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and/or (vi) demultiplexing to select the desired stream of data packets. The RF portion of various examples includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs various of these functions, including, for example, downconverting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box example, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various examples rearrange the order of the above-described (and other) elements, remove some of these elements, and/or add other elements performing similar or different functions. Adding elements can include inserting elements in between existing elements, such as, for example, inserting amplifiers and an analog-to-digital converter. In various examples, the RF portion includes an antenna.
[0108] The USB and/or HDMI terminals can include respective interface processors for connecting system 400 to other electronic devices across USB and/or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, can be implemented, for example, within a separate input processing IC or within processor 410 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface ICs or within processor 410 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410, and encoder/decoder 430 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.
[0109] Various elements of system 400 can be provided within an integrated housing, Within the integrated housing, the various elements can be interconnected and transmit data therebetween using suitable connection arrangement 425, for example, an internal bus as known in the art, including the Inter- IC (I2C) bus, wiring, and printed circuit boards.
[0110] The system 400 includes communication interface 450 that enables communication with other devices via communication channel 460. The communication interface 450 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 460. The communication interface 450 can include, but is not limited to, a modem or network card and the communication channel 460 can be implemented, for example, within a wired and/or a wireless medium. [0111] Data is streamed, or otherwise provided, to the system 400, in various examples, using a wireless network such as a Wi-Fi network, for example IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal of these examples is received over the communications channel 460 and the communications interface 450 which are adapted for Wi-Fi communications. The communications channel 460 of these examples is typically connected to an access point or router that provides access to external networks including the Internet for allowing streaming applications and other over-the-top communications. Other examples provide streamed data to the system 400 using a set-top box that delivers the data over the HDMI connection of the input block 445. Still other examples provide streamed data to the system 400 using the RF connection of the input block 445. As indicated above, various examples provide data in a non-streaming manner. Additionally, various examples use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth® network. [0112] The system 400 can provide an output signal to various output devices, including a display 475, speakers 485, and other peripheral devices 495. The display 475 of various examples includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and/or a foldable display. The display 475 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other device. The display 475 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 495 include, in various examples, one or more of a stand-alone digital video disc (or digital versatile disc) (DVD, for both terms), a disk player, a stereo system, and/or a lighting system. Various examples use one or more peripheral devices 495 that provide a function based on the output of the system 400. For example, a disk player performs the function of playing the output of the system 400. [0113] In various examples, control signals are communicated between the system 400 and the display 475, speakers 485, or other peripheral devices 495 using signaling such as AV. Link, Consumer Electronics Control (CEC), or other communications protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 400 via dedicated connections through respective interfaces 470, 480, and 490. Alternatively, the output devices can be connected to system 400 using the communications channel 460 via the communications interface 450. The display 475 and speakers 485 can be integrated in a single unit with the other components of system 400 in an electronic device such as, for example, a television. In various examples, the display interface 470 includes a display driver, such as, for example, a timing controller (T Con) chip.
[0114] The display 475 and speakers 485 can alternatively be separate from one or more of the other components, for example, if the RF portion of input 445 is part of a separate set-top box. In various examples in which the display 475 and speakers 485 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0115] The examples can be carried out by computer software implemented by the processor 410 or by hardware, or by a combination of hardware and software. As a non-limiting example, the examples can be implemented by one or more integrated circuits. The memory 420 can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 410 can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
[0116] Various implementations include decoding. "Decoding”, as used in this application, can encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various examples, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various examples, such processes also, or alternatively, include processes performed by a decoder of various implementations described in this application.
[0117] As further examples, in one example "decoding” refers only to entropy decoding, in another example "decoding” refers only to differential decoding, and in another example "decoding” refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0118] Various implementations include encoding. In an analogous way to the above discussion about "decoding”, "encoding” as used in this application can encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream. In various examples, such processes include one or more of the processes typically performed by an encoder, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various examples, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this application.
[0119] As further examples, in one example "encoding” refers only to entropy encoding, in another example "encoding” refers only to differential encoding, and in another example "encoding” refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process” is intended to refer specifically to a subset of operations or generally to the broader encoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0120] Note that syntax elements as used herein are descriptive terms. As such, they do not preclude the use of other syntax element names.
[0121] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method/process.
[0122] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable/personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0123] Reference to "one example” or "an example” or "one implementation” or "an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the example is included in at least one example. Thus, the appearances of the phrase "in one example” or "in an example” or "in one implementation” or "in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same example. [0124] Additionally, this application can refer to "determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory. Obtaining can include receiving, retrieving, constructing, generating, and/or determining.
[0125] Further, this application can refer to "accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0126] Additionally, this application can refer to "receiving” various pieces of information. Receiving is, as with "accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, "receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0127] It is to be appreciated that the use of any of the following "and/or”, and "at least one of, for example, in the cases of “A/B”, "A and/or B” and "at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of "A, B, and/or C” and "at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This can be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
[0128] Also, as used herein, the word "signal” refers to, among other things, indicating something to a corresponding decoder. In this way, in an example the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various examples. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various examples. While the preceding relates to the verb form of the word "signal”, the word "signal” can also be used herein as a noun.
[0129] As will be evident to one of ordinary skill in the art, implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described example. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on, or accessed or received from, a processor-readable medium.
[0130] Many examples are described herein. Features of examples can be provided alone or in any combination, across various claim categories and types. Further, examples can include one or more of the features, devices, or aspects described herein, alone or in any combination, across various claim categories and types. For example, features described herein can be implemented in a bitstream or signal that includes information generated as described herein. The information can allow a decoder to decode a bitstream, the encoder, bitstream, and/or decoder according to any of the embodiments described. For example, features described herein can be implemented by creating and/or transmitting and/or receiving and/or decoding a bitstream or signal. For example, features described herein can be implemented a method, process, apparatus, medium storing instructions (e.g., computer-readable medium), medium storing data, or signal. For example, features described herein can be implemented by a TV, set-top box, cell phone, tablet, or other electronic device that performs decoding. The TV, set-top box, cell phone, tablet, or other electronic device can display (e.g., using a monitor, screen, or other type of display) a resulting image (e.g., an image from residual reconstruction of the video bitstream). The TV, set-top box, cell phone, tablet, or other electronic device can receive a signal including an encoded image and perform decoding.
[0131] Motion information associated with video encoding/decoding (e.g., in a merge mode) may be used in multiple (e.g., two) modes, such as, e.g., a skip mode and a merge mode. In these modes, a syntax element (e.g., a single field) may be signaled (e.g., by an encoder) to enable a decoder to retrieve the motion information of a prediction unit (PU). The syntax element may include a merge index. The merge index may indicate which motion vector predictor (MVP) in a list of merge motion information predictors may be used to derive the motion information of a current PU. In the present disclosure, a list of motion information predictors may be referred to as a merge list, or a merge candidate list, and a candidate motion information predictor may be referred to as a merge candidate.
[0132] Operations in a merge mode may include deriving inter prediction information (which may also be referred to as motion information) of a given prediction unit from a selected motion information predictor candidate. The motion information may include one or more inter prediction parameters of a PU, such as a uni-directional or bi-directional temporal prediction type, a reference picture index within a (e.g., each) reference picture list, and/or one or more motion vectors.
[0133] A merge candidate list may include multiple (e.g., five) merge candidates. The merge candidate list may be constructed on an encoder side and/or a decoder side. Multiple positions (e.g., up to 5 spatial positions) may be considered to retrieve potential candidates. The positions may be examined, for example, according to the following order: 1-left (A1), 2-above (B1), 3-above right (BO), 4-left bottom (AO), and 5-above left (B2), wherein the symbols AO, A1 , BO, B1 , B2 may denote the spatial positions shown in FIG. 5. Spatial candidates, whose associated motion information may be different from each other, may be selected. As shown in FIG. 5, a temporal predictor (e.g., which may be noted as TMVP) may be selected according to the temporal motion information located at position H and "center” in FIG. 5 may refer to a candidate at position H (e.g., if a reference picture is not available). A pruning process may be performed (e.g., as illustrated by FIG. 6) to ensure that a selected set of spatial and/or temporal candidates do not include redundant candidates.
[0134] With at least B-slices, combined candidates (e.g., candidates of different types) may be pushed to a merge list if the list is not full. This operation may include forming a candidate by combining the motion information (e.g., associated with a first reference picture list such as L0) of a candidate present in the merge list, with the motion information (e.g., associated with another reference picture list such as L1) of another candidate present in the merge list. If the merge list is not full (e.g., with fewer than 5 elements), zero motion vectors may be pushed to the back of the merge list until it is full. FIG. 7 illustrates an example process of merge list construction.
[0135] Motion data representations may be enriched. The representations may be divided into multiple (e.g., two) categories, such as, e.g., whole-block-based motion representations and sub-block-based motion representations, as illustrated by FIG. 8. In one or more of these categories (e.g., in each category), multiple (e.g., two) modes may be used to code motion information, such as, e.g., a merge/skip mode and/or an advanced motion vector prediction (AMVP) mode. [0136] A whole-block-based motion representation may be derived by assigning a set of motion information, which may be made of one or multiple (e.g., two) motion vectors and associated reference picture(s), to an inter block. The motion information of the inter block may be represented in the form of a single motion vector for the whole block. A sub-block-based motion representation may be derived by dividing a block into subblocks (e..g., 4x4 or 8x8 luma sample subblocks) and assigning an individual set of motion information to a (e.g., each) subblock.
[0137] A video encoding or decoding device may be configured to use a whole-block-based merge mode, which may also be referred to here as a regular merge mode. Compared to legacy merge modes, the whole-block-based merge mode may adopt the same or a different merge MVP candidate list construction process, and/or may include additional merge coding modes, such as, e.g., MMVD (merge with motion vector difference), GPM (geometric partitioning mode) and/or CUP (combined intra/inter prediction).
[0138] The video encoding or decoding device may be configured to construct a merge MVP candidate list with one or more of the following types of MVP candidates. For example, the video encoding or decoding device may construct the merge MVP candidate list based on one or more spatial candidates (e.g., similar to a legacy construction process except that the two first candidates may be swapped). As another example, the video encoding or decoding device may construct the merge MVP candidate list based on one or more temporal MVP candidates (e.g., similar to a legacy construction process). As another example, the video encoding or decoding device may construct the merge MVP candidate list based on one or more history-based MVP (HMVP) candidates, in which case multiple HMVP candidates may be inserted into the merge list so that the merge list may reach a maximum allowed number of MVP candidates minus 1 . As yet another example, the video encoding or decoding device may construct the merge MVP candidate list based on a pairwise average candidate. In this case (e.g., when a pairwise average candidate may be added to the merge candidate list), the pairwise candidate may be determined as follows. The first two MVP candidates present in the list may be considered and their motion vectors may be averaged. The averaging may be computed separately for each reference picture list. If both MVP candidates are bi-directional, motion vectors related to both lists (e.g., L0 and L1) may be averaged. If one (e.g., only one) motion vector is present in a reference picture list, it may be selected as the pairwise candidate. As yet another example, the video encoding or decoding device may construct the merge MVP candidate list based on zero-valued MV candidates (e.g., MVs with values set to zeros), for example, by considering pairwise candidate(s) first, and then, if the list is not full, adding zero-valued motion vectors to the end of the list until it is full. [0139] FIG. 9 illustrates an example of a whole-block-based (e.g., also referred to non-subblock-based) merge list construction process.
[0140] The MMVD mode may allow for coding a limited motion vector difference (MVd or MVD) on top of a selected merge MVP candidate, for example, to represent the motion information of a CU. MMVD coding may be limited to a number of (e.g., 4) vector directions and a number of (e.g., 8) magnitude values. MMVD coding may be performed from 14 luma samples to 32-luma samples. MMVD may provide an intermediate accuracy level, which may lead to a trade-off between rate costs and MV accuracy for signaling motion information. FIG. 10 illustrates examples of allowed motion vector differences (MVd) in the MMVD mode.
[0141] MMVD offsets may be extended for MMVD and/or affine MMVD modes. Additional refinement positions (e.g., along kxrr/8 diagonal angles) may be added, as shown in FIG. 11 , for example, to increase the number of directions from 4 to 16. Based on a sum of absolute difference (SAD) cost between a template (e.g., one row above and one column to the left of a current block) and its reference for a (e.g., each) refinement position, one or more (e.g., all) possible MMVD refinement positions (e.g., 16x6) for a (e.g., each) base candidate may be reordered. Refinement positions (e.g., 8 refinement positions) with the smallest template SAD costs may be kept as available positions for MMVD index coding. The MMVD index may be binarized by a rice code, for example, with a binarization parameter set to 2. The affine MMVD reordering may be extended, in which case additional refinement positions (e.g., along kxrr/4 diagonal angles) may be added. After the reordering, the top refinement positions (e.g., top half of refinement positions) with the smallest template SAD costs may be kept.
[0142] The first N motion candidates in a candidate list before reordering may be utilized as the base candidates for MMVD and/or affine MMVD. For example, N may be set to 3 for MMVD, and [1 , 3] for affine MMVD depending on the neighboring block affine flags. Multiple (e.g., two) ways of adding MMVD offsets may be allowed, including, e.g., a two-sided way and a one-sided way, depending on whether the offset of the other reference picture list is mirrored or directly set to zero. Which of these ways are applied to a block may be dependent on a template matching (TM) cost.
[0143] Intra block copy (IBC) may be used for screen content coding. IBC may improve the coding efficiency of screen content materials. Since the IBC mode may be implemented as a block level coding mode, block matching (BM) may be performed (e.g., at an encoder) to find the optimal block vector (or motion vector) for a (e.g., each) CU. A block vector may indicate the displacement from a current block to a reference block, which may be reconstructed inside a current picture. A luma block vector of an IBC-coded CU may be in integer precision. A chroma block vector may be rounded to integer precision as well. When combined with an adaptive motion vector resolution (AMVR), the IBC mode may switch between 1 -pel and 4-pel motion vector precisions. An IBC-coded CU may be treated as a third prediction mode (e.g., in addition to intra and inter prediction modes). The IBC mode may be applicable to CUs with width and/or height smaller than or equal to 64 luma samples.
[0144] The activation of the IBC mode may be signaled (e.g., with a flag), for example, at a CU level, and it may be signaled as an IBC AMVP mode and/or an IBC skip/merge mode. With the IBC skip/merge mode, a merge candidate index may be used to indicate which of the block vectors in a merge list from neighboring candidate IBC coded blocks may be used to predict a current block. The merge list may include spatial, HMVP, and/or pairwise candidates. With the IBC AMVP mode, a block vector difference may be coded in the same way as a motion vector difference. A block vector prediction method may use multiple (e.g., two) candidates as predictors, such as, e.g., a candidate from a left neighbor and a candidate from an above neighbor (e.g., if the neighbor is IBC-coded). If a neighbor is not available, a default block vector may be used as a predictor. A syntax element (e.g., a flag) may be signaled to indicate the block vector predictor index.
[0145] The IBC mode may allow (e.g., only allow) for use of a reconstructed portion of a predefined area (e.g., including a region of a current coding tree unit (CTU) and/or a region of a left CTU), for example, to limit memory consumption and/or decoder complexity. FIG. 12 illustrates examples of reference regions in the IBC Mode, where a (e.g., each) block may represent a 64x64 luma sample unit.
[0146] The following may be true, for example, depending on the location of a current CU within a current CTU. If a current block falls into a top-left 64x64 block of the current CTU, the current block may refer to the reference samples in the bottom-right 64x64 block of the left, in addition to the reconstructed samples in the current CTU. The current block may refer to the reference samples in the bottom-left 64x64 block of the left CTU and the reference samples in the top-right 64x64 block of the left CTU.
[0147] If the current block falls into the top-right 64x64 block of the current CTU, and/or if luma location (0, 64) relative to the current CTU has not yet been reconstructed, the current block may refer to the reference samples in the bottom-left 64x64 block and bottom-right 64x64 block of the left, in addition to the already reconstructed samples in the current CTU. Otherwise, the current block may refer to reference samples in the bottom-right 64x64 block of the left CTU.
[0148] If the current block falls into the bottom-left 64x64 block of the current CTU, and/or if luma location (64, 0) relative to the current CTU has not yet been reconstructed, the current block may refer to the reference samples in the top-right 64x64 block and bottom-right 64x64 block of the left CTU, in addition to the already reconstructed samples in the current CTU. Otherwise, the current block may refer to the reference samples in the bottom-right 64x64 block of the left CTU. [0149] If the current block falls into the bottom-right 64x64 block of the current CTU, it may refer to the already reconstructed samples in the current CTU.
[0150] The aforementioned conditions and/or restrictions may allow the IBC mode to be implemented using a local on-chip memory.
[0151] The construction process of an IBC merge or AMVP list may be modified. For example, if an IBC merge/AMVP candidate is valid, it may be inserted into the IBC merge/AMVP candidate list. As another example, an above-right, bottom-left, or above-left spatial candidate, and/or a pairwise average candidate may be added into the IBC merge/AMVP candidate list. As yet another example, template based adaptive reordering (ARMC-TM) may be applied to the IBC/AMVP merge list.
[0152] A HMVP table size for the IBC mode may be increased, for example, to a larger number of (e.g., 25) entries. After multiple (e.g., up to 20) IBC merge candidates are derived with full pruning, they may be reordered. After the reordering, multiple (e.g., the first six) candidates with the lowest template matching costs may be selected as the final candidates in the IBC merge list.
[0153] Zero vector candidates that may be used to pad the IBC Merge/AMVP list may be replaced with a set of BVP candidates located in an IBC reference region. A zero vector may be invalid as a block vector in the IBC merge mode. Such a zero vector may be discarded as a block vector prediction (BVP) candidate in the IBC candidate list.
[0154] Multiple (e.g., three) candidates may be located on the corners (e.g., nearest corners) of a reference region. Additional (e.g., three) candidates may be determined in the middle of one or more (e.g., three) sub-regions (e.g., A, B, and C of FIG. 13). The coordinates of these sub-regions and/or the candidates may be determined by the width and height of the current block, and/or AX and AY parameters, as depicted in FIG. 13.
[0155] A reference in the IBC mode may be extended to multiple (e.g., two) CTU rows above the CTU being processed by an encoder or a decoder. FIG. 14 illustrates an example of a reference area for coding CTU (m,n). As shown in FIG. 14, the reference area may include CTUs with index (m-2,n-2) ... (W,n-2), (0,n— 1) ... (W,n— 1), (0,n) ... (m,n), where W may denote a maximum horizontal index within a current tile, slice or picture. A per-sample block vector search (e.g., a local search) range may be limited to [-(C « 1), C » 2] horizontally and [-C, C » 2] vertically, for example, to adapt to the reference area extension, where C may denote the CTU size.
[0156] Template matching (TM) based motion search and refinement may be applied in the IBC mode, which may be referred to as an IBC-TM merge mode. Such a merge mode may use a merge candidate list for block vector (BV) prediction, for example, in a different manner from other IBC merge modes (e.g., from a regular IBC merge mode). Candidates of the merge candiate list in the IBC-TM merge mode may be selected using a pruning method, with a motion distance between the candidates (e.g., similar to a regular TM merge mode). Zero motion candidates may be replaced by (-W, 0), (0, -H), (-W, -H) MVs.
[0157] In the IBC-TM merge mode, the selected candidates may be refined with a template matching method. A TM-merge syntax element (e.g., a flag) may be signaled to indicate that the IBC-TM merge mode is enabled. In an IBC-TM AMVP mode, multiple (e.g., up to three) candidates may be selected from the IBC-TM merge list. One or more of those candidates (e.g., each of those candidates) may be refined using a template matching method and may be sorted according to the resulting TM costs.
[0158] When used for IBC, TM refinement may be performed at integer pel positions. In the IBC-TM AMVP mode, TM refinement may be performed with integer precision or fractional (e.g., 4-pel) precision, for example, depending on the relevant AMVR value. The refinement may be performed within an existing IBC reference area.
[0159] The interaction between IBC mode and/or other inter coding concepts or tools, such as pairwise merge candidates, history-based motion vector predictors (HMVP), combined intra/inter prediction mode (CUP), merge mode with motion vector difference (MMVD), and/or geometric partitioning mode (GPM) may be as follows. For example, IBC may be used with pairwise merge candidates and/or HMVP. A pairwise IBC merge candidate may be generated by averaging multiple (e.g., two) IBC merge candidates. For HMVP, IBC motion may be inserted into a history buffer for future reference. As another example, IBC may not be used in combination with certain inter tools such as affine motion. As yet another example, IBC may be used in combination with CUP, MMVD, and/or GPM. As yet another example, IBC may not be allowed for chroma coding blocks if DUAL_TREE partition is used.
[0160] In examples, a picture may not be included as a reference picture in a reference picture list (e.g., L0) for IBC prediction. A derivation process of motion vectors for the IBC mode may exclude neighboring blocks coded in the inter mode, and vice versa. One or more of the following IBC design aspects may be applied. For example, IBC may share a same process as a regular MV merge. For example, an IBC process may allow the use of pairwise merge candidates and/or history-based motion predictors, but may disallow TMVP and/or zero vectors (e.g., because they may be invalid for the IBC mode). As another example, separate HMVP buffers (e.g., five candidates each) may be used for a conventional MV and/or IBC. As yet another example, block vector constraints may be implemented in the form of bitstream conformance constraints. An encoder may ensure that no invalid vectors are present in a bit-stream, and merge may not be used if the merge candidates are invalid (e.g., out of range or 0). Such bitstream conformance constraints may be expressed in terms of a virtual buffer, for example, as described herein. As yet another example, for deblocking, IBC may be handled as an inter mode. As yet another example, if the current block is coded using the IBC mode, AMVR may not use quarter-pel and AMVR may be signaled to indicate whether an MV has integer precision (e.g., inter-pel) or fractional precision (e.g., 4 integer-pel). As yet another example, the number of IBC merge candidates may be signaled in a slice header, for example, separately from the number of regular candidates, number of subblocks, and/or number of geometric merge candidates.
[0161] An merge with motion vector differences (MMVD) mode may have one or more of the following characteristics. An Affine-MMVD and/or GPM-MMVD may be adopted as an extension of a regular MMVD mode. The MMVD mode may be extended to the IBC merge mode and an IBC merge mode with block vector differences (IBC-MBVD) may be derived. In the IBC-MBVD, a distance set associated with block vector differences (BVD) may be {1 -pel, 2-pel, 4-pel, 8-pel, 12-pel, 16-pel, 24-pel, 32-pel, 40-pel, 48-pel, 56-pel, 64-pel, 72-pel, 80-pel, 88-pel, 96-pel, 104-pel, 112-pel, 120-pel, 128-pel}, and BVD directions may include horizontal and vertical directions (e.g., two horizontal and two vertical directions).
[0162] A base candidate may be selected from a number of (e.g., first five) candidates in a reordered IBC merge list. Based on the SAD cost between a template (e.g., one row above and one column to left of the current block) and its reference for a (e.g., each) refinement position, MBVD refinement positions (e.g., 20x4 refinement positions) for a (e.g., each) base candidate may be reordered. The top (e.g., eight) refinement positions with the lowest template SAD costs may be selected as available positions (e.g., for MBVD index coding). An MBVD index may be binarized by a rice code with a binarization parameter set to 1.
[0163] The IBC coding mode may be used for camera-captured contents. Multiple adaptations may be adopted to improve the IBC mode (e.g., in terms of compression efficiency). For example, high level tool control may be adopted, in which case one or more IBC merge modes may be disabled for natural contents. A syntax element such as an SPS level flag may be used to disable associated CU level signaling. An IBC AMVP mode may be activated (e.g., if indicated by an SPS flag). As another example, RR-IBC, TM-IBC, and IBC-CI IP may be disabled for natural contents. For at least RR-IBC and TM-IBC, corresponding syntax elements (e.g., SPS flags) may be implemented. As yet another example, IBC may be applied to intra slices for natural contents. This may be indicated by a high-level syntax element (e.g., to remove the signaling of an IBC flag at the CU level). For screen contents, IBC may be applied to multiple types (e.g., all types) of slices.
[0164] At an encoder, an IBC block vector search may be optimized for natural contents. The RDO process may be skipped for an IBC AMVP mode if its SAD cost is worse than the lowest SAD cost of other (e.g., all) Intra modes. The IBC AMVP mode may not be evaluated if the best Intra mode has fewer than a pre-determined number of (e.g., three) nonzero coefficients. Certain partitioning depths in an inter slice may be skipped depending on the picture order count (POC) distance between the current picture and its nearest reference picture. If IBC is enabled at an SPS level, the POC distance may be set equal to 0, which may not align with configurations used for random access and/or low delay. In some examples, a non-zero POC distance may be used.
[0165] A fractional-pel extension may be adopted for IBC, in which case the representation of IBC block vectors may be extended to fractional-pel resolution. An interpolation filter may be used to derive the prediction samples located at a non-integer phase in a reconstructed area of a current frame. Block vector resolutions may be extended to include quarter-pel resolution (e.g., in addition to full-pel and 4-pel resolutions). The first bin (e.g., binary element) of syntax representing an AMVR index may be signaled to indicate whether a BV is in quarter-pel resolution, while a second bin may be signaled to switch between full-pel and 4-pel resolutions. The interpolation filters applied to the luma and chroma components of an IBC block may be 8-tap luma filters and the same (e.g., substantially similar) chroma filters as used in motion compensation, respectively. An exception may be that a 2-tap bilinear interpolation filter may be used to generate template prediction blocks in IBC-related coding tools. Reference sample padding may be performed where reference samples are not available or are located outside a valid IBC reference area in the current frame. In examples, the reference sample padding may be performed in the horizontal direction first and then in the vertical direction.
[0166] An IBC-MBVD merge mode may be improved as follows. Adaptive BVD offsets may be allowed along MBVD directions. An MBVD list of K candidates with lowest template SAD costs may be derived via the following steps. At a first step (e.g., step 1), the largest offset may be denoted as N-pel (e.g., N=256), the number of directions may be denoted as D (e.g., D=4: left, right, top, and bottom), the starting search interval may be denoted as M-pel (e.g., M=8), and the number of candidates in an MBVD list may be denoted as K (e.g., K=8). At a second step (e.g., step 2), the TM SAD cost for an offset of every M-th position may be checked (e.g., along one or more directions) to ensure that it does not exceed N. The K lowest TM SAD cost candidates may be kept in the list. At a third step (e.g., step 3), the TM SAD cost of multiple (e.g., two) candidates with an offset equal to +-M/2 along a direction may be checked (e.g., for one or more candidates in the list). The K lowest TM SAD cost candidates may be kept in the list. The third step (e.g., a refinement step) may be repeated (e.g., as a fourth step or step 4), while reducing the interval M (e.g., by half each time), for example, until M reaches 1 -pel.
[0167] If an MBVD candidate has a left or above template outside of the BV search area, the MBVD candidate may be set to invalid candidate. If multiple (e.g., two) BVP candidates have the same BV component in at least one of the horizontal direction or vertical direction, the BV difference in the other direction may be larger than a threshold T (e.g., T = 8-pel). The amount of MBVD candidates and MBVD index signaling may be kept the same as in a legacy system (e.g., as in ECM-8.0). [0168] FIG. 15 illustrates the MBVD improvements described herein (e.g., including the first, second and third steps described above).
[0169] The IBC-MBVD block vector coding mode of IBC may be activated (e.g., enabled) to compress camera-captured video contents. The MBVD block vector representation described herein (e.g., with reference to FIG. 15) may be used when coding camera-captured video contents (e.g., these contents may also be referred to as natural videos in the examples provided herein).
[0170] Extended block vector directions may be supported in the IBC-MBVD coding mode described herein. For example, diagonal directions may be supported in a block vector difference representation of the MBVD mode (e.g., including the IBC-MBVD mode). In a merge with motion vector difference (MMVD) mode, intermediate angles between the vertical and horizontal directions may be considered. Allowing diagonal MBVD directions may increase the compression performance of IBC-MBVD coding tools, for example, for at least natural video contents coding.
[0171] BV difference based coding may include angular predictions (e.g., in the IBC-MBVD coding mode), for example, in addition to vertical and/or horizontal directions. For example, block vector difference (BVD) offsets may be extended to BV offsets in top-left, top-right, bottom-left and/or bottom-right directions. The number of BVD magnitudes in the additional directions (e.g., diagonal directions) may be reduced (e.g., compared to the number of BV magnitudes supported in the vertical and/or horizontal directions). The additional directions may include diagonal directions, and/or any non-vertical and non-horizontal directions. The use of these additional directions (e.g., IBC-MBVD directions) may be controlled and/or signaled, for example, at a sequence level (e.g., via one or more sequence parameter set (SPS) dedicated syntax elements). The use of these additional directions (e.g., IBC-MBVD directions) may be controlled and/or signaled, for example, at a picture, slice, sub-picture, tile, or tile group level (e.g., via one or more dedicated syntax elements). In examples, step 2 of the process illustrated by FIG. 15 (e.g., with respect to TM SAD cost determination during BVD candidate search) may be extended to consider diagonal directions (e.g., including any non-vertical and non-horizontal direction). In examples, when performing step 3 of the process described above (e.g., with respect to refining positions around the K best candidates derived from step 2), positions not located in the same direction as the candidate BVDs may be considered. For instance, multiple (e.g., eight) positions around a candidate (e.g., each of the K best candidates), such as those with relative coordinates of (-M/2, 0), (M/2, 0), (0, -M/2), (0, M/2), (-M/2, -M/2), (- M/2, M/2), (M/2, -M/2), and (M/2, -/2)), may be considered (e.g., in the refinement operations of step 3). As another example, an extended template matching (TM) search in a diagonal direction (e.g., including any non-vertical and non-horizontal direction) for IBC-MBVD may be performed for (e.g., only for) camera- captured video contents. As yet another example (e.g., if no template matching applies in the IBC-MBVD mode), the signaling of the IBC-MBVD coding mode may be adapted to allow the use of diagonal (e.g., including any non-vertical and non-horizontal) BV difference, for example, when coding natural video contents.
[0172] BV difference representations and/or coding techniques in the IBC-MBVD mode may be extended to support directions other than vertical and horizontal directions. The additional directions may include multiples of TT/4 and/or multiples of TT/8.
[0173] FIG. 16 illustrates an example of a BV decoding process in the IBC-MBVD coding mode. Two stages are shown in this figure: BV parsing and BV derivation. A syntax element such as a flag may be used to indicate if a coding unit (CU) may be parsed and further processed in the IBC-MBVD mode. Decoding of non-IBC-MBVD CUs may be performed, but is not shown in FIG. 16.
[0174] As shown in FIG. 16, MBVD merge information may be parsed, for example, if the CU is in the IBC-MBVD mode. This may take the form of a base BV predictor index, which may identify a BVP merge candidate in an IBC merge list. A parsed BV difference index may lead to (e.g., be represented by) index BVDJdx. Block vector(s) for the CU may be derived as follows. An IBC merge candidate list may be constructed. The candidate list may be reordered, for example, according to increases in template matching costs. A base BV predictor may be determined as the BV candidate located at the rank corresponding to the parsed base BV predictor index in the IBC merge candidate list. BV positions (e.g., all possible BV positions) that may be reachable by supported BV offset values may be ordered, for example, according to increasing template matching costs. The BV difference (e.g., also referred to as BV offset) associated with a current CU may be obtained as the BVD candidate ranked at position BVDJdx obtained from the parsing stage. A BV of the CU may be derived as the sum of the base BV predictor and the obtained BV difference, for example, based on BV = mvbvd_ba.se + bvd.
[0175] MBVD offsets associated with vectors bvdistance x bvdir, e.g., as illustrated below, may be supported. bvdistance E BV _Distance_Set
BV -Distance _Set
= { 1,2,4,8,12,16,24,32,40,48,56,64,72,80,88,96,104,112,120,128} bvdir E BV_Dir_ Set = {(1,0), (-1,0), (0,1), (0, -1)}
[0176] FIG. 17 illustrates examples of BVD (e.g., BV difference or BV offset) directions that may be supported in the IBC-MBVD mode. As shown, the supported BVD directions may include positions marked in black (e.g., 1702) in FIG. 17. The supported BVD directions may also include (e.g., be extended to) positions marked in gray (e.g., 1704) in FIG. 17, such as the following set: BV_Dir_Set = {(1,0), (-1,0), (0,1), (0, -1), (1,1), (1, -1), (-1,1), (-1, -1)}
[0177] The extended BVD directions may be used, for example, in at least some of the steps illustrated in FIG. 16 (e.g., the steps associated with sorting, ordering, and/or determining IBC-MBVD BVD position candidates).
[0178] The extended BVD directions may also include the following set (e.g., positions marked in light gray such as 1706 in FIG. 17):
BV_Dir_Set = {(1,0), (-1,0), (0,1), (0, -1), (1,1), (1, -1), (-1,1), (-1, -1), (2,1), (1,2), (2, -1), (1, —2), (-2,1), (-1,2), (—2, -1), (-1, -2)}
[0179] The magnitudes of block vector differences or offsets (e.g., BV_Distance_Set shown herein) supported in the IBC-MBVD mode may be adaptive to the direction of the block vector offsets. For example, fewer magnitudes may be supported for diagonal positions (e.g., non-vertical and non-horizontal positions) compared to vertical and horizontal positions. Multiple (e.g., two) sets of BV distance (BVD) magnitudes may be specified and/or used. For example, one set of BVD magnitudes may be used for the vertical/horizontal directions, while another set of distances may be used for the diagonal directions (e.g., including any non-vertical and non-horizontal directions), e.g., as follows:
BV -Distance _Set_HorV er
= { 1,2,4,8,12,16,24,32,40,48,56,64,72,80,88,96,104,112,120,128}
BV_Distance_Set_NonHorVer = { 1,2,4,8,12,16,24,32,40,48,56,64}
[0180] BVD magnitudes may be provided and/or used according to the directions of the BVD (e.g., whether a direction is a multiple of TT/2, TT/4, or TT/8), e.g., as illustrated below:
BV -Distance _Set_HorV er
= { 1,2,4,8,12,16,24,32,40,48,56,64,72,80,88,96,104,112,120,128}
BV -Distance _Set_pi_4 = { 1,2,4,8,12,16,24,32,40,48,56,64}
BV_Distance_Set_pi_8 = { 1,2,4,8,12,16,24,32}
[0181] Having adaptive (e.g., direction-adaptive) sets of BVD magnitudes may provide a trade-off between increased decoder/encoder complexity and increased compression efficiency.
[0182] MBVD representations having diagonal (e.g., including non-vertical and non-horizontal) directions may be supported for the coding of camera-captured video contents. As described herein, an IBC-MBVD coding mode may be supported in natural content coding (e.g., an IBC merge mode may be deactivated for such natural content coding). The extended IBC-MBVD mode (e.g., including extended BVD directions) may be enabled for camera contents. One or more syntax elements (e.g., high-level syntax elements) may be used to indicate the use of an extended set of IBC-MBVD directions. For example, the use of extended IBC-MBVD directions may be controlled and/or signaled at a sequence level (e.g., via one or more SPS dedicated syntax elements). As yet another example, the use of extended IBC-MBVD directions may be controlled and/or signaled at a picture, slice, sub-picture, tile, or tile group level (e.g., via one or more dedicated syntax elements).
[0183] As an example, extended BVD directions and/or associated magnitudes in the IBC-MBVD mode may be applied to one or more operations illustrated in FIG. 15. For instance, steps 1 and 2 shown in FIG. 15 may be extended to consider diagonal (e.g., non-vertical and non-horizontal) directions. These extended directions may, for example, correspond to the diagonal positions shown in FIG. 18, which, in turn, may correspond to the positions associated with multiples of (M,M), (M,-M), (-M,M), (-M.-M), where M may be equal to 8. The refinement operations described with reference to FIG. 15 (e.g., step 3 and/or step 4) may consider vertical and/or horizontal positions around the K best positions in an (e.g., each) extended direction as starting positions for a BVD search.
[0184] As another example, when refining positions (e.g., in step 3 described herein) around K best candidates (e.g., derived from step 2 described herein), one or more positions not located in the same direction as the candidate BVDs may be considered. For instance, multiple (e.g., eight) positions around a candidate (e.g., each of the K best candidates), such as those associated with relative coordinates (-M/2, 0), (M/2,0), (0.-M/2), (O.M/2), (-M/2.-M/2), (-M/2.M/2), (M/2, -M/2), (M/2.-/2), may be considered in the refinement operation (e.g., step 3 described herein).
[0185] FIG. 19 illustrates the use of an extended set of BVD candidates, for example, in step 3 and/or step 4 of the IBC-MBVD operations described herein.
[0186] One or more of the techniques described herein may be integrated in a video coding (e.g., encoding and/or decoding) system that may not use a template matching process while in an IBC-MBVD coding mode. In such a system, one or more aspects of MBVD signaling may be modified. For example, a signaling process in such a system may include a bvd_base_idx syntax element, a bvd_direction_idx syntax element, and a bvd_distance_idx syntax element. These syntax elements may respectively indicate the base IBC merge candidate to which a BV offset may be applied, the direction (e.g., vertical, horizontal, or diagonal) of the BV offset, and the magnitude of the BV offset. The bvd_direction_idx syntax element may be used to indicate a vertical direction, a horizontal direction, and/or a diagonal direction (e.g., as a multiple of TT/4 or a multiple of TT/8). The value of bvd_distance_idx may be dependent on the value of bvd_direction_idx. For example, for a direction index corresponding to a diagonal direction, the BVD distance index may derive its value in a reduced set of possible values (e.g., compared to vertical/horizontal directions).

Claims

Claims
1 . A video decoding device, comprising: a processor configured to: obtain a base block vector (BV) associated with a current video block coded in a block vector-based coding mode; determine a plurality of candidate BV offsets associated with the base BV, wherein the plurality of candidate BV offsets includes a diagonal BV offset; determine a refined BV based on the base BV and the plurality of candidate BV offsets, wherein the determination is made at least partially based on the diagonal BV offset; and decode the current video block based on the refined BV.
2. The video decoding device of claim 1 , wherein the diagonal BV offset points at a 45-degree angle from a horizontal direction or a vertical direction.
3. The video decoding device of claim 1 , wherein the diagonal BV offset points at a 22.5-degree angle from a horizontal direction or a vertical direction.
4. The video decoding device of claim 1 , wherein the plurality of candidate BV offsets further includes a horizontal or vertical BV offset, and wherein the processor is further configured to determine a first set of candidate magnitudes associated with the diagonal BV offset and a second set of candidate magnitudes associated with the horizontal or vertical BV offset, the first set of candidate magnitudes includes fewer candidates than the second set of candidate magnitudes.
5. The video decoding device of claim 4, wherein the number of candidates included in the first set of candidate magnitudes is dependent on an angle of the diagonal BV offset.
6. The video decoding device of claim 1 , wherein the current video block includes camera-captured video content.
7. The video decoding device of claim 1 , wherein the current video block is decoded using an intra block copy (IBC) prediction mode.
8. The video decoding device of claim 7, wherein the current video block is decoded using an IBC merge with block vector differences (IBC-MBVD) prediction mode.
9. A video decoding method, comprising: obtaining a base block vector (BV) associated with a current video block coded in a block vectorbased coding mode; determining a plurality of candidate BV offsets associated with the base BV, wherein the plurality of candidate BV offsets includes a diagonal BV offset; determining a refined BV based on the base BV and the plurality of candidate BV offsets, wherein the determination is made at least partially based on the diagonal BV offset; and decoding the current video block based on the refined BV.
10. The video decoding method of claim 9, wherein the diagonal BV offset points at a 45-degree angle or a 22.5-degree angle from a horizontal direction or a vertical direction.
11 . The video decoding method of claim 9, wherein the plurality of candidate BV offsets further includes a horizontal or vertical BV offset, and wherein the video decoding method further comprises determine a first set of candidate magnitudes associated with the diagonal BV offset and a second set of candidate magnitudes associated with the horizontal or vertical BV offset, the first set of candidate magnitudes includes fewer candidates than the second set of candidate magnitudes.
12. The video decoding method of claim 11 , wherein the number of candidates included in the first set of candidate magnitudes is dependent on an angle of the diagonal BV offset.
13. The video decoding method of claim 9, wherein the current video block includes camera-captured video content.
14. The video decoding method of claim 9, wherein the current video block is decoded using an intra block copy merge with block vector differences prediction mode.
15. A video encoding device, comprising: a processor configured to: obtain a base block vector (BV) associated with a current video block coded in a block vector-based coding mode; determine a plurality of candidate BV offsets associated with the base BV, wherein the plurality of candidate BV offsets includes a diagonal BV offset; determine a refined BV based on the base BV and the plurality of candidate BV offsets, wherein the determination is made at least partially based on the diagonal BV offset; and encode the current video block based on the refined BV.
16. The video encoding method of claim 15, wherein the diagonal BV offset points at a 45-degree angle or a 22.5-degree angle from a horizontal direction or a vertical direction.
17. A video encoding method, comprising: obtaining a base block vector (BV) associated with a current video block coded in a block vectorbased coding mode; determining a plurality of candidate BV offsets associated with the base BV, wherein the plurality of candidate BV offsets includes a diagonal BV offset; determining a refined BV based on the base BV and the plurality of candidate BV offsets, wherein the determination is made at least partially based on the diagonal BV offset; and encoding the current video block based on the refined BV.
18. A computer program product which is stored on a non-transitory computer readable medium and comprises program code instructions for implementing the steps of a method according to any one of claims 9-14 or claim 17 when executed by a processor.
19. A computer program comprising program code instructions for implementing the steps of a method according to any one of claims 9-14 or claim 17 when executed by a processor.
20. Video data comprising information representative of a current video block encoded using a method according to claim 17.
EP24732322.3A 2023-06-29 2024-06-14 Improvements to intra block copy Pending EP4736416A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP23306059 2023-06-29
PCT/EP2024/066638 WO2025002858A1 (en) 2023-06-29 2024-06-14 Improvements to intra block copy

Publications (1)

Publication Number Publication Date
EP4736416A1 true EP4736416A1 (en) 2026-05-06

Family

ID=87429446

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24732322.3A Pending EP4736416A1 (en) 2023-06-29 2024-06-14 Improvements to intra block copy

Country Status (3)

Country Link
EP (1) EP4736416A1 (en)
CN (1) CN121359443A (en)
WO (1) WO2025002858A1 (en)

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118285100A (en) * 2021-09-15 2024-07-02 抖音视界有限公司 Method, device and medium for video processing

Also Published As

Publication number Publication date
WO2025002858A1 (en) 2025-01-02
CN121359443A (en) 2026-01-16

Similar Documents

Publication Publication Date Title
WO2025133069A1 (en) Intra tmp-lic merge mode
WO2025003037A1 (en) Intra tmp and lic combination
EP4732532A1 (en) Generalized intra prediction fusion
EP4629620A1 (en) Motion vectors as merge candidates for intra blocks
EP4629619A1 (en) Generalized usage of auto relocated block vector
EP4633142A1 (en) Interaction between adaptive dual tree and chroma direct block vector prediction mode
EP4661392A1 (en) Adaptive dual tree and chroma intra prediction modes reordering
EP4633138A1 (en) Interaction between adaptive dual tree and intra-block copy
EP4736416A1 (en) Improvements to intra block copy
EP4736417A1 (en) Bi-predictive merge list for intra block copy coding
WO2025003042A1 (en) Block vector guided chroma direct mode
WO2025131654A1 (en) Intratmp lic extended template and probing
WO2025003313A1 (en) Adaptive ibc/intra tmp filtering
WO2025073519A1 (en) Filtering applied to chroma direct block vector
WO2024133579A1 (en) Gpm combination with inter tools
WO2025003029A1 (en) Sbt applied to ibc and intratmp
WO2025073548A1 (en) Sub-partitioning associated with template matching
WO2025002871A1 (en) Affine block vector model for intra block copy
EP4702743A1 (en) Gpm boundary prediction
WO2025002778A1 (en) Amvr interactions with filtered prediction
WO2024133007A1 (en) Hmvp candidate reordering
WO2025003104A1 (en) Intra block copy geometric partitioning mode (ibc-gpm) with bi-predictive block vectors
EP4736425A1 (en) Derivation of coding parameters
EP4725194A1 (en) Gpm with inter and ibc prediction
WO2025003493A1 (en) Bi-predictive intra block copy with weighted averaging

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE