WO2025201358A1 - 用于管理媒体项的方法、装置、设备和介质 - Google Patents

用于管理媒体项的方法、装置、设备和介质

Info

Publication number
WO2025201358A1
WO2025201358A1 PCT/CN2025/084840 CN2025084840W WO2025201358A1 WO 2025201358 A1 WO2025201358 A1 WO 2025201358A1 CN 2025084840 W CN2025084840 W CN 2025084840W WO 2025201358 A1 WO2025201358 A1 WO 2025201358A1
Authority
WO
WIPO (PCT)
Prior art keywords
target
media items
group
acquisition
live broadcast
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/084840
Other languages
English (en)
French (fr)
Inventor
邢伟科
颜志武
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Publication of WO2025201358A1 publication Critical patent/WO2025201358A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/21Server components or server architectures
    • H04N21/218Source of audio or video content, e.g. local disk arrays
    • H04N21/2187Live feed
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/25Management operations performed by the server for facilitating the content distribution or administrating data related to end-users or client devices, e.g. end-user or client device authentication, learning user preferences for recommending movies
    • H04N21/258Client or end-user data management, e.g. managing client capabilities, user preferences or demographics, processing of multiple end-users preferences to derive collaborative data
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/25Management operations performed by the server for facilitating the content distribution or administrating data related to end-users or client devices, e.g. end-user or client device authentication, learning user preferences for recommending movies
    • H04N21/258Client or end-user data management, e.g. managing client capabilities, user preferences or demographics, processing of multiple end-users preferences to derive collaborative data
    • H04N21/25808Management of client data
    • H04N21/25841Management of client data involving the geographical location of the client
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/45Management operations performed by the client for facilitating the reception of or the interaction with the content or administrating data related to the end-user or to the client device itself, e.g. learning user preferences for recommending movies, resolving scheduling conflicts
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/45Management operations performed by the client for facilitating the reception of or the interaction with the content or administrating data related to the end-user or to the client device itself, e.g. learning user preferences for recommending movies, resolving scheduling conflicts
    • H04N21/4508Management of client data or end-user data
    • H04N21/4524Management of client data or end-user data involving the geographical location of the client

Definitions

  • Exemplary implementations of the present disclosure relate generally to media management, and more particularly to methods, apparatuses, devices, and computer-readable storage media for managing media items in a live broadcast application.
  • live streaming applications With the advancement of computer and network technologies, a variety of live streaming applications have been developed. These applications can provide live streaming services to a large audience in real time. Furthermore, multi-location live streaming technology solutions have been proposed. For example, multiple acquisition devices can be deployed in a live streaming environment to provide viewers with live video from multiple viewpoints. However, because the deployment locations of multiple acquisition devices are fixed, viewers cannot choose viewing positions outside of these fixed locations.
  • a device for managing media items includes: a location acquisition module configured to acquire a target location for viewing a live broadcast object; a media acquisition module configured to acquire a set of media items based on the target location, the set of media items being from a set of acquisition devices associated with the live broadcast object; and a determination module configured to determine, based on the set of media items, a target media item that matches the target location.
  • an electronic device in a third aspect of the present disclosure, includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.
  • a computer-readable storage medium on which a computer program is stored.
  • the processor implements the method according to the first aspect of the present disclosure.
  • FIG1A shows a block diagram of an application environment according to an exemplary implementation of the present disclosure
  • FIG2 illustrates a block diagram for managing media items according to some implementations of the present disclosure
  • FIG9 illustrates a block diagram of an apparatus for managing media items according to some implementations of the present disclosure.
  • the term "in response to” refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.
  • media items may include at least any one of images and/or videos.
  • the process of media item management will be described using only videos as an example. In this way, it is possible to support the provision of richer media data to users.
  • a group of media items 230 can be obtained, each of which comes from a group of acquisition devices associated with the live broadcast object.
  • a group of acquisition devices may be at least a portion of a plurality of acquisition devices.
  • a plurality of acquisition devices e.g., acquisition devices 124, 126, ..., and 128) are associated with the target object and are deployed within a predetermined spatial range of the second user (i.e., the target object).
  • the predetermined spatial range may represent a pre-specified spatial range around the live broadcast object, for example, the room where the live broadcast object is located, or a spatial range within 3 meters (or other values) around the live broadcast object.
  • users can be provided with more flexible control methods, allowing them to select any desired viewpoint.
  • media items from relevant capture devices among multiple capture devices can be used to generate a video at that viewpoint.
  • the determined media items are generated using a set of media items captured in real time, and thus can include more accurate and rich visual information. In this way, users can view the live broadcast from any desired viewpoint during the live broadcast.
  • the virtual space can present the positions of multiple acquisition devices and the position of the second user 120 in three-dimensional coordinates.
  • the second user 120 can be located at the coordinate origin O, and the multiple acquisition devices can be located at multiple points P1, P2, P3, P4, and P5 on the x, y, and z coordinate axes, respectively.
  • the user can perform an interactive operation 320 to select a target position 310 (for example, represented by coordinates (x 0 , y 0 , z 0 )) at which the live video is desired to be viewed.
  • the first user can determine the target position 310 by single-clicking, double-clicking, dragging, or the like.
  • a selection page can be provided to the user in a visual manner so that the user can select the desired viewpoint position in a more convenient and accurate manner.
  • FIG3 schematically illustrates an example in which multiple acquisition devices are deployed at respective coordinate axes
  • the acquisition devices can be deployed at locations other than the coordinate axes.
  • acquisition devices can be deployed at any location on the sphere shown in FIG3 along a direction toward the origin. In this way, media items from more angles can be provided, thereby improving the accuracy of generating target media items.
  • the distances between the multiple acquisition devices and the origin O can be different. In this way, long-distance images and/or close-up images can be presented according to user needs.
  • a second client device can transmit multiple media items from multiple capture devices to a server device.
  • the multiple media items can be transmitted individually to the server device.
  • encoding operations can be performed on the multiple media items so that the multiple media items are transmitted as a single data stream. In this way, the bandwidth involved in the transmission process can be reduced, transmission delay can be reduced, and transmission efficiency can be improved.
  • a server device of a live broadcast application can determine which media items to provide to a first client device based on a target location. In this way, the first client device can directly receive the required media items from the server device, thereby reducing the workload of the first client device.
  • a server device can obtain a set of directions for a group of collection devices.
  • the first direction of the first collection device is from the first position of the first collection device to the position of the second user.
  • the directions of each collection device at points P1 to P5 all point to the origin O.
  • the space defined by the target location and the directions of each collection device can be compared to determine a group of collection devices.
  • a vector passing through the position of the acquisition device can be used to represent the direction of the acquisition device.
  • the target position may be located in a one-dimensional space (i.e., a straight line) defined by the position of the acquisition device and the position of the second user
  • the target position may be located in a two-dimensional space (i.e., a plane) defined by the positions of two acquisition devices and the position of the second user
  • the target position may be located in a three-dimensional space (i.e., a volume space) defined by the positions of three acquisition devices and the position of the second user.
  • FIG. 4 which illustrates a block diagram 400 for determining a group of collection devices according to some implementations of the present disclosure.
  • the collection device at P1 e.g., referred to as the first collection device
  • the collection device at P2 e.g., referred to as the second collection device
  • the collection device at P5 e.g., referred to as the third collection device
  • the target position is at position 410
  • the target position is located in a one-dimensional space (i.e., the x-axis) defined by a first position of a first acquisition device (i.e., point P1) and a position of a second user (i.e., origin O).
  • a set of media items may include media items acquired by the first acquisition device.
  • FIG4 is merely illustrative, and the target position may be located elsewhere. Assuming the target position is located on the y-axis, a set of media items may include media items from the capture device at point P2, and so on. In this way, by comparing the positional relationship between the target position and the straight line defined by the position of the capture device and the position of the second user, the data required to generate the target media item can be determined in a simple and accurate manner.
  • the target position is located in a two-dimensional space (i.e., the xoy plane) defined by a first position of a first acquisition device (i.e., point P1), a second position of a second acquisition device (i.e., point P2), and a position of a second user (i.e., origin O).
  • a set of media items may include media items collected by the first acquisition device and media items collected by the second acquisition device.
  • FIG4 is merely illustrative, and the target location can be located elsewhere.
  • a set of media items can include media items from a capture device at point P1 and media items from a capture device at point P5, and so on.
  • the data required to generate the target media item can be determined in a simple and accurate manner.
  • FIG5 illustrates a block diagram 500 for determining a group of capture devices according to some implementations of the present disclosure. As shown in FIG5 , assume that the target location is at position 510.
  • position 510 is located in a three-dimensional space defined by a first position of a first capture device (i.e., point P1), a second position of a second capture device (i.e., point P2), a third position of a third capture device (i.e., point P5), and a position of a second user (i.e., origin O) (i.e., the first quadrant defined by origin O, +x-axis, +y-axis, and +z-axis).
  • a group of media items may include media items captured by the first capture device, media items captured by the second capture device, and media items captured by the third capture device.
  • FIG5 is merely schematic, and the target position may be located elsewhere. Assuming the target position is located in a space defined by the origin O, the -x axis, the +y axis, and the +z axis, a set of media items may include media items from capture devices at points P2, P3, and P5, and so on. In this way, by comparing the positional relationship of the target position and the plane defined by the positions of the three capture devices and the position of the second user, the data required to generate the target media item can be determined in a simple and accurate manner.
  • a server device may generate target media items based on a target location and a set of media items and transmit the generated media items to client devices of respective viewers. It should be understood that during a live broadcast, there may be a large number of viewers, and generating target media items at the server device may result in an excessive workload on the server device.
  • a server device may transmit a determined set of media items to a first client device.
  • each media item may be transmitted separately to the server device.
  • each media item may be encoded so as to be transmitted as a single data stream.
  • the client device may receive a set of media items from the server device and generate target media items using the set of media items. In this way, the workload of the server device may be reduced, and the server device may be able to focus on other computational processing processes of the live broadcast process.
  • a first client device can determine a set of weights based on the distance between a target location and a set of directions of a set of capture devices. Subsequently, target media items can be determined based on the set of weights and the set of media items.
  • the server device only transmits media items from the capture device at point P1 to the first client device. At this point, the distance between the target location and the x-axis is zero, so the media items from the capture device at point P1 can be directly used as target media items.
  • the server device can transmit two media items from the capture devices at points P1 and P2 to the first client device.
  • the corresponding weights can be determined based on the ratio between location 420 and the x-axis and y-axis.
  • location 420 is at the angle bisector of the x-axis and y-axis (i.e., 45 degrees)
  • the contributions from the two media items are equal, and the specific content of the target media item can be determined based on a weighted summation.
  • the distance between position 510 and the directions of each acquisition device i.e., +x axis, +y axis, and +z axis
  • the distance between position 510 and the directions of each acquisition device can be determined to determine the corresponding weights.
  • the distance 416 between position 510 and the +x axis is expressed as The distance 412 between the position 510 and the +y axis is represented as The distance 414 between the position 510 and the +z axis is represented as In this way, determining the contribution of each media item to the final target media item can be converted into a mathematical operation process, thereby determining the target media item in a simpler and more accurate manner.
  • color represents the color data of pixel 652
  • color x represents the color data of pixel 612
  • color y represents the color data of pixel 622
  • color z represents the color data of pixel 632.
  • the color data of each pixel in the target media item can be determined based on simple mathematical operations. It should be understood that the above formula is merely exemplary, and alternatively and/or additionally, a media item with a viewpoint at a desired position and/or orientation can be generated based on multiple media items from capture devices deployed at different positions and/or orientations based on various technical solutions currently known and/or to be developed in the future.
  • a group of images when the media item is an image, a group of images can be used to generate a final image based on the method described above.
  • a group of images with the same timestamp can be used to generate an image associated with that timestamp.
  • a group of images at time point t can be used to generate an image at time point t
  • a group of images at time point t+1 can be used to generate an image at time point t+1.
  • the images can then be presented in chronological order.
  • the target position is located at position 720, that is, the distance between the target position and the origin O is less than the distance between the acquisition device at point P2 and the origin O, at this time, the image content in the target media item can be reduced. In this way, users can be supported to select viewpoint positions in a more flexible manner, thereby obtaining more information about live broadcasts.
  • the first user can be a viewer in a live broadcast application
  • the live broadcast object can be a second user in the live broadcast application, such as at least one of the host or guest.
  • ordinary viewers can select any viewpoint to watch the live broadcast.
  • the generated video is generated using a set of media items collected in real time, and thus can include more accurate and rich visual information. In this way, users can watch the live broadcast from any desired viewpoint during the live broadcast.
  • obtaining a target position includes: presenting a virtual space, the virtual space indicating multiple virtual device positions corresponding to multiple device positions of multiple acquisition devices, respectively, the multiple acquisition devices are associated with a live broadcast object, and a group of acquisition devices is at least a part of the multiple acquisition devices; and in response to an interactive operation of a first user in the virtual space, determining a target position corresponding to the interactive operation.
  • determining a target media item includes: obtaining a set of directions of a set of acquisition devices, wherein, for a first acquisition device in the set of acquisition devices, a first direction of the first acquisition device points from a first position of the first acquisition device to a position of a live broadcast object; determining a set of weights based on a distance between the target position and a set of straight lines defined by the set of directions; and generating a target media item based on the set of weights and a set of media items.
  • a media item includes at least any one of an image and a video
  • generating a target media item includes: determining, for a target pixel in the target media item, a set of pixels in a set of media items that correspond to the target pixel; and determining color information of the target pixel based on a set of weights and color information of a set of pixels.
  • the target position is located in a one-dimensional space defined by the first position of the first acquisition device and the position of the live broadcast object.
  • a group of collection devices further includes a second collection device, and the target position is located in a two-dimensional space defined by a first position of the first collection device, a second position of the second collection device, and a position of the live object.
  • a group of acquisition devices further includes a second acquisition device and a third acquisition device, and the target position is located in a three-dimensional space defined by a first position of the first acquisition device, a second position of the second acquisition device, a third position of the third acquisition device, and a position of the live broadcast object.
  • obtaining a target location includes: determining the target location based on the input of a first user in a live broadcast application, the first user is a viewer in the live broadcast application, the live broadcast object is a second user in the live broadcast application, and the second user is at least any one of a host or a guest, and the method further includes: presenting a target media item at a first client device.
  • computing device 1000 is in the form of a general-purpose computing device.
  • Components of computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060.
  • Processing unit 1010 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 1000.

Landscapes

  • Engineering & Computer Science (AREA)
  • Databases & Information Systems (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computer Graphics (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)

Abstract

提供了用于管理媒体项的方法、装置、设备和介质。在一种方法中,获取用于观看直播对象的目标位置。基于目标位置,获取一组媒体项,一组媒体项分别来自与直播对象相关联的一组采集设备。基于一组媒体项来确定与目标位置相匹配的目标媒体项。利用本公开的示例实现方式,确定的媒体项是利用实时采集的一组媒体项来生成的,因而可以包括更为准确和丰富的视觉信息。以此方式,用户在直播过程中可以在任何期望的视点位置处观看直播。

Description

用于管理媒体项的方法、装置、设备和介质
本申请要求2024年3月29日递交的、标题为“用于管理媒体项的方法、装置、设备和介质”、申请号为2024103847032的中国发明专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
技术领域
本公开的示例性实现方式总体涉及媒体管理,特别地涉及用于在直播应用中管理媒体项的方法、装置、设备和计算机可读存储介质。
背景技术
随着计算机技术和网络技术的发展,目前已经开发出了多种直播应用。可以经由直播应用来实时地向大量观众提供直播服务。进一步,已经提出了基于多位置的直播技术方案,例如,在直播环境中可以部署多个采集设备,并且多个采集设备可以向观众提供多视点的直播视频。然而,由于多个采集设备的部署位置是固定的,观众并不能选择这些固定位置之外的观看位置。
发明内容
在本公开的第一方面,提供了一种用于管理媒体项的方法。在该方法中,确定用于观看直播对象的目标位置。基于目标位置,获取一组媒体项,一组媒体项分别来自与直播对象相关联的一组采集设备。基于一组媒体项来确定与目标位置相匹配的目标媒体项。
在本公开的第二方面,提供了一种用于管理媒体项的装置。该装置包括:位置获取模块,被配置用于获取用于观看直播对象的目标位置;媒体获取模块,被配置用于基于目标位置,获取一组媒体项,一组媒体项分别来自与直播对象相关联的一组采集设备;以及确定模块,被配置用于基于一组媒体项来确定与目标位置相匹配的目标媒体项。
在本公开的第三方面,提供了一种电子设备。该电子设备包括:至少一个处理单元;以及至少一个存储器,至少一个存储器被耦合到至少一个处理单元并且存储用于由至少一个处理单元执行的指令,指令在由至少一个处理单元执行时使电子设备执行根据本公开第一方面的方法。
在本公开的第四方面,提供了一种计算机可读存储介质,其上存储有计算机程序,计算机程序在被处理器执行时使处理器实现根据本公开第一方面的方法。
应当理解,本内容部分中所描述的内容并非旨在限定本公开的实现方式的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
附图说明
在下文中,结合附图并参考以下详细说明,本公开各实现方式的上述和其他特征、优点及方面将变得更加明显。在附图中,相同或相似的附图标注表示相同或相似的元素,其中:
图1A示出了根据本公开的一个示例性实现方式的应用环境的框图;
图1B示出了根据本公开的一个示例性实现方式的多个采集设备的部署位置的框图;
图2示出了根据本公开的一些实现方式的用于管理媒体项的框图;
图3示出了根据本公开的一些实现方式的用于确定目标位置的框图;
图4示出了根据本公开的一些实现方式的用于确定一组采集设备的框图;
图5示出了根据本公开的一些实现方式的用于确定一组采集设备的框图;
图6示出了根据本公开的一些实现方式的用于确定目标媒体项中的像素的框图;
图7示出了根据本公开的一些实现方式的用于缩放目标媒体项的框图;
图8示出了根据本公开的一些实现方式的用于管理媒体项的方法的流程图;
图9示出了根据本公开的一些实现方式的用于管理媒体项的装置的框图;以及
图10示出了能够实施本公开的多个实现方式的设备的框图。
具体实施方式
下面将参照附图更详细地描述本公开的实现方式。虽然附图中示出了本公开的某些实现方式,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实现方式,相反,提供这些实现方式是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实现方式仅用于示例性作用,并非用于限制本公开的保护范围。
在本公开的实现方式的描述中,术语“包括”及其类似用语应当理解为开放性包含,即“包括但不限于”。术语“基于”应当理解为“至少部分地基于”。术语“一个实现方式”或“该实现方式”应当理解为“至少一个实现方式”。术语“一些实现方式”应当理解为“至少一些实现方式”。下文还可能包括其他明确的和隐含的定义。如本文中所使用的,术语“模型”可以表示各个数据之间的关联关系。例如,可以基于目前已知的和/或将在未来开发的多种技术方案来获取上述关联关系。
可以理解的是,本技术方案所涉及的数据(包括但不限于数据本身、数据的获取或使用)应当遵循相应法律法规及相关规定的要求。
可以理解的是,在使用本公开各实施例公开的技术方案之前,均应当根据相关法律法规通过适当的方式对本公开所涉及个人信息的类型、使用范围、使用场景等告知用户并获得用户的授权。
例如,在响应于接收到用户的主动请求时,向用户发送提示信息,以明确地提示用户,其请求执行的操作将需要获取和使用到用户的个人信息。从而,使得用户可以根据提示信息来自主地选择是否向执行本公开技术方案的操作的电子设备、应用程序、服务器或存储介质等软件或硬件提供个人信息。
作为一种可选的但非限制性的实现方式,响应于接收到用户的主动请求,向用户发送提示信息的方式,例如可以是弹出窗口的方式,弹出窗口中可以以文字的方式呈现提示信息。此外,弹出窗口中还可以承载供用户选择“同意”或“不同意”向电子设备提供个人信息的选择控件。
可以理解的是,上述通知和获取用户授权过程仅是示意性的,不对本公开的实现方式构成限定,其他满足相关法律法规的方式也可应用于本公开的实现方式中。
在此使用的术语“响应于”表示相应的事件发生或者条件得以满足的状态。将会理解,响应于该事件或者条件而被执行的后续动作的执行时机,与该事件发生或者条件成立的时间,二者之间未必是强关联的。例如,在某些情况下,后续动作可在事件发生或者条件成立时立即被执行;而在另一些情况下,后续动作可在事件发生或者条件成立后经过一段时间才被执行。
示例环境
目前已经开发出了多种直播应用。可以经由直播应用来实时地向大量观众提供直播服务。参见图1A描述根据本公开的一个示例实现方式的应用环境,该图1A示出了根据本公开的一个示例性实现方式的应用环境的框图100A。如图1A所示,经由网络140,直播应用的第一用户110的第一客户端设备112可以访问服务器设备130,并且第二用户120的利用第二客户端设备122可以访问服务器设备130。
在此,第一用户110例如可以是直播应用中的观众,并且第二用户120例如可以是直播应用中的主播和/或嘉宾。根据本公开的一个示例实现方式,第二用户120可以展示直播对象(例如,某个商品)并且介绍该商品的多方面信息。备选地和/或附加地,第二用户120可以作为直播对象并且直播才艺表演(例如,唱歌、跳舞,等)。为了便于描述,在下文中,将仅以第二用户作为直播对象为示例描述直播过程。
如图1A所示,可以在第二用户120周围的各个位置处部署多个采集设备124、126、…、128。第二客户端设备122可以将来自多个采集设备的多个媒体项经由网络140发送至服务器设备130,以便服务器设备130向一个或者多个第一用户110的第一客户端设备112提供各个媒体项。
参见图1B描述有关部署采集设备的更多细节,该图1B示出了根据本公开的一个示例性实现方式的多个采集设备的部署位置的框图100B。如图1B所示,可以以第二用户120的位置作为原点O来建立xyz坐标系,并且xyz轴的方向如图1B所示。可以在点P1、P2、P3、P4、P5处分别部署多个采集设备,并且每个采集设备的方向为朝向原点。例如,P1处的采集设备的方向为P1->O所示方向,P2处的采集设备的方向为P2->O所示方向,等等。以此方式,多个采集设备可以向观众提供多角度的直播视频。此时,第一用户110可以选择观看来自某个(或者某些)采集设备的媒体项。然而,由于多个采集设备的部署位置是固定的,观众并不能选择这些固定位置之外的观看位置。
目前已经提出了基于全景图的漫游技术方案,可以向第一用户110提供基于多个媒体项生成的全景图像,并且允许用户选择不同的视点位置。然而,全景图像的视觉效果并不理想,并且在某些位置处呈现的图像可能存在严重变形。此时,期望可以向用户提供在任意视点位置处的视频,并且期望可以缓解该视频中的图像变形问题并且提供更加准确和逼真的视频。
管理媒体项的概要
为了至少部分地解决现有技术中的不足,根据本公开的一个示例性实现方式,提出了一种用于管理媒体项的方法。参见图2描述根据本公开的一个示例性实现方式的概要,该图2示出了根据本公开的一些实现方式的用于管理媒体项的框图200。如图2所示,可以获取用于观看直播对象的目标位置210,该目标位置210还可以被称为用户的观看位置或者视点位置。
根据本公开的一个示例实现方式,媒体项可以包括图像和/或视频中的至少任一项。在下文中,将仅以视频为示例描述有关媒体项管理的过程。以此方式,可以支持向用户提供更为丰富的媒体数据。基于所述目标位置,可以获取一组媒体项230,一组媒体项分别来自与直播对象相关联的一组采集设备。一组采集设备可以是多个采集设备中的至少一部分采集设备。多个采集设备(例如,采集设备124、126、…、以及128)与目标对象相关联,并且被部署在第二用户(也即,目标对象)的预定空间范围内。在此,预定空间范围可以表示该直播对象周围的预先指定的空间范围,例如,该直播对象所在的房间,或者该直播对象周围的距离为3米(或者其他数值)内的空间范围。
例如,可以基于多个采集设备的位置与目标位置210之间的距离,或者基于多个采集设备的方向所在直线与目标位置210之间的距离,来确定一组采集设备。可以利用一组媒体项来确定与目标位置210相匹配的目标媒体项240。进一步,可以向第一客户端设备112提供目标媒体项240,以便第一用户110可以观看在任意视点处的直播视频。
利用本公开的示例性实现方式,可以向用户提供更为灵活的控制方式,用户可以选择任意期望的视点位置。进一步,可以利用来自多个采集设备中的相关采集设备的媒体项,来生成该视点位置的视频。利用本公开的示例实现方式,所确定的媒体项是利用实时采集的一组媒体项来生成的,因而可以包括更为准确和丰富的视觉信息。以此方式,用户在直播过程中可以在任何期望的视点位置处观看直播。
管理媒体项的详细过程
已经描述了根据本公开的一个示例实现方式的概要,在下文中,参见附图描述有关管理媒体项的更多细节。根据本公开的一个示例实现方式,目标位置可以指示第一用户将要观看直播应用的第二用户的位置(也即,第一用户的视点位置)。例如,可以在第一客户端设备112处提供选择页面,以便第一用户110经由该选择页面来选择期望的视点位置。此时,可以基于第一用户的输入来确定目标位置。
在获取目标位置的过程中,可以在第一客户端设备112处向第一用户110呈现虚拟空间,虚拟空间可以指示分别对应于多个采集设备的多个设备位置的多个虚拟设备位置。进一步,第一用户可以与虚拟空间进行交互,以便基于交互操作来确定与交互操作相对应的目标位置。
参见图3描述确定目标位置的更多细节,该图3示出了根据本公开的一些实现方式的用于确定目标位置的框图300。如图3所示,虚拟空间可以在三维坐标中呈现多个采集设备的位置以及第二用户120的位置。例如,第二用户120可以位于坐标原点O处,并且多个采集设备可以分别位于x、y、z坐标轴处的多个点P1、P2、P3、P4和P5。进一步,用户可以执行交互操作320,以便选择期望观看直播视频的目标位置310(例如,以坐标(x0,y0,z0)表示)。例如,第一用户可以通过单击、双击、拖拽等方式来确定目标位置310。利用本公开的示例实现方式,可以以可视化方式向用户提供选择页面,以便用户以更为方便并且准确的方式来选择期望的视点位置。
应当理解,尽管图3示意性示出了多个采集设备分别被部署在各个坐标轴处的示例。备选地和/或附加地,采集设备可以被部署在坐标轴以外的其他位置,例如,可以在图3所示球面的任意位置处沿着朝向原点的方向部署采集设备。以此方式,可以提供来自更多角度的媒体项,从而提高生成目标媒体项的准确性。备选地和/或附加地,多个采集设备与原点O之间的距离可以有所不同,以此方式,可以按照用户需求来呈现远距离画面和/或近距离画面。
根据本公开的一个示例实现方式,第二客户端设备可以向服务器设备传输分别来自多个采集设备的多个媒体项。例如,可以单独向服务器设备传输多个媒体项。备选地和/或附加地,可以将多个媒体项执行编码操作以便以单一数据流的方式来传输多个媒体项。以此方式,可以降低传输过程所涉及的带宽、降低传输延迟并且提高传输效率。
根据本公开的一个示例实现方式,直播应用的服务器设备可以基于目标位置来确定的向第一客户端设备提供哪些媒体项。以此方式,第一客户端设备可以直接从服务器设备接收所需的媒体项,由此可以降低第一客户端设备的工作负载。
根据本公开的一个示例实现方式,服务器设备可以获取一组采集设备的一组方向。在此,针对一组采集设备中的第一采集设备,第一采集设备的第一方向为从第一采集设备的第一位置指向第二用户的位置。具体地,如图3所示,点P1至P5处的各个采集设备的方向均指向原点O。进一步,可以比较目标位置和各个采集设备的方向所定义的空间,进而确定一组采集设备。
在本公开的上下文中,可以利用通过采集设备的位置的向量来表示采集设备的方向。目标位置与采集设备的位置之间可能会存在多种空间关系:例如,目标位置可以位于由采集设备的位置和第二用户的位置所定义的一维空间(也即,直线)中,目标位置可以位于由两个采集设备的位置和第二用户的位置所定义的二维空间(也即,平面)中,目标位置可以位于由三个采集设备的位置和第二用户的位置所定义的三维空间(也即,体空间)中。
参见图4描述确定一组采集设备的更多细节,该图4示出了根据本公开的一些实现方式的用于确定一组采集设备的框图400。为了便于描述,假设在位置P1、P2和P5处分别部署了三个采集设备。P1处的采集设备(例如,称为第一采集设备)的方向为-x方向,P2处的采集设备(例如,称为第二采集设备)的方向为-y方向,并且P5处的采集设备(例如,称为第三采集设备)的方向为-z方向。
根据本公开的一个示例实现方式,假设目标位置位于位置410处,此时目标位置位于由第一采集设备的第一位置(也即点P1)和第二用户的位置(也即原点O)所定义的一维空间(也即,x轴)中。此时,一组媒体项可以包括由该第一采集设备所采集的媒体项。
应当理解,图4仅仅是示意性的,目标位置可以位于其他位置。假设目标位置位于y轴,则可以一组媒体项可以包括来自点P2处的采集设备的媒体项,等等。以此方式,可以通过比较目标位置以及由采集设备的位置和第二用户的位置所定义的直线的位置关系,来以简单并且准确的方式确定生成目标媒体项所需的数据。
根据本公开的一个示例实现方式,假设目标位置位于位置420处,此时目标位置位于由第一采集设备的第一位置(也即点P1)、第二采集设备的第二位置(也即点P2)和第二用户的位置(也即原点O)所定义的二维空间(也即,xoy平面)中。此时,一组媒体项可以包括由该第一采集设备所采集的媒体项和由该第二采集设备所采集的媒体项。
应当理解,图4仅仅是示意性的,目标位置可以位于其他位置。假设目标位置位于xoz平面,则可以一组媒体项可以包括来自点P1处的采集设备的媒体项以及来自点P5处的采集设备的媒体项,等等。以此方式,可以通过比较目标位置以及由两个采集设备的位置和第二用户的位置所定义的平面的位置关系,来以简单并且准确的方式确定生成目标媒体项所需的数据。
参见图5描述目标位置位于坐标平面以外的其他位置的更多示例,该图5示出了根据本公开的一些实现方式的用于确定一组采集设备的框图500。如图5所示,假设目标位置位于位置510处,此时位置510位于由第一采集设备的第一位置(也即点P1)、第二采集设备的第二位置(也即点P2)、第三采集设备的第三位置(也即点P5)、和第二用户的位置(也即原点O)所定义的三维空间(也即,由原点O、+x轴、+y轴和+z轴所定义的第一象限)中。此时,一组媒体项可以包括由该第一采集设备所采集的媒体项、由该第二采集设备所采集的媒体项、以及由该第三采集设备所采集的媒体项。
应当理解,图5仅仅是示意性的,目标位置可以位于其他位置。假设目标位置位于由原点O、-x轴、+y轴和+z轴所定义的空间中,则可以一组媒体项可以包括来自点P2、P3和P5处的采集设备的媒体项,等等。以此方式,可以通过比较目标位置以及由三个采集设备的位置和第二用户的位置所定义的平面的位置关系,来以简单并且准确的方式确定生成目标媒体项所需的数据。
根据本公开的一个示例实现方式,服务器设备可以基于目标位置和一组媒体项,来生成目标媒体项并且向各个观众的客户端设备传输所生成的媒体项。应当理解,在直播过程中可能会存在大量观众,在服务器设备处生成目标媒体项可能会导致服务器设备的工作负载过高。
根据本公开的一个示例实现方式,服务器设备可以向第一客户端设备传输所确定的一组媒体项。例如,可以单独向服务器设备传输各个媒体项。备选地和/或附加地,可以将各个媒体项执行编码操作以便以单一数据流的方式来传输各个媒体项。以此方式,可以降低传输过程所涉及的带宽、降低传输延迟并且提高传输效率。客户端设备可以接收来自服务器设备的一组媒体项,并且利用一组媒体项来生成目标媒体项。以此方式,可以降低服务器设备的工作负载,并且使得服务器设备可以关注于直播过程的其他计算处理过程。
根据本公开的一个示例实现方式,第一客户端设备可以基于目标位置与一组采集设备的一组方向之间的距离来确定一组权重,继而,可以基于一组权重和一组媒体项,确定目标媒体项。返回图4,假设目标位置位于图4中的位置410,此时服务器设备仅向第一客户端设备传输来自点P1处的采集设备的媒体项。此时,目标位置与x轴之间的距离为零,因而可以直接将来自点P1处的采集设备的媒体项作为目标媒体项。
假设目标位置位于图4中的位置420,此时服务器设备可以向第一客户端设备传输来自点P1和P2处的采集设备的两个媒体项。此时,可以基于位置420与x轴和y轴之间的比例,来确定相应的权重。假设位置420位于x轴和y轴的角平分线(也即,45度)处,则来自两个媒体项的贡献相等,由此可以基于加权求和的方式来确定目标媒体项的具体内容。
假设目标位置位于图5中的位置510,可以确定位置510与各个采集设备的方向(也即,+x轴、+y轴和+z轴)之间的距离,来确定相应的权重。具体地,假设位置510的坐标为(x0,y0,z0),则位置510与+x轴之间的距离416表示为位置510与+y轴之间的距离412表示为位置510与+z轴之间的距离414表示为以此方式,可以将确定各个媒体项对于最终的目标媒体项的贡献转化为数学运算过程,由此以更为简单并且准确的方式确定目标媒体项。
根据本公开的一个示例实现方式,可以按照上述距离之间的比例关系,确定各个媒体项的权重。具体地,在生成目标媒体项的过程中,针对目标媒体项中的目标像素,确定一组媒体项中的对应于目标像素的一组像素。进一步,可以基于一组权重和一组像素的颜色信息,确定目标像素的颜色信息。参见图6描述更多信息,该图6示出了根据本公开的一些实现方式的用于确定目标媒体项中的像素的框图600。
如图6所示,假设来自三个采集设备的媒体项分别表示为媒体项610(来自点P1处的采集设备的媒体项)、620(来自点P2处的采集设备的媒体项)和630(来自点P5处的采集设备的媒体项),在确定目标媒体项650中的像素652(例如,像素坐标为(i,j))的颜色的过程中,可以针对媒体项610中的像素612、媒体项620中的像素622、以及媒体项630中的像素632执行加权求和640。应当理解,在此像素612、622以及624的像素坐标均为(i,j),可以遍历每个像素位置,进而确定相应的颜色。例如,可以基于如下公式1来确定相应的颜色。
在上述公式中,color表示像素652的颜色数据,colorx表示像素612的颜色数据、colory表示像素622的颜色数据、以及colorz表示像素632的颜色数据。以此方式,可以基于简单的数学运算来确定目标媒体项中的各个像素的颜色数据。应当理解,上文的公式仅仅是示例性的,备选地和/或附加地,可以基于目前已知的和/或将在未来开发的多种技术方案,基于来自部署在不同位置和/或方向的采集设备的多个媒体项来生成具有期望位置和/或方向的视点的媒体项。
根据本公开的一个示例实现方式,当媒体项为图像时,可以基于上文描述的方式来利用一组图像生成最终的图像。当媒体项为视频时,可以利用具有相同时间戳的一组图像来生成与该时间戳相关联的图像。例如,可以利用时间点t的一组图像来生成时间点t的图像,可以利用时间点t+1的一组图像来生成时间点t+1的图像,继而可以按照时间顺序来呈现各个图像。
根据本公开的一个示例实现方式,在生成目标媒体项的过程中,可以基于目标位置与第二用户的位置之间的距离,缩放目标媒体项。参见图7描述有关缩放操作的更多细节,该图7示出了根据本公开的一些实现方式的用于缩放目标媒体项的框图700。如图7所示,假设目标位置位于位置710,也即,目标位置与原点O之间的距离小于点P2处的采集设备与原点O之间的距离,此时,可以放大目标媒体项中的图像内容。又例如,假设目标位置位于位置720,也即,目标位置与原点O之间的距离小于点P2处的采集设备与原点O之间的距离,此时,可以缩小目标媒体项中的图像内容。以此方式,可以支持用户以更为灵活的方式来选择视点位置,进而获得有关直播的更多信息。
应当理解,上文仅以示例方式描述了目标位置位于三维空间中的第一象限的处理过程,备选地和/或附加地,目标位置可以位于三维空间中的其他象限。可以基于各个点的位置之间的空间几何关系来确定相应的权重,在下文中将不再赘述。
应当理解,上文仅以示例方似乎描述了三个采集设备分别位于+x轴、+y轴和+z轴的情况,备选地和/或附加地,采集设备可以位于三维空间中的其他位置。此时,可以以第二用户的位置为原点来建立相应的坐标系,并且基于各个采集设备的位置与原点之间的空间位置关系,来确定各个采集设备的方向和位置的具体表达方式。进一步,可以按照上文描述的原理,从多个采集设备中确定一组采集设备,进而基于来自一组采集设备的媒体项,来确定最终的目标媒体项。
根据本公开的一个示例实现方式,第一用户可以是直播应用中的观众,直播对象可以是直播应用中的第二用户,例如主播或者嘉宾中的至少任一项。此时,普通观众可以选择任意的视点位置,来观看直播过程。利用本公开的示例实现方式,所生成的视频是利用实时采集的一组媒体项来生成的,因而可以包括更为准确和丰富的视觉信息。以此方式,用户在直播过程中可以在任何期望的视点位置处观看直播。
示例过程
图8示出了根据本公开的一些实现方式的用于管理媒体项的方法800的流程图。在框810处,获取用于观看直播对象的目标位置。在框820处,基于目标位置,获取一组媒体项,一组媒体项分别来自与直播对象相关联的一组采集设备。在框830处,基于一组媒体项来确定与目标位置相匹配的目标媒体项。
根据本公开的一个示例实现方式,获取目标位置包括:呈现虚拟空间,虚拟空间指示分别对应于多个采集设备的多个设备位置的多个虚拟设备位置,多个采集设备与直播对象相关联,并且一组采集设备是多个采集设备中的至少一部分采集设备;以及响应于第一用户在虚拟空间中的交互操作,确定与交互操作相对应的目标位置。
根据本公开的一个示例实现方式,确定目标媒体项包括:获取一组采集设备的一组方向,针对一组采集设备中的第一采集设备,第一采集设备的第一方向从第一采集设备的第一位置指向直播对象的位置;基于目标位置与一组方向所定义的一组直线之间的距离来确定一组权重;以及基于一组权重和一组媒体项,生成目标媒体项。
根据本公开的一个示例实现方式,媒体项包括图像和视频中的至少任一项,并且生成目标媒体项包括:针对目标媒体项中的目标像素,确定一组媒体项中的对应于目标像素的一组像素;以及基于一组权重和一组像素的颜色信息,确定目标像素的颜色信息。
根据本公开的一个示例实现方式,生成目标媒体项进一步包括:基于目标位置与直播对象的位置之间的距离,缩放目标媒体项。
根据本公开的一个示例实现方式,目标位置位于由一组采集设备的一组方向所定义的空间之内。
根据本公开的一个示例实现方式,目标位置位于由第一采集设备的第一位置和直播对象的位置所定义的一维空间中。
根据本公开的一个示例实现方式,一组采集设备进一步包括第二采集设备,并且目标位置位于由第一采集设备的第一位置、第二采集设备的第二位置、以及和直播对象的位置所定义的二维空间中。
根据本公开的一个示例实现方式,一组采集设备进一步包括第二采集设备以及第三采集设备,并且目标位置位于由第一采集设备的第一位置、第二采集设备的第二位置、第三采集设备的第三位置、以及和直播对象的位置所定义的三维空间中。
根据本公开的一个示例实现方式,该方法在第一用户的第一客户端设备处被执行,并且一组媒体项是由直播应用的服务器设备基于目标位置来确定的。
根据本公开的一个示例实现方式,获取目标位置包括:基于直播应用中的第一用户的输入来确定目标位置,第一用户是直播应用中的观众,直播对象是直播应用中的第二用户,并且第二用户是主播或者嘉宾中的至少任一项,并且该方法进一步包括:在第一客户端设备处呈现目标媒体项。
示例装置和设备
图9示出了根据本公开的一些实现方式的用于管理媒体项的装置900的框图。该装置包括:位置获取模块910,被配置用于获取用于观看直播对象的目标位置;媒体获取模块920,被配置用于基于目标位置,获取一组媒体项,一组媒体项分别来自与直播对象相关联的一组采集设备;以及确定模块930,被配置用于基于一组媒体项来确定与目标位置相匹配的目标媒体项。
根据本公开的一个示例实现方式,位置获取模块包括:呈现模块,被配置用于呈现虚拟空间,虚拟空间指示分别对应于多个采集设备的多个设备位置的多个虚拟设备位置,多个采集设备与直播对象相关联,并且一组采集设备是多个采集设备中的至少一部分采集设备;以及位置确定模块,被配置用于响应于第一用户在虚拟空间中的交互操作,确定与交互操作相对应的目标位置。
根据本公开的一个示例实现方式,确定模块包括:方向获取模块,被配置用于获取一组采集设备的一组方向,针对一组采集设备中的第一采集设备,第一采集设备的第一方向从第一采集设备的第一位置指向直播对象的位置;权重确定模块,被配置用于基于目标位置与一组方向所定义的一组直线之间的距离来确定一组权重;以及生成模块,被配置用于基于一组权重和一组媒体项,生成目标媒体项。
根据本公开的一个示例实现方式,媒体项包括图像和视频中的至少任一项,并且生成模块包括:像素确定模块,被配置用于针对目标媒体项中的目标像素,确定一组媒体项中的对应于目标像素的一组像素;以及颜色确定模块,被配置用于基于一组权重和一组像素的颜色信息,确定目标像素的颜色信息。
根据本公开的一个示例实现方式,生成模块进一步包括:缩放模块,被配置用于基于目标位置与直播对象的位置之间的距离,缩放目标媒体项。
根据本公开的一个示例实现方式,目标位置位于由一组采集设备的一组方向所定义的空间之内。
根据本公开的一个示例实现方式,目标位置位于由第一采集设备的第一位置和直播对象的位置所定义的一维空间中。
根据本公开的一个示例实现方式,一组采集设备进一步包括第二采集设备,并且目标位置位于由第一采集设备的第一位置、第二采集设备的第二位置、以及和直播对象的位置所定义的二维空间中。
根据本公开的一个示例实现方式,一组采集设备进一步包括第二采集设备以及第三采集设备,并且目标位置位于由第一采集设备的第一位置、第二采集设备的第二位置、第三采集设备的第三位置、以及和直播对象的位置所定义的三维空间中。
根据本公开的一个示例实现方式,装置在第一用户的第一客户端设备处被实现,并且一组媒体项是由直播应用的服务器设备基于目标位置来确定的。
根据本公开的一个示例实现方式,位置获取模块进一步被配置用于:基于直播应用中的第一用户的输入来确定目标位置,第一用户是直播应用中的观众,直播对象是直播应用中的第二用户,并且第二用户是主播或者嘉宾中的至少任一项,并且装置进一步包括:呈现模块,被配置用于在第一客户端设备处呈现目标媒体项。
图10示出了能够实施本公开的多个实现方式的设备1000的框图。应当理解,图10所示出的计算设备1000仅仅是示例性的,而不应当构成对本文所描述的实现方式的功能和范围的任何限制。图10所示出的计算设备1000可以用于实现上文描述的方法。
如图10所示,计算设备1000是通用计算设备的形式。计算设备1000的组件可以包括但不限于一个或多个处理器或处理单元1010、存储器1020、存储设备1030、一个或多个通信单元1040、一个或多个输入设备1050以及一个或多个输出设备1060。处理单元1010可以是实际或虚拟处理器并且能够根据存储器1020中存储的程序来执行各种处理。在多处理器系统中,多个处理单元并行执行计算机可执行指令,以提高计算设备1000的并行处理能力。
计算设备1000通常包括多个计算机存储介质。这样的介质可以是计算设备1000可访问的任何可以获得的介质,包括但不限于易失性和非易失性介质、可拆卸和不可拆卸介质。存储器1020可以是易失性存储器(例如寄存器、高速缓存、随机访问存储器(RAM))、非易失性存储器(例如,只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、闪存)或它们的某种组合。存储设备1030可以是可拆卸或不可拆卸的介质,并且可以包括机器可读介质,诸如闪存驱动、磁盘或者任何其他介质,其可以能够用于存储信息和/或数据(例如用于训练的训练数据)并且可以在计算设备1000内被访问。
计算设备1000可以进一步包括另外的可拆卸/不可拆卸、易失性/非易失性存储介质。尽管未在图10中示出,可以提供用于从可拆卸、非易失性磁盘(例如“软盘”)进行读取或写入的磁盘驱动和用于从可拆卸、非易失性光盘进行读取或写入的光盘驱动。在这些情况中,每个驱动可以由一个或多个数据介质接口被连接至总线(未示出)。存储器1020可以包括计算机程序产品1025,其具有一个或多个程序模块,这些程序模块被配置为执行本公开的各种实现方式的各种方法或动作。
通信单元1040实现通过通信介质与其他计算设备进行通信。附加地,计算设备1000的组件的功能可以以单个计算集群或多个计算机器来实现,这些计算机器能够通过通信连接进行通信。因此,计算设备1000可以使用与一个或多个其他服务器、网络个人计算机(PC)或者另一个网络节点的逻辑连接来在联网环境中进行操作。
输入设备1050可以是一个或多个输入设备,例如鼠标、键盘、追踪球等。输出设备1060可以是一个或多个输出设备,例如显示器、扬声器、打印机等。计算设备1000还可以根据需要通过通信单元1040与一个或多个外部设备(未示出)进行通信,外部设备诸如存储设备、显示设备等,与一个或多个使得用户与计算设备1000交互的设备进行通信,或者与使得计算设备1000与一个或多个其他计算设备通信的任何设备(例如,网卡、调制解调器等)进行通信。这样的通信可以经由输入/输出(I/O)接口(未示出)来执行。
根据本公开的示例性实现方式,提供了一种计算机可读存储介质,其上存储有计算机可执行指令,其中计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,还提供了一种计算机程序产品,计算机程序产品被有形地存储在非瞬态计算机可读介质上并且包括计算机可执行指令,而计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,提供了一种计算机程序产品,其上存储有计算机程序,程序被处理器执行时实现上文描述的方法。
这里参照根据本公开实现的方法、装置、设备和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理单元,从而生产出一种机器,使得这些指令在通过计算机或其他可编程数据处理装置的处理单元执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
可以把计算机可读程序指令加载到计算机、其他可编程数据处理装置、或其他设备上,使得在计算机、其他可编程数据处理装置或其他设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其他可编程数据处理装置、或其他设备上执行的指令实现流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实现的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或指令的一部分,模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实现,上述说明是示例性的,并非穷尽性的,并且也不限于所公开的各实现。在不偏离所说明的各实现的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实现的原理、实际应用或对市场中的技术的改进,或者使本技术领域的其他普通技术人员能理解本文公开的各个实现方式。

Claims (14)

  1. 一种用于管理媒体项的方法,包括:
    获取用于观看直播对象的目标位置;
    基于所述目标位置,获取一组媒体项,所述一组媒体项分别来自与所述直播对象相关联的一组采集设备;以及
    基于所述一组媒体项来确定与所述目标位置相匹配的目标媒体项。
  2. 根据权利要求1所述的方法,其中获取所述目标位置包括:
    呈现虚拟空间,所述虚拟空间指示分别对应于多个采集设备的多个设备位置的多个虚拟设备位置,所述多个采集设备与所述直播对象相关联,并且所述一组采集设备是所述多个采集设备中的至少一部分采集设备;以及
    响应于所述第一用户在所述虚拟空间中的交互操作,确定与所述交互操作相对应的所述目标位置。
  3. 根据权利要求1所述的方法,其中确定所述目标媒体项包括:
    获取所述一组采集设备的一组方向,针对所述一组采集设备中的第一采集设备,所述第一采集设备的第一方向从所述第一采集设备的第一位置指向所述直播对象的位置;
    基于所述目标位置与所述一组方向所定义的一组直线之间的距离来确定一组权重;以及
    基于所述一组权重和所述一组媒体项,生成所述目标媒体项。
  4. 根据权利要求3所述的方法,其中所述媒体项包括图像和视频中的至少任一项,并且生成所述目标媒体项包括:
    针对所述目标媒体项中的目标像素,确定所述一组媒体项中的对应于所述目标像素的一组像素;以及
    基于所述一组权重和所述一组像素的颜色信息,确定所述目标像素的颜色信息。
  5. 根据权利要求3所述的方法,其中生成所述目标媒体项进一步包括:基于所述目标位置与所述直播对象的位置之间的距离,缩放所述目标媒体项。
  6. 根据权利要求3所述的方法,其中所述目标位置位于由所述一组采集设备的所述一组方向所定义的空间之内。
  7. 根据权利要求6所述的方法,其中所述目标位置位于由所述第一采集设备的第一位置和所述直播对象的位置所定义的一维空间中。
  8. 根据权利要求6所述的方法,其中所述一组采集设备进一步包括第二采集设备,并且所述目标位置位于由所述第一采集设备的第一位置、所述第二采集设备的第二位置、以及和所述直播对象的位置所定义的二维空间中。
  9. 根据权利要求6所述的方法,其中所述一组采集设备进一步包括第二采集设备以及第三采集设备,并且所述目标位置位于由所述第一采集设备的第一位置、所述第二采集设备的第二位置、所述第三采集设备的第三位置、以及和所述直播对象的位置所定义的三维空间中。
  10. 根据权利要求1所述的方法,其中所述方法在所述第一用户的第一客户端设备处被执行,并且所述一组媒体项是由所述直播应用的服务器设备基于所述目标位置来确定的。
  11. 根据权利要求10所述的方法,其中获取所述目标位置包括:基于所述直播应用中的第一用户的输入来确定所述目标位置,所述第一用户是所述直播应用中的观众,所述直播对象是所述直播应用中的第二用户,并且所述第二用户是主播或者嘉宾中的至少任一项,并且所述方法进一步包括:在所述第一客户端设备处呈现所述目标媒体项。
  12. 一种用于管理媒体项的装置,包括:
    位置获取模块,被配置用于获取用于观看直播对象的目标位置;
    媒体获取模块,被配置用于基于所述目标位置,获取一组媒体项,一组媒体项分别来自与直播对象相关联的一组采集设备;以及
    确定模块,被配置用于基于所述一组媒体项来确定与所述目标位置相匹配的目标媒体项。
  13. 一种电子设备,包括:
    至少一个处理单元;以及
    至少一个存储器,所述至少一个存储器被耦合到所述至少一个处理单元并且存储用于由所述至少一个处理单元执行的指令,所述指令在由所述至少一个处理单元执行时使所述电子设备执行根据权利要求1至11中任一项所述的方法。
  14. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序在被处理器执行时使所述处理器实现根据权利要求1至11中任一项所述的方法。
PCT/CN2025/084840 2024-03-29 2025-03-25 用于管理媒体项的方法、装置、设备和介质 Pending WO2025201358A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410384703.2A CN120730085A (zh) 2024-03-29 2024-03-29 用于管理媒体项的方法、装置、设备和介质
CN202410384703.2 2024-03-29

Publications (1)

Publication Number Publication Date
WO2025201358A1 true WO2025201358A1 (zh) 2025-10-02

Family

ID=97165836

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/084840 Pending WO2025201358A1 (zh) 2024-03-29 2025-03-25 用于管理媒体项的方法、装置、设备和介质

Country Status (2)

Country Link
CN (1) CN120730085A (zh)
WO (1) WO2025201358A1 (zh)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101321299A (zh) * 2007-06-04 2008-12-10 华为技术有限公司 视差生成方法、生成单元以及三维视频生成方法及装置
JP2020043467A (ja) * 2018-09-11 2020-03-19 株式会社アクセル 画像処理装置、画像処理方法、及び画像処理プログラム
JP2020127211A (ja) * 2020-03-31 2020-08-20 株式会社バーチャルキャスト 3次元コンテンツ配信システム、3次元コンテンツ配信方法、コンピュータプログラム
CN114092315A (zh) * 2020-08-24 2022-02-25 阿里巴巴集团控股有限公司 重建图像的方法、装置、计算机可读存储介质和处理器

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101321299A (zh) * 2007-06-04 2008-12-10 华为技术有限公司 视差生成方法、生成单元以及三维视频生成方法及装置
JP2020043467A (ja) * 2018-09-11 2020-03-19 株式会社アクセル 画像処理装置、画像処理方法、及び画像処理プログラム
JP2020127211A (ja) * 2020-03-31 2020-08-20 株式会社バーチャルキャスト 3次元コンテンツ配信システム、3次元コンテンツ配信方法、コンピュータプログラム
CN114092315A (zh) * 2020-08-24 2022-02-25 阿里巴巴集团控股有限公司 重建图像的方法、装置、计算机可读存储介质和处理器

Also Published As

Publication number Publication date
CN120730085A (zh) 2025-09-30

Similar Documents

Publication Publication Date Title
CN107113381B (zh) 时空局部变形及接缝查找的容差视频拼接方法、装置及计算机可读介质
Guan et al. DeepMix: mobility-aware, lightweight, and hybrid 3D object detection for headsets
US20170301110A1 (en) Producing three-dimensional representation based on images of an object
US9697581B2 (en) Image processing apparatus and image processing method
JP2018026064A (ja) 画像処理装置、画像処理方法、システム
WO2024104248A1 (zh) 虚拟全景图的渲染方法、装置、设备及存储介质
CN111402136B (zh) 全景图生成方法、装置、计算机可读存储介质及电子设备
US10600202B2 (en) Information processing device and method, and program
WO2023005170A1 (zh) 全景视频的生成方法和装置
Park et al. InstantXR: Instant XR environment on the web using hybrid rendering of cloud-based NeRF with 3d assets
CN115115971A (zh) 处理图像以定位新颖对象
WO2025097814A1 (zh) 一种新视点图像合成方法、系统、电子设备和存储介质
JP5378883B2 (ja) 画像処理装置および画像処理方法
WO2025201358A1 (zh) 用于管理媒体项的方法、装置、设备和介质
CN114339120A (zh) 沉浸式视频会议系统
CN117635886B (zh) 内容展示方法、装置及计算机可读存储介质
US20240193824A1 (en) Computing device and method for realistic visualization of digital human
WO2025213834A1 (zh) 绘制三维场景的方法、装置、设备和介质
WO2025139909A1 (zh) 直播交互的方法、装置、设备和存储介质
CN108920598B (zh) 全景图浏览方法、装置、终端设备、服务器及存储介质
CN114089836B (zh) 标注方法、终端、服务器和存储介质
WO2023160072A1 (zh) 增强现实ar场景中的人机交互方法、装置和电子设备
CN115665361A (zh) 虚拟环境中的视频融合方法和在线视频会议通信方法
CN113837978A (zh) 图像合成方法、装置、终端设备以及可读存储介质
JP2022114626A (ja) 情報処理装置、情報処理方法、およびプログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25777587

Country of ref document: EP

Kind code of ref document: A1