WO2025228337A1 - 生成媒体内容的方法、装置、设备和存储介质 - Google Patents

生成媒体内容的方法、装置、设备和存储介质

Info

Publication number
WO2025228337A1
WO2025228337A1 PCT/CN2025/091825 CN2025091825W WO2025228337A1 WO 2025228337 A1 WO2025228337 A1 WO 2025228337A1 CN 2025091825 W CN2025091825 W CN 2025091825W WO 2025228337 A1 WO2025228337 A1 WO 2025228337A1
Authority
WO
WIPO (PCT)
Prior art keywords
target
media
attribute
prompt information
prompt
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/091825
Other languages
English (en)
French (fr)
Inventor
张伟
唐娜
查心怡
张楠
马倩
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Publication of WO2025228337A1 publication Critical patent/WO2025228337A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/048Interaction techniques based on graphical user interfaces [GUI]
    • G06F3/0481Interaction techniques based on graphical user interfaces [GUI] based on specific properties of the displayed interaction object or a metaphor-based environment, e.g. interaction with desktop elements like windows or icons, or assisted by a cursor's changing behaviour or appearance
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/048Interaction techniques based on graphical user interfaces [GUI]
    • G06F3/0484Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range
    • G06F3/04845Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range for image manipulation, e.g. dragging, rotation, expansion or change of colour
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • H04N21/44016Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving splicing one content stream with another content stream, e.g. for substituting a video clip
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/488Data services, e.g. news ticker
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/488Data services, e.g. news ticker
    • H04N21/4882Data services, e.g. news ticker for displaying messages, e.g. warnings, reminders
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/85Assembly of content; Generation of multimedia applications
    • H04N21/854Content authoring

Definitions

  • the exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for generating media content.
  • AI artificial intelligence
  • a method for generating media content includes: providing a set of candidate media samples in a target interface, the set of candidate media samples corresponding to a set of preset prompt messages; based on the selection of a target media sample from the set of candidate media samples, presenting a first prompt message corresponding to the target media sample in the target interface; presenting an attribute information set of the target media sample regarding multiple preset attribute items in the target interface, the multiple preset attribute items being used to describe different aspects of the media content to be generated; modifying the first prompt message to a second prompt message corresponding to the edited second attribute information based on editing the first attribute information in the attribute information set; and generating target media content based on the second prompt message.
  • an apparatus for generating media content includes: a sample providing module configured to provide a set of candidate media samples in a target interface, the set of candidate media samples corresponding to a set of preset prompt messages; a first presentation module configured to present, in the target interface, first prompt messages corresponding to a target media sample based on the selection of a target media sample from the set of candidate media samples; a second presentation module configured to present, in the target interface, a set of attribute information of the target media sample regarding multiple preset attribute items, the multiple preset attribute items being used to describe different aspects of the media content to be generated; an information processing module configured to modify the first prompt messages to second prompt messages corresponding to the edited second attribute information based on editing the first attribute information in the attribute information set; and a content generation module configured to generate target media content based on the second prompt messages.
  • an electronic device in a third aspect of this disclosure, includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
  • a computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
  • Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented
  • FIGS. 2A to 2D illustrate example interfaces according to some embodiments of the present disclosure
  • Figure 3 illustrates a flowchart of an example process for generating media content according to some embodiments of the present disclosure
  • Figure 4 shows a schematic structural block diagram of an example media content generation apparatus according to some embodiments of the present disclosure.
  • Figure 5 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure.
  • the term “comprising” and similar terms should be understood as open-ended inclusion, i.e., “including but not limited to”.
  • the term “based on” should be understood as “at least partially based on”.
  • the term “one embodiment” or “the embodiment” should be understood as “at least one embodiment”.
  • the term “some embodiments” should be understood as “at least some embodiments”.
  • Other explicit and implicit definitions may also be included below.
  • the terms “first”, “second”, etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
  • the embodiments of this disclosure may involve user data, data acquisition, and/or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and/or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
  • any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon.
  • a user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
  • a set of candidate media samples can be provided in a target interface, each set corresponding to a set of preset prompts; based on the selection of a target media sample from the set of candidate media samples, a first prompt corresponding to the target media sample is presented in the target interface; an attribute information set of the target media sample regarding multiple preset attribute items is presented in the target interface, the multiple preset attribute items being used to describe different aspects of the media content to be generated; based on editing the first attribute information in the attribute information set, the first prompt is modified to a second prompt corresponding to the edited second attribute information; and based on the second prompt, the target media content is generated.
  • the embodiments of this disclosure can enable users to refine the prompt information by setting preset media samples and corresponding attribute adjustments, thereby improving the accuracy of the prompt information and the efficiency of editing the prompt information, and thus improving the quality of the generated media content.
  • Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented.
  • the example environment 100 may include an electronic device 110.
  • electronic device 110 can run an application 120 that supports user interface interaction.
  • Application 120 can be any suitable type of application for user interface interaction, examples of which may include, but are not limited to, video applications, editing applications, or other suitable applications.
  • User 140 can interact with application 120 via electronic device 110 and/or its attached devices.
  • electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.
  • electronic device 110 communicates with server 130 to provide services to application 120.
  • Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR/AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio/video players, digital cameras/camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof.
  • electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
  • Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.
  • Server 130 may include, for example, computing systems/servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc.
  • Server 130 can provide backend services for applications 120 that support virtual scenarios in electronic devices 110.
  • a communication connection can be established between server 130 and electronic device 110.
  • This communication connection can be established via wired or wireless means.
  • the communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect.
  • server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.
  • Video content will be used as an example of media content in the description of the example media content generation process. It should be understood that the generation process of the present disclosure can also be applied to other types of media content, such as images, audio, etc.
  • FIGS 2A to 2D illustrate example interfaces 200A to 200D according to some embodiments of the present disclosure. Interfaces 200A to 200D can be provided by the electronic device 110 shown in Figure 1.
  • the electronic device 110 presents a target interface based on a request to generate media content.
  • FIG2A illustrates an example target interface according to some embodiments of the present disclosure. As shown in FIG2A, the target interface may include an input control 212 for inputting prompt information.
  • the user can input prompts (e.g., descriptive text) about the video content to be generated in the input control.
  • the electronic device 110 can determine the video content to be generated based on the received prompts.
  • the electronic device 110 can also acquire at least one material for generating video content via the interface 200A.
  • a user can upload video material, image material, audio material, etc., based on the control 210.
  • the user can also input text material via the control 212, such as subtitles for the text content to be generated.
  • the electronic device 110 may accept a user's selection of the generation entry 216 and generate media content based on the prompts and materials received in the interface 200A.
  • the electronic device 110 may utilize generative artificial intelligence technology to generate the corresponding media content, and this disclosure is not intended to limit the specific generation process of this media content.
  • the electronic device 110 can also support users in more efficiently editing prompts used to generate media content. As shown in FIG2A, the electronic device 110 is associated with input control 212 to provide entry point 214. Upon receiving a trigger on entry point 214, the electronic device 110 may, for example, present interface 200B as shown in FIG2B.
  • the electronic device 110 can provide a set of candidate media samples 222 in the interface 200B.
  • such candidate media samples can be associated with different preset prompt messages.
  • different candidate media samples can correspond to different media content styles.
  • electronic device 110 when entering interface 200B, can automatically recommend target media sample from a set of candidate media samples 222, and can update input control 220 accordingly based on the prompt information corresponding to the target candidate media sample.
  • the electronic device 110 can display a preset prompt message corresponding to “Media Sample 1” in the input control 220: “I want a song like Music 1, the theme is about XX, and the video length is 3-5 seconds.”
  • the electronic device 110 may determine the target media sample to be recommended from a set of candidate media samples 222 based on information obtained from the interface 200A.
  • electronic device 110 can determine the target media sample to be recommended based on the analysis of at least one piece of material received in interface 200A.
  • material may include, for example, uploaded video material, image material, music material, or text material to be used, etc.
  • the electronic device 110 may also recommend target media samples based on text content that the user has already entered in the input control 212.
  • the electronic device 110 may determine the selected target media sample based on the user's selection of a media sample. In some examples, the electronic device 110 may, for example, not recommend any media samples or recommend other media samples when presenting the interface 200B. Further, the electronic device 110 may, for example, receive the user's selection of "media sample one" and determine "media sample one" as the target media sample.
  • the electronic device 110 may also present an attribute information set corresponding to multiple preset attribute items 224 (e.g., music, style, theme, video duration, etc.) in the interface 200B based on a target media sample (e.g., media sample one).
  • an attribute information set corresponding to multiple preset attribute items 224 e.g., music, style, theme, video duration, etc.
  • a target media sample e.g., media sample one
  • multiple preset attribute items are used to describe different aspects of the video content to be generated. That is, the multiple preset attribute items can correspond to multiple different video attributes of the video content to be generated, such as background music, video style, video theme, video duration, etc.
  • the electronic device 110 presents multiple preset attribute values corresponding to one or more of a plurality of preset attribute items in a target interface. As shown in FIG2B, taking the "music" attribute item as an example, the electronic device 110 can display multiple preset attribute values corresponding to the "music" attribute item in the interface 200B, such as "music one", “music two", etc.
  • the music used in "Media Sample 1” could be “Music 1". Accordingly, the attribute value "Music 1” can be displayed distinctly in the interface 200B. For example, the electronic device 110 can highlight the attribute value "Music 1" by bolding, highlighting, etc., to indicate that the music corresponding to "Media Sample 1" is "Music 1".
  • the electronic device 110 can also provide multiple preset attribute values corresponding to other attribute items, and can distinguish and display one or more attribute values among the multiple preset attribute values that correspond to "media sample one".
  • the electronic device 110 may further receive user editing operations on attribute information corresponding to a target media sample (e.g., media sample one).
  • a target media sample e.g., media sample one
  • the electronic device 110 may receive user adjustment operations on attribute values, thereby completing the editing of the attribute information.
  • the electronic device 110 can, for example, change the attribute information corresponding to the "Music” attribute item from “Music 1" to “Music 2" by receiving the user's selection of the attribute value "Music 2".
  • the electronic device 110 can update the prompt information in the input control 220 based on the modified attribute information.
  • the electronic device 110 can update the prompt information to "I want a song like Music 2, the theme is about XX, and the video length is 3-5 seconds".
  • editing attribute information may also include, for example, deselecting one or more attribute values associated with the target media sample.
  • electronic device 110 can receive the user's selection of the attribute value "Music 1" and deselect "Music 1". Accordingly, electronic device 110 can delete the part corresponding to "Music 1" from the preset prompt information corresponding to "Media Sample 1". For example, electronic device 110 can update the prompt information to "The topic is about XX, and the video duration is 3-5 seconds".
  • editing attribute information may also include selecting one or more attribute values not associated with the target media sample.
  • the target media sample may not include attribute information associated with the "Style” attribute item.
  • the electronic device 110 may, for example, receive the user's selection of "Style One" in the "Style” attribute item and update the prompt information in the input control 220 accordingly.
  • the electronic device 110 may update the prompt information to "I want a song like Music One, with a theme about XX, a video length of 3-5 seconds, and a video style of Style One.”
  • embodiments of the present disclosure can help users edit prompts more efficiently by selecting media samples and modifying attributes, thereby improving the accuracy of the prompts.
  • the set of candidate media samples 222, multiple preset attribute items 224 and/or multiple preset attribute values 226 provided by the electronic device 110 may be determined based on the current creative scenario.
  • the creation scenario can indicate the type or theme of the media content to be created.
  • the media samples, attribute items, and attribute values provided for creating video content may differ from those provided for creating music content.
  • the media samples, attribute items, and attribute values provided for creating different types of video content e.g., educational video content and advertising video content may also differ.
  • the creation scenario can indicate the creation chain for generating media content.
  • different creation chains can correspond to different media samples, attribute items, attribute values, etc.
  • a text-based creation chain can instruct the matching of corresponding materials based on text content to generate video content;
  • a music effect-based creation chain can, for example, instruct the matching of corresponding materials based on music effects (e.g., beats) to generate video content.
  • text-based and music effect-based creation chains can correspond to different media samples, attribute items, attribute values, etc.
  • the electronic device 110 may also provide a preview interface corresponding to the candidate media sample.
  • the electronic device 110 may present a preview interface 200D as shown in FIG2D.
  • the preview interface 200D can display reference media content 230 and reference prompt information 232 corresponding to the selected media sample (e.g., media sample one).
  • reference media content 230 can be, for example, media content generated using reference prompt information 232, to facilitate perception of the generation result corresponding to reference prompt information 232.
  • the electronic device 110 may receive a user's selection of the access point 234, return to the interface 200C, select "Media Sample 1" as the target media sample, and update the input control 220 accordingly based on the reference prompt information 232. For example, the electronic device 110 may display the reference prompt information 232 in the input control 220.
  • the electronic device 110 can receive a swipe operation 236 in the interface 200D and can accordingly display reference media content and reference prompt information of another candidate media sample. In this way, embodiments of the present disclosure can help users more effectively perceive the generation effect of each candidate media sample.
  • the electronic device 110 may return to the target interface 200A shown in FIG2A and may accordingly display the prompt information determined based on FIG2B or FIG2C in the input control 212.
  • the electronic device 110 may also allow users to edit the prompts automatically generated by the electronic device 110 via input control 212 or input control 220.
  • a user can edit the prompt information generated by the electronic device 110 based on the selection of the attribute value "Music Two" in the input control 220.
  • Such editing can include, but is not limited to, adding, modifying, and deleting content.
  • the electronic device 110 can receive the user's additional content "Video style is style two".
  • the electronic device 110 can generate target media content based on prompts in the input control 212 (e.g., descriptive text about the media content to be generated) triggered by the user at the generation entry 218.
  • prompts in the input control 212 e.g., descriptive text about the media content to be generated
  • the electronic device 110 can generate a video with "Music 2" as background music, related to the theme "XX", and with a duration of 3 to 5 seconds.
  • the embodiments of this disclosure allow users to refine the prompts by using preset templates and corresponding attribute adjustments, thereby improving the accuracy of the prompts and the efficiency of editing them, and ultimately improving the quality of the generated media content.
  • FIG 3 shows a flowchart of an example process 300 for generating media content according to some embodiments of the present disclosure.
  • Process 300 can be implemented at electronic device 110.
  • Process 300 will now be described with reference to Figure 1.
  • electronic device 110 provides a set of candidate media samples in the target interface, and the set of candidate media samples corresponds to a set of preset prompt information.
  • electronic device 110 presents a first prompt message corresponding to the target media sample in the target interface.
  • electronic device 110 presents a set of attribute information of the target media sample with respect to multiple preset attribute items in the target interface.
  • the multiple preset attribute items are used to describe different aspects of the media content to be generated.
  • the electronic device 110 modifies the first prompt information to a second prompt information corresponding to the edited second attribute information based on the editing of the first attribute information in the attribute information set.
  • electronic device 110 generates target media content based on the second prompt information.
  • presenting the attribute information set in the target interface includes: presenting multiple preset attribute values corresponding to the target attribute item among multiple preset attribute items in the target interface; and distinguishingly displaying a set of attribute values among the multiple preset attribute values that correspond to the first attribute information.
  • editing the first attribute information includes: deselecting a first attribute value from a set of attribute values; or selecting a second attribute value from a plurality of preset attribute values that is different from a set of attribute values.
  • modifying the first prompt information to a second prompt information corresponding to the edited second attribute information includes: deleting the first part of the first prompt information corresponding to the first attribute value; or adding a second part of the first prompt information corresponding to the second attribute value.
  • generating target media content based on the second prompt information includes: receiving an edit of the second prompt information to determine a third prompt information; and generating target media content based on the third prompt information.
  • process 300 further includes: acquiring at least one media material for generating target media content; and determining a target media sample from a set of candidate media samples based on the at least one media material.
  • determining a target media sample from a set of candidate media samples based on at least one media material includes: in response to receiving a fourth prompt message from a user, determining a target media sample from a set of candidate media samples based on at least one media material and the fourth prompt message.
  • presenting a first prompt message corresponding to a target media sample in a set of candidate media samples on the target interface includes: presenting a first prompt message corresponding to a target media sample on the target interface based on the selection of a target media sample in a set of candidate targets.
  • process 300 further includes: presenting a preview interface based on a first preset operation on the target media sample, the preview interface displaying reference media content and reference prompt information corresponding to the target media sample; and presenting a first prompt information corresponding to the target media sample in the target interface based on the selection of a first entry point in the preview interface, the first prompt information corresponding to the reference prompt information.
  • process 300 further includes: based on a second preset operation received in the preview interface, presenting another reference media content and another reference prompt information in the preview interface that corresponds to another media sample in a set of candidate media samples.
  • providing a set of candidate media samples for controlling the generation of media content in a target interface includes: presenting the target interface based on a request to generate media content, the target interface including an input control for inputting prompt information; providing a second entry point associated with the input control; and providing a set of candidate media samples in the target interface based on the selection of the second entry point.
  • a set of candidate media samples and/or multiple preset attribute items are determined based on the creation scenario corresponding to the target interface.
  • the first prompt information includes first text content corresponding to the first attribute information; or the second prompt information includes second text content corresponding to the second attribute information.
  • the target media content includes video content
  • multiple preset attribute items correspond to multiple different video attributes of the video content.
  • FIG4 shows a schematic structural block diagram of an example media content generation apparatus 400 according to certain embodiments of this disclosure.
  • Apparatus 400 may be implemented as or included in electronic device 110.
  • the various modules/components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
  • the device 400 includes a sample providing module 410, configured to provide a set of candidate media samples in a target interface, the set of candidate media samples corresponding to a set of preset prompt information; a first presentation module 420, configured to present a first prompt information corresponding to a target media sample in the target interface based on the selection of a target media sample from the set of candidate media samples; a second presentation module 430, configured to present an attribute information set of the target media sample about multiple preset attribute items in the target interface, the multiple preset attribute items being used to describe different aspects of the media content to be generated; an information processing module 440, configured to modify the first prompt information to a second prompt information corresponding to the edited second attribute information based on editing the first attribute information in the attribute information set; and a content generation module 450, configured to generate target media content based on the second prompt information.
  • a sample providing module 410 configured to provide a set of candidate media samples in a target interface, the set of candidate media samples corresponding to a set of preset prompt information
  • the second presentation module 430 is specifically configured to present multiple preset attribute values corresponding to a target attribute item among multiple preset attribute items in a target interface; and to distinguishably display a set of attribute values among the multiple preset attribute values that correspond to the first attribute information.
  • editing the first attribute information includes: deselecting a first attribute value from a set of attribute values; or selecting a second attribute value from a plurality of preset attribute values that is different from a set of attribute values.
  • the information processing module 440 is specifically configured to delete the first part of the first prompt information corresponding to the first attribute value; or to add a second part of the first prompt information corresponding to the second attribute value.
  • the content generation module 450 is specifically configured to receive editing of the second prompt information to determine the third prompt information; and to generate target media content based on the third prompt information.
  • the apparatus 400 further includes a material acquisition module configured to acquire at least one media material for generating target media content; and to determine a target media sample from a set of candidate media samples based on the at least one media material.
  • the media acquisition module is specifically configured to, in response to receiving a fourth prompt message input by the user, determine a target media sample from a set of candidate media samples based on at least one media material and the fourth prompt message.
  • the first presentation module 420 is specifically configured to present a first prompt message corresponding to the target media sample in the target interface based on the selection of the target media sample from a set of candidate targets.
  • the device 400 further includes a third presentation module, configured to present a preview interface based on a first preset operation on the target media sample, the preview interface displaying reference media content and reference prompt information corresponding to the target media sample; and to present a first prompt information corresponding to the target media sample in the target interface based on the selection of a first entry point in the preview interface, the first prompt information corresponding to the reference prompt information.
  • a third presentation module configured to present a preview interface based on a first preset operation on the target media sample, the preview interface displaying reference media content and reference prompt information corresponding to the target media sample; and to present a first prompt information corresponding to the target media sample in the target interface based on the selection of a first entry point in the preview interface, the first prompt information corresponding to the reference prompt information.
  • the device 400 further includes a fourth presentation module configured to present, in the preview interface, another reference media content and another reference prompt information corresponding to another media sample in a set of candidate media samples, based on a second preset operation received in the preview interface.
  • a fourth presentation module configured to present, in the preview interface, another reference media content and another reference prompt information corresponding to another media sample in a set of candidate media samples, based on a second preset operation received in the preview interface.
  • the media sample providing module 410 is specifically configured to present a target interface based on a request to generate media content, the target interface including an input control for inputting prompt information; providing a second entry point associated with the input control; and providing a set of candidate media samples in the target interface based on the selection of the second entry point.
  • a set of candidate media samples and/or multiple preset attribute items are determined based on the creation scenario corresponding to the target interface.
  • the first prompt information includes first text content corresponding to the first attribute information; or the second prompt information includes second text content corresponding to the second attribute information.
  • the target media content includes video content
  • multiple preset attribute items correspond to multiple different video attributes of the video content.
  • the modules included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof.
  • one or more units can be implemented using software and/or firmware, such as machine-executable instructions stored on a storage medium.
  • some or all of the modules in device 400 can be implemented at least partially by one or more hardware logic components.
  • exemplary types of hardware logic components include field-programmable gate arrays (FPDAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
  • Figure 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used to implement the electronic device 110 of Figure 1.
  • the electronic device 500 is in the form of a general-purpose electronic device.
  • Components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560.
  • the processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.
  • Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media.
  • Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof.
  • Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and/or data and can be accessed within electronic device 500.
  • Electronic device 500 may further include additional removable/non-removable, volatile/non-volatile storage media.
  • disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided.
  • each drive may be connected to a bus (not shown) via one or more data media interfaces.
  • Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
  • Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
  • PCs network personal computers
  • Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc.
  • Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc.
  • Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input/output (I/O) interface (not shown).
  • I/O input/output
  • a computer-readable storage medium that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above.
  • a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
  • These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and/or other device to operate in a particular manner.
  • the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions/actions specified in one or more boxes of a flowchart and/or block diagram.
  • each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function.
  • the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
  • each block in the block diagrams and/or flowcharts, and combinations of blocks in the block diagrams and/or flowcharts may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Landscapes

  • Engineering & Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Security & Cryptography (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

本公开的实施例涉及生成媒体内容的方法、装置、设备和存储介质。在此提出的方法包括:在目标界面中提供一组候选媒体样本,一组候选媒体样本对应于一组预设提示信息;基于对一组候选媒体样本中的目标媒体样本的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息;在目标界面中呈现目标媒体样本关于多个预设属性项的属性信息集,多个预设属性项用于描述待生成的媒体内容的不同方面;基于对属性信息集中的第一属性信息的编辑,将第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息;以及基于第二提示信息,生成目标媒体内容。

Description

生成媒体内容的方法、装置、设备和存储介质
本申请要求2024年4月29日递交的、标题为“生成媒体内容的方法、装置、设备和存储介质”、申请号为202410534129.4的中国发明专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
技术领域
本公开的示例实施例总体涉及计算机领域,特别地涉及生成媒体内容的方法、装置、设备和计算机可读存储介质。
背景技术
随着计算机技术的不断发展,生成式人工智能技术逐渐被应用各个领域。人们例如可以利用生成式人工智能技术来创作各种类型的媒体内容,例如,图片、视频等。在这样的创作过程中,如何更好地引用媒体内容的生成过程以提高媒体内容的生成质量是人们关注的焦点问题。
发明内容
在本公开的第一方面,提供了一种生成媒体内容的方法。该方法包括:在目标界面中提供一组候选媒体样本,一组候选媒体样本对应于一组预设提示信息;基于对一组候选媒体样本中的目标媒体样本的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息;在目标界面中呈现目标媒体样本关于多个预设属性项的属性信息集,多个预设属性项用于描述待生成的媒体内容的不同方面;基于对属性信息集中的第一属性信息的编辑,将第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息;以及基于第二提示信息,生成目标媒体内容。
在本公开的第二方面,提供了一种用于生成媒体内容的装置。该装置包括:样本提供模块,被配置为在目标界面中提供一组候选媒体样本,一组候选媒体样本对应于一组预设提示信息;第一呈现模块,被配置为基于对一组候选媒体样本中的目标媒体样本的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息;第二呈现模块,被配置为在目标界面中呈现目标媒体样本关于多个预设属性项的属性信息集,多个预设属性项用于描述待生成的媒体内容的不同方面;信息处理模块,被配置为基于对属性信息集中的第一属性信息的编辑,将第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息;以及内容生成模块,被配置为基于第二提示信息,生成目标媒体内容。
在本公开的第三方面,提供了一种电子设备。该设备包括至少一个处理单元;以及至少一个存储器,至少一个存储器被耦合到至少一个处理单元并且存储用于由至少一个处理单元执行的指令。指令在由至少一个处理单元执行时使设备执行第一方面的方法。
在本公开的第四方面,提供了一种计算机可读存储介质。该计算机可读存储介质上存储有计算机程序,计算机程序可由处理器执行以实现第一方面的方法。
应当理解,本内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其他特征将通过以下的描述而变得容易理解。
附图说明
结合附图并参考以下详细说明,本公开各实施例的上述和其他特征、优点及方面将变得更加明显。在附图中,相同或相似的附图标记表示相同或相似的元素,其中:
图1示出了其中可以实施根据本公开的实施例的示例环境的示意图;
图2A至图2D示出了根据本公开的一些实施例的示例界面;
图3示出了根据本公开的一些实施例的示例生成媒体内容的过程的流程图;
图4示出了根据本公开的一些实施例的示例生成媒体内容装置的示意性结构框图;以及
图5示出了能够实施本公开的多个实施例的电子设备的框图。
具体实施方式
下面将参照附图更详细地描述本公开的实施例。虽然附图中示出了本公开的某些实施例,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例,相反,提供这些实施例是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实施例仅用于示例性作用,并非用于限制本公开的保护范围。
需要注意的是,本文中所提供的任何节/子节的标题并不是限制性的。本文通篇描述了各种实施例,并且任何类型的实施例都可以包括在任何节/子节下。此外,在任一节/子节中描述的实施例可以以任何方式与同一节/子节和/或不同节/子节中描述的任何其他实施例相结合。
在本公开的实施例的描述中,术语“包括”及其类似用语应当理解为开放性包含,即“包括但不限于”。术语“基于”应当理解为“至少部分地基于”。术语“一个实施例”或“该实施例”应当理解为“至少一个实施例”。术语“一些实施例”应当理解为“至少一些实施例”。下文还可能包括其他明确的和隐含的定义。术语“第一”、“第二”等可以指待不同的或相同的对象。下文还可能包括其他明确的和隐含的定义。
本公开的实施例中可能涉及用户的数据、数据的获取和/或使用等。这些方面均遵循相应的法律法规及相关规定。在本公开的实施例中,所有数据的采集、获取、处理、加工、转发、使用等,都是在用户知晓并且确认的前提下进行的。相应地,在实现本公开的各实施例时,均应根据相关法律法规通过适当的方式,将可能所涉及的数据或信息的类型、使用范围、使用场景等告知用户并获得用户的授权。具体的告知和/或授权方式可以根据实际情况和应用场景而变化,本公开的范围在此方面不受限制。
本说明书及实施例中方案,如涉及个人信息处理,则均会在具备合法性基础(例如征得个人信息主体同意,或者为履行合同所必需等)的前提下进行处理,且仅会在规定或者约定的范围内进行处理。用户拒绝处理基本功能所需必要信息以外的个人信息,不会影响用户使用基本功能。
如上文所介绍的,在人们使用生成式人工智能技术来进行媒体创作时,准确的提示项是影响媒体内容的生成质量的关键因素。传统的方案通常依赖于用户输入文本内容作为提示项,或者只是简单地指定预设的模板。这都限制了提示项的表达准确性,并影响了媒体内容的生成质量。
为此,本公开的实施例提出了一种用于生成媒体内容的方案。根据该方案,可以在目标界面中提供一组候选媒体样本,一组候选媒体样本对应于一组预设提示信息;基于对一组候选媒体样本中的目标媒体样本的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息;在目标界面中呈现目标媒体样本关于多个预设属性项的属性信息集,多个预设属性项用于描述待生成的媒体内容的不同方面;基于对属性信息集中的第一属性信息的编辑,将第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息;以及基于第二提示信息,生成目标媒体内容。
以此方式,本公开的实施例能够支持用户通过预设媒体样本和对应的属性调整来实现对提示信息的细化编辑,从而能够提高提示信息的准确性和编辑提示信息的效率,进而提高所生成的媒体内容的质量。
以下进一步结合附图来详细描述该方案的各种示例实现。
示例环境
图1示出了本公开的实施例能够在其中实现的示例环境100的示意图。如图1所示,示例环境100可以包括电子设备110。
在该示例环境100中,电子设备110可以运行有支持界面交互的应用120。应用120可以是用于界面交互的任何适当类型应用,其示例可以包括但不限于:视频应用、剪辑应用或其他适当的应用。用户140可以经由电子设备110和/或其附接设备来与应用120进行交互。
在图1的环境100中,如果应用120处于活动状态,电子设备110可以通过应用120呈现用于支持界面交互的界面150。
在一些实施例中,电子设备110与服务器130通信,以实现对应用120的服务的供应。电子设备110可以是任意类型的移动终端、固定终端或便携式终端,包括移动手机、台式计算机、膝上型计算机、笔记本计算机、上网本计算机、平板计算机、媒体计算机、多媒体平板、掌上电脑、便携式游戏终端、VR/AR设备、个人通信系统(Personal Communication System,PCS)设备、个人导航设备、个人数字助理(Personal DiDital Assistant,PDA)、音频/视频播放器、数码相机/摄像机、定位设备、电视接收器、无线电广播接收器、电子书设备、游戏设备或者前述各项的任意组合,包括这些设备的配件和外设或者其任意组合。在一些实施例中,电子设备110也能够支持任意类型的针对用户的接口(诸如“可佩戴”电路等)。
服务器130可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络、以及大数据和人工智能平台等基础云计算服务的云服务器。服务器130例如可以包括计算系统/服务器,诸如大型机、边缘计算节点、云环境中的计算设备,等等。服务器130可以为电子设备110中支持虚拟场景的应用120提供后台服务。
服务器130与电子设备110之间可以建立有通信连接。通信连接可以通过有线方式或无线方式建立。通信连接可以包括但不限于蓝牙连接、移动网络连接、通用串行总线(Universal Serial Bus,USB)连接、无线保真(Wireless Fidelity,WiFi)连接等,本公开的实施例在此方面不受限制。在本公开的实施例中,服务器130与电子设备110可以通过二者之间的通信连接实现信令交互。
应当理解,仅出于示例性的目的描述环境100中各个元素的结构和功能,而不暗示对于本公开的范围的任何限制。
以下将继续参考附图描述本公开的一些示例实施例。
示例交互
以下将参考附图来描述根据本公开的实施例的示例生成媒体内容的过程。下文将以视频内容作为媒体内容的示例来描述媒体内容的示例生成过程。应当理解的是,本公开的生成过程也可以适用于其他类型的媒体内容,例如,图像、音频等。
图2A至图2D示出了根据本公开的一些实施例的示例界面200A至200D。界面200A至界面200D可以由图1所示的电子设备110所提供。
在一些实施例中,电子设备110基于生成媒体内容的请求,呈现目标界面。图2A示出了根据本公开的一些实施例的示例目标界面。如图2A所示,目标界面可以包括用于输入提示信息的输入控件212。
如图2A所示,用户可以在输入控件中输入关于待生成视频内容的提示信息(例如,一段描述文本),即用户关于待生成视频内容的描述。电子设备110可以基于所接收到的提示信息来确定待生成的视频内容。
在一些实施例中,电子设备110还可以经由界面200A来获取用于生成视频内容的至少一项素材。例如,用户可以基于控件210上传视频素材、图像素材、音频素材等。备选地,用户也可以通过控件212来输入文案素材,以例如作为待生成的文案内容的字幕等。
在一些实施例中,电子设备110可以接受用户对生成入口216的选择,并基于界面200A中接收到的提示信息和素材来生成媒体内容。作为示例,电子设备110例如可以利用生成式人工智能技术来生成对应的媒体内容,本公开不旨在对此媒体内容的具体生成过程进行限定。
在一些实施例中,电子设备110还可以支持用户更为高效地编辑用于生成媒体内容的提示信息。如图2A所示,电子设备110关联于输入控件212来提供入口214。在接收到针对入口214的触发后,电子设备110例如可以呈现如图2B所示的界面200B。
如图2B所示,电子设备110可以在界面200B中提供一组候选媒体样本222。在一些实施例中,这样的候选媒体样本可以关联于不同的预设提示信息。作为示例,不同的候选媒体样本可以对应于不同的媒体内容样式。
在一些实施例中,在进入到界面200B时,电子设备110可以自动推荐一组候选媒体样本222中的目标媒体样本,并可以相应地基于该目标候选媒体样本所对应的提示信息来更新输入控件220。
例如,在“媒体样本一”被确定为目标媒体样本的情况下,电子设备110可以在输入控件220中显示与“媒体样本一”对应的预设提示信息“我想要像音乐一的歌曲,主题是关于XX,视频时长3-5秒”。
在一些实施例中,电子设备110可以基于界面200A中获取到的信息来从一组候选媒体样本222中确定待推荐的目标媒体样本。
作为示例,电子设备110可以基于对界面200A中所接收到的至少一项素材的分析来确定待推荐的目标媒体样本。这样的素材例如可以包括所上传的视频素材、图片素材、音乐素材,也可以包括待使用的文案素材等等。
在又一些实施例中,电子设备110还可以基于输入控件212中用户已经输入的文本内容来推荐目标媒体样本。
在一些实施例中,电子设备110例如也可以基于用户对媒体样本的选择来确定所选择的目标媒体样本。在一些示例中,电子设备110在呈现界面200B时例如可以不推荐任何媒体样本,或者推荐其他的媒体样本。进一步地,电子设备110例如可以接收用户对“媒体样本一”的选择,并将“媒体样本一”确定为目标媒体样本。
在一些实施例中,如图2B所示,电子设备110还可以基于目标媒体样本(例如,媒体样本一),来在界面200B中呈现与多个预设属性项224(例如,音乐、风格、主题、视频时长等)对应的属性信息集。
在一些实施例中,以媒体内容为视频内容为例,多个预设属性项用于描述待生成视频内容的不同方面。也即,多个预设属性项可以对应于待生成的视频内容的多个不同视频属性,例如,背景音乐、视频风格、视频主题、视频时长等。
在一些实施例中,电子设备110在目标界面中呈现与多个预设属性项中的一个或多个属性项对应的多个预设属性值。如图2B所示,以“音乐”属性项作为示例,电子设备110可以在界面200B中显示与“音乐”属性项对应的多个预设属性值,例如,“音乐一”、“音乐二”等等。
作为示例,“媒体样本一”所使用的音乐例如可以为“音乐一”。相应地,属性值“音乐一”可以在界面200B中被区别地显示。例如,电子设备110可以通过加粗、高亮等方式来突出显示属性值“音乐一”,以表征“媒体样本一”所对应的音乐为“音乐一”。
类似地,电子设备110还可以提供与其他属性项所对应的多个预设属性值,并可以区别显示多个预设属性值中与“媒体样本一”所对应的一个或多个属性值。
在一些实施例中,电子设备110可以进一步接收用户对与目标媒体样本(例如,媒体样本一)对应的属性信息的编辑操作。作为示例,电子设备110可以接收用户对于属性值的调整操作,从而完成对属性信息的编辑。
如图2C所示,电子设备110例如可以通过接收用户对于属性值“音乐二”的选择,从而将“音乐”属性项所对应的属性信息从“音乐一”修改为“音乐二”。
相应地,电子设备110可以基于修改后的属性信息来更新输入控件220中的提示信息。作为示例,电子设备110可以将提示信息更新为“我想要像音乐二的歌曲,主题是关于XX,视频时长3-5秒”。
在一些实施例中,对属性信息的编辑例如还可以包括取消对与目标媒体样本相关联的一个或多个属性值的选择。
例如,电子设备110可以接收用户对属性值“音乐一”的选择,并取消“音乐一”的选中状态。相应地,电子设备110可以从与“媒体样本一”对应的预设提示信息中删除与“音乐一”对应的部分。例如,电子设备110可以将提示信息更新为“主题是关于XX,视频时长3-5秒”。
在一些实施例中,对属性信息的编辑例如还可以包括选择未关联于目标媒体样本的一个或多个属性值。例如,目标媒体样本可能不包括与“风格”属性项相关联的属性信息。进一步地,电子设备110例如可以接收用户对于“风格”属性项中的“风格一”的选择,并可以相应地更新输入控件220中的提示信息。例如,电子设备110可以将提示信息更新为“我想要像音乐一的歌曲,主题是关于XX,视频时长3-5秒,视频风格为风格一”。
以此方式,本公开的实施例可以帮助用户通过选择媒体样本和修改属性的方式来更为高效地编辑提示信息,从而提高提示信息的准确性。
在一些实施例中,电子设备110所提供的一组候选媒体样本222、多个预设属性项224和/或多个预设属性值226可以是基于当前的创作场景所确定的。
在一些实施例中,创作场景可以指示待创作的媒体内容的类型或主题。例如,创作视频内容所提供的媒体样本、属性项、属性值可以不同于创作音乐内容所提供的媒体样本、属性项、属性值等。又例如,创作不同类型的视频内容(例如,科普视频内容和广告视频内容)所提供的媒体样本、属性项、属性值也可以不同。
在一些实施例中,创作场景可以指示生成媒体内容的创作链路。例如,不同的创作链路可以对应于不同的媒体样本、属性项、属性值等。作为示例,基于文案的创作链路可以指示基于文案内容来匹配对应的素材来生成视频内容;基于音乐特效的创作链路例如可以指示基于音乐特效(例如,节拍)来匹配对应的素材来生成视频内容。在一些示例中,基于文案的创作链路和基于音乐特效的创作链路可以对应于不同的媒体样本、属性项、属性值等。
在一些实施例中,为了方便用户更好地感知各媒体样本的效果,电子设备110还可以提供与候选媒体样本对应的预览界面。在一些实施例中,电子设备110在接收到对候选媒体样本222的预设操作(例如,点击、双击、长按等)后,可以呈现如图2D所示的预览界面200D。
如图2D所示,预览界面200D可以显示与所选择的媒体样本(例如,媒体样本一)对应的参考媒体内容230和参考提示信息232。这样的参考媒体内容230例如可以是利用参考提示信息232所生成的媒体内容,以方便用于感知参考提示信息232所对应的生成结果。
在一些实施例中,电子设备110可以接收用户对使用入口234的选择,并可以返回至界面200C,并将“媒体样本一”选择为目标媒体样本,并相应地根据参考提示信息232来更新输入控件220。例如,电子设备110可以将参考提示信息232显示在输入控件220中。
在一些实施例中,如图2D所示,电子设备110可以接收界面200D中的滑动操作236,并可以相应地显示另一候选媒体样本的参考媒体内容和参考提示信息。以此方式,本公开的实施例可以帮助用户更为有效地感知各候选媒体样本的生成效果。
在一些实施例中,在接收到对于图2B或图2C所示的“确认”按钮的选择后,电子设备110可以返回至如图2A所示的目标界面200A,并可以相应地在输入控件212中显示基于图2B或图2C所确定的提示信息。
在一些实施例中,电子设备110还可以支持用户通过输入控件212或输入控件220来对电子设备110所自动生成的提示信息进行编辑。
以图2C作为示例,用户例如可以在输入控件220对电子设备110基于对属性值“音乐二”的选择所生成的提示信息进行编辑。这样的编辑可以包括但不限于:增加内容、修改内容、删除内容等。例如,电子设备110可以接收用户所补充的内容“视频风格为风格二”。
作为又一示例,用户例如通过点击图2C所示的确认按钮返回至界面200A时,并可以通过输入控件212来类似地对提示信息进行相应地编辑。
进一步地,电子设备110可以基于用户对生成入口218的触发,以基于输入控件212中的提示信息(例如,关于待生成媒体内容的描述文本)来生成目标媒体内容。
例如,以图2C的输入控件220所示的提示信息作为示例,电子设备110例如可以相应地生成以“音乐二”作为背景音乐、与主题“XX””相关,并且时长在3秒到5秒的一段视频。
以此方式,本公开的实施例支持用户通过预设模板和对应的属性调整来实现对提示项的细化编辑,从而能够提高提示项的准确性和编辑提示项的效率,进而提高所生成的媒体内容的质量。
示例过程
图3示出了根据本公开的一些实施例的生成媒体内容示例过程300的流程图。过程300可以被实现在电子设备110处。下面参考图1来描述过程300。
如图所示,在框310,电子设备110在目标界面中提供一组候选媒体样本,一组候选媒体样本对应于一组预设提示信息。
在框320,电子设备110基于对一组候选媒体样本中的目标媒体样本的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息。
在框330,电子设备110在所述目标界面中呈现所述目标媒体样本关于多个预设属性项的属性信息集,所述多个预设属性项用于描述待生成的媒体内容的不同方面。
在框340,电子设备110基于对所述属性信息集中的第一属性信息的编辑,将所述第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息。
在框350,电子设备110基于第二提示信息,生成目标媒体内容。
在一些实施例中,在目标界面中呈现属性信息集包括:在目标界面中呈现与多个预设属性项中的目标属性项对应的多个预设属性值;以及区别地显示多个预设属性值中与第一属性信息对应的一组属性值。
在一些实施例中,对第一属性信息的编辑包括:取消对一组属性值中的第一属性值的选择;或选择多个预设属性值中不同于一组属性值的第二属性值。
在一些实施例中,将第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息包括:删除第一提示信息中与第一属性值对应的第一部分;或向第一提示信息增加与第二属性值对应的第二部分。
在一些实施例中,基于第二提示信息生成目标媒体内容包括:接收对第二提示信息的编辑,以确定第三提示信息;以及基于第三提示信息,生成目标媒体内容。
在一些实施例中,过程300还包括:获取用于生成目标媒体内容的至少一项媒体素材;以及基于至少一项媒体素材,从一组候选媒体样本中确定目标媒体样本。
在一些实施例中,基于至少一项媒体素材从一组候选媒体样本中确定目标媒体样本包括:响应于获取到用户输入的第四提示信息,基于至少一项媒体素材和第四提示信息,从一组候选媒体样本中确定目标媒体样本。
在一些实施例中,在目标界面中呈现与一组候选媒体样本中的目标媒体样本对应的第一提示信息包括:基于对一组候选目标中的目标媒体样本的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息。
在一些实施例中,过程300还包括:基于对目标媒体样本的第一预设操作,呈现预览界面,预览界面显示与目标媒体样本对应的参考媒体内容和参考提示信息;以及基于对预览界面中的第一入口的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息,第一提示信息对应于参考提示信息。
在一些实施例中,过程300还包括:基于预览界面中接收到的第二预设操作,在预览界面中呈现与一组候选媒体样本中的另一媒体样本对应的另一参考媒体内容和另一参考提示信息。
在一些实施例中,在目标界面中提供用于控制媒体内容的生成的一组候选媒体样本包括:基于生成媒体内容的请求,呈现目标界面,目标界面包括用于输入提示信息的输入控件;关联于输入控件提供第二入口;以及基于对第二入口的选择,在目标界面中提供一组候选媒体样本。
在一些实施例中,一组候选媒体样本和/或多个预设属性项是基于与目标界面对应的创作场景所确定。
在一些实施例中,第一提示信息包括与第一属性信息对应的第一文本内容;或第二提示信息包括与第二属性信息对应的第二文本内容。
在一些实施例中,目标媒体内容包括视频内容,并且多个预设属性项对应于视频内容的多个不同视频属性。
示例装置和设备
本公开的实施例还提供了用于实现上述方法或过程的相应装置。图4示出了根据本公开的某些实施例的用于示例生成媒体内容装置400的示意性结构框图。装置400可以被实现为或者被包括在电子设备110中。装置400中的各个模块/组件可以由硬件、软件、固件或者它们的任意组合来实现。
如图4所示,装置400包括样本提供模块410,被配置为在目标界面中提供一组候选媒体样本,一组候选媒体样本对应于一组预设提示信息;第一呈现模块420,被配置为基于对一组候选媒体样本中的目标媒体样本的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息;第二呈现模块430,被配置为在目标界面中呈现目标媒体样本关于多个预设属性项的属性信息集,多个预设属性项用于描述待生成的媒体内容的不同方面;信息处理模块440,被配置为基于对属性信息集中的第一属性信息的编辑,将第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息;以及内容生成模块450,被配置为基于第二提示信息,生成目标媒体内容。
在一些实施例中,第二呈现模块430,被具体配置为在目标界面中呈现与多个预设属性项中的目标属性项对应的多个预设属性值;以及区别地显示多个预设属性值中与第一属性信息对应的一组属性值。
在一些实施例中,对第一属性信息的编辑包括:取消对一组属性值中的第一属性值的选择;或选择多个预设属性值中不同于一组属性值的第二属性值。
在一些实施例中,信息处理模块440,被具体配置为删除第一提示信息中与第一属性值对应的第一部分;或向第一提示信息增加与第二属性值对应的第二部分。
在一些实施例中,内容生成模块450,被具体配置为接收对第二提示信息的编辑,以确定第三提示信息;以及基于第三提示信息,生成目标媒体内容。
在一些实施例中,装置400还包括素材获取模块,被配置为获取用于生成目标媒体内容的至少一项媒体素材;以及基于至少一项媒体素材,从一组候选媒体样本中确定目标媒体样本。
在一些实施例中,素材获取模块,被具体配置为响应于获取到用户输入的第四提示信息,基于至少一项媒体素材和第四提示信息,从一组候选媒体样本中确定目标媒体样本。
在一些实施例中,第一呈现模块420,被具体配置为基于对一组候选目标中的目标媒体样本的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息。
在一些实施例中,装置400还包括第三呈现模块,被配置为基于对目标媒体样本的第一预设操作,呈现预览界面,预览界面显示与目标媒体样本对应的参考媒体内容和参考提示信息;以及基于对预览界面中的第一入口的选择,在目标界面中呈现与目标媒体样本对应的第一提示信息,第一提示信息对应于参考提示信息。
在一些实施例中,装置400还包括第四呈现模块,被配置为基于预览界面中接收到的第二预设操作,在预览界面中呈现与一组候选媒体样本中的另一媒体样本对应的另一参考媒体内容和另一参考提示信息。
在一些实施例中,媒体样本提供模块410,被具体配置为基于生成媒体内容的请求,呈现目标界面,目标界面包括用于输入提示信息的输入控件;关联于输入控件提供第二入口;以及基于对第二入口的选择,在目标界面中提供一组候选媒体样本。
在一些实施例中,一组候选媒体样本和/或多个预设属性项是基于与目标界面对应的创作场景所确定。
在一些实施例中,第一提示信息包括与第一属性信息对应的第一文本内容;或第二提示信息包括与第二属性信息对应的第二文本内容。
在一些实施例中,目标媒体内容包括视频内容,并且多个预设属性项对应于视频内容的多个不同视频属性。
装置400中所包括的模块可以利用各种方式来实现,包括软件、硬件、固件或其任意组合。在一些实施例中,一个或多个单元可以使用软件和/或固件来实现,例如存储在存储介质上的机器可执行指令。除了机器可执行指令之外或者作为替待,装置400中的部分或者全部模块可以至少部分地由一个或多个硬件逻辑组件来实现。作为示例而非限制,可以使用的示范类型的硬件逻辑组件包括现场可编程门阵列(FPDA)、专用集成电路(ASIC)、专用标准品(ASSP)、片上系统(SOC)、复杂可编程逻辑器件(CPLD),等等。
图5示出了其中可以实施本公开的一个或多个实施例的电子设备500的框图。应当理解,图5所示出的电子设备500仅仅是示例性地,而不应当构成对本文所描述的实施例的功能和范围的任何限制。图5所示出的电子设备500可以用于实现图1的电子设备110。
如图5所示,电子设备500是通用电子设备的形式。电子设备500的组件可以包括但不限于一个或多个处理器或处理单元510、存储器520、存储设备530、一个或多个通信单元540、一个或多个输入设备550以及一个或多个输出设备560。处理单元510可以是实际或虚拟处理器并且能够根据存储器520中存储的程序来执行各种处理。在多处理器系统中,多个处理单元并行执行计算机可执行指令,以提高电子设备500的并行处理能力。
电子设备500通常包括多个计算机存储介质。这样的介质可以是电子设备500可访问的任何可以获取的介质,包括但不限于易失性和非易失性介质、可拆卸和不可拆卸介质。存储器520可以是易失性存储器(例如寄存器、高速缓存、随机访问存储器(RAM))、非易失性存储器(例如,只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、闪存)或它们的某种组合。存储设备530可以是可拆卸或不可拆卸的介质,并且可以包括机器可读介质,诸如闪存驱动、磁盘或者任何其他介质,其可以能够用于存储信息和/或数据并且可以在电子设备500内被访问。
电子设备500可以进一步包括另外的可拆卸/不可拆卸、易失性/非易失性存储介质。尽管未在图5中示出,可以提供用于从可拆卸、非易失性磁盘(例如“软盘”)进行读取或写入的磁盘驱动和用于从可拆卸、非易失性光盘进行读取或写入的光盘驱动。在这些情况中,每个驱动可以由一个或多个数据介质接口被连接至总线(未示出)。存储器520可以包括计算机程序产品525,其具有一个或多个程序模块,这些程序模块被配置为执行本公开的各种实施例的各种方法或动作。
通信单元540实现通过通信介质与其他电子设备进行通信。附加地,电子设备500的组件的功能可以以单个计算集群或多个计算机器来实现,这些计算机器能够通过通信连接进行通信。因此,电子设备500可以使用与一个或多个其他服务器、网络个人计算机(PC)或者另一个网络节点的逻辑连接来在联网环境中进行操作。
输入设备550可以是一个或多个输入设备,例如鼠标、键盘、追踪球等。输出设备560可以是一个或多个输出设备,例如显示器、扬声器、打印机等。电子设备500还可以根据需要通过通信单元540与一个或多个外部设备(未示出)进行通信,外部设备诸如存储设备、显示设备等,与一个或多个使得用户与电子设备500交互的设备进行通信,或者与使得电子设备500与一个或多个其他电子设备通信的任何设备(例如,网卡、调制解调器等)进行通信。这样的通信可以经由输入/输出(I/O)接口(未示出)来执行。
根据本公开的示例性实现方式,提供了一种计算机可读存储介质,其上存储有计算机可执行指令,其中计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,还提供了一种计算机程序产品,计算机程序产品被有形地存储在非瞬态计算机可读介质上并且包括计算机可执行指令,而计算机可执行指令被处理器执行以实现上文描述的方法。
这里参照根据本公开实现的方法、装置、设备和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理单元,从而生产出一种机器,使得这些指令在通过计算机或其他可编程数据处理装置的处理单元执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
可以把计算机可读程序指令加载到计算机、其他可编程数据处理装置、或其他设备上,使得在计算机、其他可编程数据处理装置或其他设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其他可编程数据处理装置、或其他设备上执行的指令实现流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实现的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以待表一个模块、程序段或指令的一部分,模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实现,上述说明是示例性地,并非穷尽性的,并且也不限于所公开的各实现。在不偏离所说明的各实现的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实现的原理、实际应用或对市场中的技术的改进,或者使本技术领域的其他普通技术人员能理解本文公开的各个实现方式。

Claims (17)

  1. 一种生成媒体内容的方法,包括:
    在目标界面中提供一组候选媒体样本,所述一组候选媒体样本对应于一组预设提示信息;
    基于对所述一组候选媒体样本中的目标媒体样本的选择,在所述目标界面中呈现与所述目标媒体样本对应的第一提示信息;
    在所述目标界面中呈现所述目标媒体样本关于多个预设属性项的属性信息集,所述多个预设属性项用于描述待生成的媒体内容的不同方面;
    基于对所述属性信息集中的第一属性信息的编辑,将所述第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息;以及
    基于所述第二提示信息,生成目标媒体内容。
  2. 根据权利要求1所述的方法,其中在所述目标界面中呈现所述属性信息集包括:
    在所述目标界面中呈现与所述多个预设属性项中的目标属性项对应的多个预设属性值;以及
    区别地显示所述多个预设属性值中与所述第一属性信息对应的一组属性值。
  3. 根据权利要求2所述的方法,其中对所述第一属性信息的编辑包括:
    取消对所述一组属性值中的第一属性值的选择;或
    选择所述多个预设属性值中不同于所述一组属性值的第二属性值。
  4. 根据权利要求3所述的方法,其中将所述第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息包括:
    删除所述第一提示信息中与所述第一属性值对应的第一部分;或
    向所述第一提示信息增加与所述第二属性值对应的第二部分。
  5. 根据权利要求1所述的方法,其中基于所述第二提示信息生成目标媒体内容包括:
    接收对所述第二提示信息的编辑,以确定第三提示信息;以及
    基于所述第三提示信息,生成所述目标媒体内容。
  6. 根据权利要求1所述的方法,还包括:
    获取用于生成所述目标媒体内容的至少一项媒体素材;以及
    基于所述至少一项媒体素材,从所述一组候选媒体样本中确定所述目标媒体样本。
  7. 根据权利要求6所述的方法,其中基于所述至少一项媒体素材从所述一组候选媒体样本中确定所述目标媒体样本包括:
    响应于获取到用户输入的第四提示信息,基于所述至少一项媒体素材和所述第四提示信息,从所述一组候选媒体样本中确定所述目标媒体样本。
  8. 根据权利要求1所述的方法,其中在所述目标界面中呈现与所述一组候选媒体样本中的目标媒体样本对应的第一提示信息包括:
    基于对所述一组候选媒体样本中的所述目标媒体样本的选择,在所述目标界面中呈现与所述目标媒体样本对应的第一提示信息。
  9. 根据权利要求8所述的方法,还包括:
    基于对所述目标媒体样本的第一预设操作,呈现预览界面,所述预览界面显示与所述目标媒体样本对应的参考媒体内容和参考提示信息;以及
    基于对所述预览界面中的第一入口的选择,在所述目标界面中呈现与所述目标媒体样本对应的所述第一提示信息,所述第一提示信息对应于所述参考提示信息。
  10. 根据权利要求9所述的方法,还包括:
    基于所述预览界面中接收到的第二预设操作,在所述预览界面中呈现与所述一组候选媒体样本中的另一媒体样本对应的另一参考媒体内容和另一参考提示信息。
  11. 根据权利要求1所述的方法,其中在目标界面中提供用于控制媒体内容的生成的一组候选媒体样本包括:
    基于生成媒体内容的请求,呈现所述目标界面,所述目标界面包括用于输入提示信息的输入控件;
    关联于所述输入控件提供第二入口;以及
    基于对所述第二入口的选择,在所述目标界面中提供所述一组候选媒体样本。
  12. 根据权利要求1所述的方法,其中所述一组候选媒体样本和/或所述多个预设属性项是基于与所述目标界面对应的创作场景所确定。
  13. 根据权利要求1所述的方法,其中:
    所述第一提示信息包括与所述第一属性信息对应的第一文本内容;或
    所述第二提示信息包括与所述第二属性信息对应的第二文本内容。
  14. 根据权利要求1所述的方法,其中所述目标媒体内容包括视频内容,并且所述多个预设属性项对应于视频内容的多个不同视频属性。
  15. 一种用于生成媒体内容的装置,包括:
    样本提供模块,被配置为在目标界面中提供一组候选媒体样本,所述一组候选媒体样本对应于一组预设提示信息;
    第一呈现模块,被配置为基于对所述一组候选媒体样本中的目标媒体样本的选择,在所述目标界面中呈现与所述目标媒体样本对应的第一提示信息;
    第二呈现模块,被配置为在所述目标界面中呈现所述目标媒体样本关于多个预设属性项的属性信息集,所述多个预设属性项用于描述待生成的媒体内容的不同方面;
    信息处理模块,被配置为基于对所述属性信息集中的第一属性信息的编辑,将所述第一提示信息修改为与编辑后的第二属性信息对应的第二提示信息;以及
    内容生成模块,被配置为基于所述第二提示信息,生成目标媒体内容。
  16. 一种电子设备,包括:
    至少一个处理单元;以及
    至少一个存储器,所述至少一个存储器被耦合到所述至少一个处理单元并且存储用于由所述至少一个处理单元执行的指令,所述指令在由所述至少一个处理单元执行时使所述电子设备执行根据权利要求1至14中任一项所述的方法。
  17. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序可由处理器执行以实现根据权利要求1至14中任一项所述的方法。
PCT/CN2025/091825 2024-04-29 2025-04-28 生成媒体内容的方法、装置、设备和存储介质 Pending WO2025228337A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410534129.4A CN120881328A (zh) 2024-04-29 2024-04-29 生成媒体内容的方法、装置、设备和存储介质
CN202410534129.4 2024-04-29

Publications (1)

Publication Number Publication Date
WO2025228337A1 true WO2025228337A1 (zh) 2025-11-06

Family

ID=97459641

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/091825 Pending WO2025228337A1 (zh) 2024-04-29 2025-04-28 生成媒体内容的方法、装置、设备和存储介质

Country Status (2)

Country Link
CN (1) CN120881328A (zh)
WO (1) WO2025228337A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200335132A1 (en) * 2019-04-18 2020-10-22 Kristin Fahy Systems and Methods for Automated Generation of Video
CN116433825A (zh) * 2023-05-24 2023-07-14 北京百度网讯科技有限公司 图像生成方法、装置、计算机设备及存储介质
CN117332118A (zh) * 2023-10-24 2024-01-02 科大讯飞股份有限公司 一种故事视频的生成方法、装置、存储介质及设备
CN117493013A (zh) * 2023-11-16 2024-02-02 广州商研网络科技有限公司 提示文本生成方法及其装置、设备、介质
CN117726716A (zh) * 2023-11-29 2024-03-19 抖音视界有限公司 一种多媒体数据处理方法、装置、电子设备及存储介质

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200335132A1 (en) * 2019-04-18 2020-10-22 Kristin Fahy Systems and Methods for Automated Generation of Video
CN116433825A (zh) * 2023-05-24 2023-07-14 北京百度网讯科技有限公司 图像生成方法、装置、计算机设备及存储介质
CN117332118A (zh) * 2023-10-24 2024-01-02 科大讯飞股份有限公司 一种故事视频的生成方法、装置、存储介质及设备
CN117493013A (zh) * 2023-11-16 2024-02-02 广州商研网络科技有限公司 提示文本生成方法及其装置、设备、介质
CN117726716A (zh) * 2023-11-29 2024-03-19 抖音视界有限公司 一种多媒体数据处理方法、装置、电子设备及存储介质

Also Published As

Publication number Publication date
CN120881328A (zh) 2025-10-31

Similar Documents

Publication Publication Date Title
WO2025168004A1 (zh) 创建虚拟对象的方法、装置、设备和存储介质
WO2025251807A1 (zh) 生成音乐的方法、装置、设备和存储介质
WO2025092766A1 (zh) 用于显示作品的方法、装置、设备和存储介质
WO2026040862A1 (zh) 创建模板的方法、装置、设备和存储介质
WO2026041028A1 (zh) 界面交互的方法、装置、设备和存储介质
WO2026067727A1 (zh) 处理媒体内容的方法、装置、设备、存储介质和程序产品
WO2025256532A1 (zh) 发布内容的方法、装置、设备和存储介质
US20250271981A1 (en) Method, apparatus, device, and storage medium for media item input
WO2026046251A1 (zh) 用于信息搜索方法、装置、设备和存储介质
WO2026051838A1 (zh) 用于媒体编辑方法、装置、设备和存储介质
WO2026007522A1 (zh) 生成媒体内容的方法、装置、设备和存储介质
US20250272335A1 (en) Method, appartus, device and storage medium for media item generation
WO2025218280A1 (zh) 发布作品的方法、装置、设备和存储介质
CN119847401A (zh) 媒体编辑的方法、装置、设备和存储介质
CN118550625A (zh) 生成媒体内容方法、装置、设备和存储介质
WO2025228337A1 (zh) 生成媒体内容的方法、装置、设备和存储介质
WO2026001452A1 (zh) 生成媒体内容的方法、装置、设备和存储介质
WO2025228334A1 (zh) 媒体内容生成方法、装置、设备和存储介质
CN119336203B (zh) 信息处理的方法、装置、设备和存储介质
WO2026046223A1 (zh) 基于模板的训练方法、装置、设备和存储介质
WO2025245783A1 (zh) 媒体内容生成方法、装置、设备和存储介质
WO2026012153A1 (zh) 生成媒体内容的方法、装置、设备和存储介质
WO2025218620A1 (zh) 媒体编辑方法、装置、设备和存储介质
WO2025232728A1 (zh) 生成媒体内容的方法、装置、设备和存储介质
WO2026002244A1 (zh) 交互方法、装置、设备和存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25797458

Country of ref document: EP

Kind code of ref document: A1