WO2021180155A1 - 对图片、视频进行语音标记的方法及装置 - Google Patents

对图片、视频进行语音标记的方法及装置 Download PDF

Info

Publication number
WO2021180155A1
WO2021180155A1 PCT/CN2021/080145 CN2021080145W WO2021180155A1 WO 2021180155 A1 WO2021180155 A1 WO 2021180155A1 CN 2021080145 W CN2021080145 W CN 2021080145W WO 2021180155 A1 WO2021180155 A1 WO 2021180155A1
Authority
WO
WIPO (PCT)
Prior art keywords
voice
mark
interface
response
text
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/080145
Other languages
English (en)
French (fr)
Inventor
王中
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba Group Holding Ltd
Original Assignee
Alibaba Group Holding Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba Group Holding Ltd filed Critical Alibaba Group Holding Ltd
Publication of WO2021180155A1 publication Critical patent/WO2021180155A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/78Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/7867Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using information manually generated, e.g. tags, keywords, comments, title and artist information, manually generated time, location and usage information, user ratings
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/78Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • GPHYSICS
    • G11INFORMATION STORAGE
    • G11BINFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
    • G11B27/00Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
    • G11B27/02Editing, e.g. varying the order of information signals recorded on, or reproduced from, record carriers
    • G11B27/031Electronic editing of digitised analogue information signals, e.g. audio or video signals
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/84Generation or processing of descriptive data, e.g. content descriptors

Definitions

  • One or more embodiments of this specification relate to the field of computer technology, and in particular to a method and device for voice tagging a picture, a method and device for voice tagging a video, and a method and device for text tagging a picture , A method and device for text tagging a video, a method and device for viewing pictures, and a method and device for viewing videos.
  • One or more embodiments of this specification describe a method for voice tagging a picture. By recording any position on the picture, that is, the voice tagging function, the picture can be explained and illustrated conveniently and quickly.
  • a method for voice tagging a picture comprising: displaying a recording interface including a target picture; continuously collecting voice signals in response to a recording start instruction issued based on the recording interface; The recording end instruction issued by the recording interface stores the collected voice signal as an audio file; and the voice mark associated with the audio file is added to the target picture.
  • a picture viewing method comprising: displaying a picture with a voice mark, the voice mark being added to the picture by the method provided in the first aspect; Trigger the instruction to play the audio file associated with the voice tag.
  • a method for voice tagging a video including: displaying a recording interface including a first video, the first video including a first video frame; Select the instruction to determine the first video frame as the target video frame; in response to the recording start instruction issued based on the recording interface, continue to collect voice signals; in response to the recording end instruction issued based on the recording interface, the The voice signal of is stored as an audio file; the voice mark associated with the audio file is added to the target video frame.
  • a video viewing method comprising: displaying a video with a voice mark, the voice mark is added to the video by the method provided in the third aspect, and the video includes A first video frame of a voice mark; in response to a triggering instruction to the first voice mark, an audio file associated with the first voice mark is played.
  • a method for text tagging a picture comprising: displaying an editing interface for a target picture; continuously collecting voice signals in response to a recording start instruction issued based on the editing interface; The recording end instruction issued by the editing interface performs voice recognition on the collected voice signal to obtain a recognized text; and a text mark associated with the recognized text is added to the target picture.
  • a picture viewing method comprising: displaying a picture with a text mark, the text mark being added to the picture by the method provided in the fifth aspect;
  • the trigger instruction displays the recognized text associated with the text mark.
  • a method for text tagging a video including: displaying an editing interface including a first video, the first video including a first video frame; Select the instruction to determine the first video frame as the target video frame; in response to the recording start instruction issued based on the editing interface, continue to collect voice signals; in response to the recording end instruction issued based on the editing interface, Perform voice recognition on the voice signal to obtain a recognized text; add a text mark associated with the recognized text on the target video frame.
  • a video viewing method comprising: displaying a video with a text mark, the voice mark is added to the video by the method provided in the seventh aspect, and the video includes A first video frame of a text mark; in response to a trigger instruction to the first text mark, the recognized text associated with the text mark is displayed.
  • a device for voice tagging a picture comprising: a display unit configured to display a recording interface including a target picture; and a collection unit configured to respond to a recording start instruction issued based on the recording interface , Continuously collecting voice signals; a storage unit, configured to store the collected voice signals as an audio file in response to a recording end instruction issued based on the recording interface; and an adding unit, configured to add the associated voice signal to the target picture The voice tag of the audio file.
  • a picture viewing device comprising: a display unit configured to display a picture with a voice mark, the voice mark being added to the picture by the device provided in the ninth aspect; a playing unit, It is configured to play the audio file associated with the voice mark in response to a trigger instruction to the voice mark.
  • an apparatus for voice tagging a video comprising: a display unit configured to display a recording interface including a first video, the first video including a first video frame; a determining unit, configured In response to a selection instruction for the first video frame, determining the first video frame as a target video frame; the acquisition unit is configured to continuously collect voice signals in response to a recording start instruction issued based on the recording interface; A storage unit configured to store the collected voice signal as an audio file in response to a recording end instruction issued based on the recording interface; an adding unit configured to add a voice mark associated with the audio file on the target video frame .
  • a video viewing device comprising: a display unit configured to display a video with a voice mark, the voice mark being added to the video by the device provided in the eleventh aspect, and The video includes a first video frame with a first voice mark; the playback unit is configured to play an audio file associated with the first voice mark in response to a trigger instruction to the first voice mark.
  • a device for text marking a picture comprising: a display unit configured to display an editing interface for a target picture; and a collection unit configured to respond to the start of a recording based on the editing interface Instruction to continuously collect the voice signal; the recognition unit is configured to perform voice recognition on the collected voice signal in response to the recording end instruction issued based on the editing interface to obtain the recognized text; the adding unit is configured to be on the target picture Add a text mark associated with the recognized text.
  • a picture viewing device comprising: a display unit configured to display a picture with a text mark, the text mark being added to the picture by the device provided in the thirteenth aspect;
  • the unit is configured to display the recognized text associated with the text mark in response to a trigger instruction to the text mark.
  • an apparatus for text marking a video comprising: a display unit configured to display an editing interface including a first video, the first video including a first video frame; a determining unit, configured In response to the selection instruction of the first video frame, the first video frame is determined as the target video frame; the acquisition unit is configured to continuously collect voice signals in response to the recording start instruction issued based on the editing interface; The recognition unit is configured to perform voice recognition on the collected voice signal in response to the recording end instruction issued based on the editing interface to obtain recognized text; the adding unit is configured to add the associated recognized text to the target video frame Text tag.
  • a video viewing device comprising: a display unit configured to display a video with a text mark, and the voice mark is added to the video by the device provided in the fifteenth aspect, and The video includes a first video frame with a first text mark; the display unit is configured to display the recognized text associated with the text mark in response to a trigger instruction to the first text mark.
  • a picture processing method including: displaying a chat interface, and receiving a target picture to be sent selected based on the chat interface; entering the picture editing interface in response to an editing instruction for the target picture, A voice mark icon is displayed; in response to a trigger instruction to the voice mark icon, enter the recording interface; store the voice signal collected based on the recording interface as an audio file, and add the associated audio file to the target picture Voice tag.
  • a picture processing method including: displaying a chat interface, the chat window of the chat interface contains a target picture; in response to a trigger instruction for the target picture, displaying a menu bar, the menu bar Include a voice mark icon; in response to a trigger instruction to the voice mark icon, enter the recording interface; store the voice signal collected based on the recording interface as an audio file, and add the associated audio file to the target picture Voice tag.
  • a picture processing method including: displaying a chat interface, the chat window of the chat interface contains a target picture with a first voice tag, and the first voice tag is corresponding to the chat window The current contact is added; in response to the trigger instruction to the first voice mark, display a menu bar, the menu bar includes a voice reply icon; in response to the trigger instruction to the voice reply icon, enter the recording interface;
  • the voice signal collected based on the recording interface is stored as an audio file, and a second voice mark associated with the audio file is added in the area adjacent to the first voice mark in the target picture; or, based on the The voice signal collected by the recording interface is added to the audio file corresponding to the first voice mark.
  • a picture processing method including: displaying a picture editing interface containing a target picture, and a function menu of the picture editing interface includes a voice mark icon; in response to a trigger instruction to the voice mark icon, Enter the voice mark interface; convert the input text received based on the voice mark interface into an audio file, and add a voice mark associated with the audio file on the target picture.
  • a picture processing method including: displaying a chat interface containing a target picture; in response to a trigger instruction to the target picture, displaying a menu bar, the menu bar including a voice adding emoticon icon; In response to a trigger instruction to add an emoticon icon to the voice, enter the voice emoticon interface; convert the voice signal collected based on the voice emoticon interface into text, and generate an animated emoticon based on the text; in the target picture Add the animated emoticon.
  • an image processing method the execution subject of the method is an e-commerce platform, and the method includes: displaying a product information editing interface, which includes a target image for the target product; The trigger instruction of the target picture, the menu bar is displayed, and the menu bar includes the voice mark icon; in response to the trigger instruction to the voice mark icon, enter the recording interface; store the voice signal collected based on the recording interface as an audio file , And add a voice mark associated with the audio file on the target picture.
  • an image processing method including: displaying an order evaluation interface for a first order, which includes adding a picture icon; receiving a selected target picture in response to a trigger instruction for the adding of the picture icon; In response to the voice mark instruction issued to the target picture, enter the recording interface; store the voice signal collected based on the recording interface as an audio file, and add a voice mark associated with the audio file on the target electronic file.
  • an image processing method including: displaying a product evaluation interface for a target product, which includes a first user evaluation, and the first user evaluation includes a target image with a first voice mark; response In response to the trigger instruction for the target picture, a menu bar is displayed, the menu bar includes a voice reply icon; in response to the trigger instruction to the voice reply icon, enters the recording interface; based on the voice signal collected by the recording interface It is stored as an audio file, and a second voice mark associated with the audio file is added in the area adjacent to the first voice mark in the target picture; or, the voice signal collected based on the recording interface is added to In the audio file corresponding to the first voice tag.
  • an image processing method the execution subject of the method is a customer service platform, and the method includes: receiving a conversation message sent by a user, the conversation message including a target with a first voice mark Picture; Acquire an audio file associated with the first voice tag, perform voice recognition on the audio file to obtain a recognized text; input the recognized text into a pre-trained user questioning prediction model, and output the corresponding user standard question ; Feedback to the user the answer to the question corresponding to the user's standard question.
  • an electronic file processing method including: displaying a file processing interface for a target electronic file, the function menu bar of the file processing interface includes a voice mark icon; The icon trigger instruction enters the recording interface; the voice signal collected based on the recording interface is stored as an audio file, and the voice mark associated with the audio file is added to the target electronic file.
  • an image processing method the execution subject of the method is a live broadcast platform, and the method includes: displaying a product information editing interface, which includes a target image for a target product to be put on the shelf;
  • the trigger instruction of the target picture displays a menu bar, the menu bar includes a voice mark icon; in response to the trigger instruction to the voice mark icon, enters the recording interface; and stores the voice signal collected based on the recording interface as Audio file, and adding a voice mark associated with the audio file on the target picture.
  • a picture processing device including: a display unit configured to display a chat interface; a receiving unit configured to receive a target picture to be sent selected based on the chat interface; and a first interface switching unit, Configured to enter the picture editing interface in response to an editing instruction for the target picture, in which a voice mark icon is displayed; a second interface switching unit, configured to enter the recording interface in response to a trigger instruction to the voice mark icon; marking unit And configured to store the voice signal collected based on the recording interface as an audio file, and add a voice mark associated with the audio file on the target picture.
  • an image processing device including: an interface display unit configured to display a chat interface, and a chat window of the chat interface contains a target picture; and a menu bar display unit configured to respond to the The trigger instruction of the target picture displays the menu bar, and the menu bar includes the voice mark icon; the interface switching unit is configured to enter the recording interface in response to the trigger instruction of the voice mark icon; the mark unit is configured to be based on the The voice signal collected by the recording interface is stored as an audio file, and a voice mark associated with the audio file is added to the target picture.
  • a picture processing device including: an interface display unit configured to display a chat interface, the chat window of the chat interface contains a target picture with a first voice tag, and the first voice tag Added by the current contact corresponding to the chat window; a menu bar display unit configured to display a menu bar in response to a trigger instruction to the first voice mark, the menu bar including a voice reply icon; an interface switching unit, It is configured to enter the recording interface in response to a trigger instruction to the voice reply icon; the marking unit is configured to store the voice signal collected based on the recording interface as an audio file, and is adjacent to the first image in the target picture. In the area of the voice mark, a second voice mark associated with the audio file is added; or, the voice signal collected based on the recording interface is added to the audio file corresponding to the first voice mark.
  • a picture processing device comprising: a display unit configured to display a picture editing interface containing a target picture, the function menu of the picture editing interface includes a voice mark icon; an interface switching unit , Configured to enter the voice mark interface in response to a trigger instruction to the voice mark icon; the marking unit is configured to convert input text received based on the voice mark interface into an audio file, and add an association to the target picture The voice tag of the audio file.
  • a picture processing device including: an interface display unit configured to display a chat interface containing a target picture; a menu bar display unit configured to display a menu in response to a trigger instruction to the target picture
  • the menu bar includes the voice adding emoticon icon;
  • the interface switching unit is configured to enter the voice adding emoticon interface in response to a trigger instruction for adding the emoticon icon to the voice;
  • the emoticon generating unit is configured to add emoticons based on the voice
  • the voice signal collected by the emoticon interface is converted into text, and an animated emoticon is generated based on the text;
  • the emoticon adding unit is configured to add the animated emoticon to the target picture.
  • an image processing device integrated in an e-commerce platform, the device comprising: an interface display unit configured to display a product information editing interface, which includes a target picture for the target product; and a menu;
  • the bar display unit is configured to display a menu bar in response to a trigger instruction to the target picture, and the menu bar includes a voice mark icon;
  • the interface switching unit is configured to enter in response to a trigger instruction to the voice mark icon A recording interface;
  • a marking unit configured to store the voice signal collected based on the recording interface as an audio file, and add a voice mark associated with the audio file on the target picture.
  • a picture processing device including: a display unit configured to display an order evaluation interface for a first order, including an add picture icon; a receiving unit, configured to respond to the added picture icon To receive the selected target picture; the interface switching unit is configured to enter the recording interface in response to the voice marking instruction issued to the target picture; the marking unit is configured to store the voice signal collected based on the recording interface as Audio file, and adding a voice mark associated with the audio file to the target electronic file.
  • an image processing device including: an interface display unit configured to display a product evaluation interface for a target product, which includes a first user evaluation, and the first user evaluation includes a first voice A marked target picture; a menu bar display unit configured to display a menu bar in response to a trigger instruction for the target picture, the menu bar including a voice reply icon; an interface switching unit, configured to respond to the voice reply The icon triggers the instruction to enter the recording interface; the marking unit is configured to store the voice signal collected based on the recording interface as an audio file, and add an associated location in the area adjacent to the first voice mark in the target picture The second voice mark of the audio file; or, adding the voice signal collected based on the recording interface to the audio file corresponding to the first voice mark.
  • an image processing device integrated in a customer service platform, the device comprising: a receiving unit configured to receive a conversation message sent by a user, the conversation message including a first voice A marked target picture; an obtaining unit configured to obtain an audio file associated with the first voice mark, perform voice recognition on the audio file, and obtain a recognized text; a prediction unit configured to input the recognized text into a pre-trained In the user mark prediction model, the corresponding user standard question is output; the feedback unit is configured to feed back the answer to the question corresponding to the user standard question to the user.
  • an electronic file processing device including: a display unit configured to display a file processing interface for a target electronic file, the function menu bar of the file processing interface includes a voice mark icon; interface switching Unit, configured to enter the recording interface in response to a trigger instruction to the voice mark icon; marking unit, configured to store the voice signal collected based on the recording interface as an audio file, and add an association to the target electronic file The voice tag of the audio file.
  • a picture processing device integrated in a live broadcast platform, the device comprising: an interface display unit configured to display a product information editing interface, which includes a target picture for a target product to be put on the shelf A menu bar display unit, configured to display a menu bar in response to a trigger instruction to the target picture, the menu bar including a voice mark icon; an interface switching unit, configured to respond to a trigger instruction to the voice mark icon Enter the recording interface; the marking unit is configured to store the voice signal collected based on the recording interface as an audio file, and add a voice mark associated with the audio file on the target picture.
  • a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed in a computer, the computer is caused to execute the first to eighth and seventeenth aspects above.
  • a computing device including a memory and a processor, the memory stores executable code, and when the processor executes the executable code, the first to eighth aspects are implemented.
  • Fig. 1 shows a schematic diagram of an application scenario of a voice marking function according to an embodiment
  • Fig. 2 shows a flowchart of a method for voice tagging a picture according to an embodiment
  • Fig. 3 shows a schematic diagram of switching to a recording interface according to an embodiment
  • FIG. 4 shows a schematic diagram of switching to a recording interface according to another embodiment
  • FIG. 5 shows a schematic diagram of interface switching when using the voice mark function according to an embodiment
  • FIG. 6 shows a schematic diagram of interface switching when using the voice mark function according to another embodiment
  • FIG. 7 shows a schematic diagram of interface switching of a serial number modification in a voice mark icon according to an embodiment
  • FIG. 8 shows a schematic diagram of an interface in which the voice mark icon contains the subject text according to an embodiment
  • Fig. 9 shows a schematic diagram of interface switching when playing an audio file according to an embodiment
  • FIG. 10 shows a schematic diagram of interface switching related to voice recognition according to an embodiment
  • Fig. 11 shows a flowchart of a picture viewing method according to an embodiment
  • FIG. 12 shows a product detail interface including a product picture with a voice icon according to an embodiment
  • Fig. 13 shows a flowchart of a method for voice tagging a video according to an embodiment
  • Fig. 14 shows a schematic diagram of a voice mark interface for video according to an embodiment
  • FIG. 15 shows a schematic diagram of a voice mark interface for video according to another embodiment
  • Fig. 16 shows a flowchart of a video viewing method according to an embodiment
  • FIG. 17 shows a schematic diagram of interface switching for viewing a voice-marked video according to an embodiment
  • Fig. 18 shows a flowchart of a method for text labeling a picture according to an embodiment
  • FIG. 19 shows a schematic diagram of interface switching when text marking a picture according to an embodiment
  • Fig. 20 shows a flowchart of a picture viewing method according to an embodiment
  • FIG. 21 shows a schematic diagram of interface switching for viewing text marks in a picture according to an embodiment
  • Fig. 22 shows a flowchart of a method for text tagging a video according to an embodiment
  • FIG. 23 shows a schematic diagram of interface switching when text marking a video according to an embodiment
  • FIG. 24 shows a flowchart of a video viewing method according to an embodiment
  • FIG. 25 shows a schematic diagram of interface switching for viewing text marks in a video according to an embodiment
  • Fig. 26 shows a structure diagram of an apparatus for voice tagging a picture according to an embodiment
  • Fig. 27 shows a structural diagram of a picture viewing device according to an embodiment
  • Fig. 28 shows a structural diagram of an apparatus for voice tagging a video according to an embodiment
  • Fig. 29 shows a structural diagram of a video viewing device according to an embodiment
  • Fig. 30 shows a structural diagram of an apparatus for text labeling a picture according to an embodiment
  • FIG. 31 shows a structural diagram of a picture viewing device according to an embodiment
  • Fig. 32 shows a structural diagram of an apparatus for text tagging a video according to an embodiment
  • Fig. 33 shows a structural diagram of a video viewing device according to an embodiment
  • Fig. 34 shows a flowchart of a picture processing method according to an embodiment
  • FIG. 35 shows a schematic diagram of a chat interface according to an embodiment
  • Fig. 36 shows a flowchart of a picture processing method according to another embodiment
  • Fig. 37 shows a schematic diagram of a chat interface according to another embodiment
  • Fig. 38 shows a schematic diagram of a chat interface according to yet another embodiment
  • Fig. 39 shows a flowchart of a picture processing method according to yet another embodiment
  • Fig. 40 shows a schematic diagram of a chat interface according to still another embodiment
  • FIG. 41 shows a schematic diagram of a chat interface according to still another embodiment
  • Fig. 42 shows a flowchart of a picture processing method according to still another embodiment
  • Fig. 43 shows a flowchart of a picture processing method according to an embodiment
  • Fig. 44 shows a schematic diagram of a chat interface according to an embodiment
  • FIG. 45 shows a flowchart of a picture processing method according to another embodiment
  • FIG. 46 shows a schematic diagram of a product information editing interface according to an embodiment
  • Fig. 47 shows a flowchart of a picture processing method according to another embodiment
  • Fig. 48 shows a flowchart of a picture processing method according to still another embodiment
  • Fig. 49 shows a flowchart of a picture processing method according to still another embodiment
  • FIG. 50 shows a flowchart of an electronic file processing method according to an embodiment
  • FIG. 51 shows a schematic diagram of an office software interface according to an embodiment
  • Fig. 52 shows a flowchart of a picture processing method according to an embodiment
  • Fig. 53 shows a schematic diagram of a live broadcast interface according to an embodiment
  • Fig. 54 shows a structural diagram of a picture processing apparatus according to an embodiment
  • FIG. 55 shows a structural diagram of a picture processing apparatus according to another embodiment
  • Fig. 56 shows a structural diagram of a picture processing apparatus according to still another embodiment
  • Fig. 57 shows a structural diagram of a picture processing apparatus according to still another embodiment
  • Fig. 58 shows a structural diagram of a picture processing apparatus according to still another embodiment
  • Fig. 59 shows a structural diagram of a picture processing apparatus according to an embodiment
  • Fig. 60 shows a structural diagram of a picture processing apparatus according to yet another embodiment
  • Fig. 61 shows a structural diagram of a picture processing apparatus according to another embodiment
  • Fig. 62 shows a structural diagram of a picture processing apparatus according to still another embodiment
  • Fig. 63 shows a structural diagram of an electronic file processing apparatus according to an embodiment
  • Fig. 64 shows a structural diagram of a picture processing apparatus according to still another embodiment.
  • FIG. 1 shows a schematic diagram of an application scenario of the voice tag function according to an embodiment.
  • user A may use client A (such as an instant messaging client).
  • client A such as an instant messaging client.
  • Use the voice mark function to add voice marks to several places in the picture.
  • terminal B can listen to the corresponding recording file by clicking the voice mark on it. For example, by clicking the voice mark 10, you can listen to the corresponding recording, and the display state of the voice mark 10 is switched to voice mark 13. , Used to remind the user that the corresponding recording file is currently being played.
  • the voice tagging function proposed by the inventor can fully improve the experience of all users.
  • FIG. 2 shows a flowchart of a method for voice tagging a picture according to an embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, Client (such as image processing software), system software or client plug-in (such as the plug-in in the instant messaging client).
  • Client such as image processing software
  • system software or client plug-in (such as the plug-in in the instant messaging client).
  • client plug-in such as the plug-in in the instant messaging client.
  • the method includes the following steps:
  • Step S210 display the recording interface including the target picture; step S220, continue to collect voice signals in response to the recording start instruction issued based on the recording interface; step S230, in response to the recording end instruction issued based on the recording interface, collect The received voice signal is stored as an audio file; in step S240, a voice mark associated with the audio file is added to the target picture.
  • step S210 the recording interface including the target picture is displayed.
  • this step may include: in response to an import instruction for the target picture issued based on the recording interface, displaying the target picture in the recording interface.
  • the recording interface may be an editing interface in image processing software, or a product information editing interface in an e-commerce platform, or an editing interface in an instant messaging client for pictures to be sent.
  • the import instruction may be a voice control instruction or a click instruction, such as a click instruction to an import icon.
  • the method may further include: in response to a trigger instruction for the target picture, displaying a voice mark icon; correspondingly, this step may include: responding to triggering of the voice mark icon Instruction to jump to the recording interface.
  • the trigger instruction for the target picture may include a viewing instruction, or an editing instruction, or a sending instruction.
  • the trigger instruction for the target picture may include a screenshot instruction for capturing the target picture, that is, the trigger instruction is a screenshot instruction, and the picture obtained by the screenshot is the target picture.
  • the chat interface 31 is shown in FIG. 3, in response to a further operation instruction to the picture 32 therein (which may correspond to a long press operation), an operation menu bar 33 is displayed, which includes a voice mark icon 34.
  • the click instruction on the voice mark icon 34 jump to the recording interface 35.
  • the method may further include: displaying a playback interface for the first video; further, in a more specific
  • displaying a voice mark icon may include: in response to a jump instruction issued for the first video, determining a video frame that is jumped to be displayed as the target picture, and The voice mark icon is displayed on the playback interface.
  • displaying a voice mark icon in response to a trigger instruction for the target picture, may include: in response to a pause instruction issued for the first video, determining a video frame that is paused to be displayed as the Target picture, and display the voice mark icon on the playback interface.
  • a playback interface 41 for the first video is displayed.
  • a voice mark icon 42 is displayed in the playback interface.
  • the click instruction of the mark icon 42 jumps to the recording interface 43 (or called the voice mark interface 43), which includes the target picture 44, which is the video frame in the pause playback interface. In this way, it is possible to jump to the above-mentioned recording interface in response to the triggering instruction of the voice mark icon.
  • this step may further include: displaying prompt information in the recording interface for prompting the user of an operation mode of adding a voice mark to the target picture.
  • the prompt message 36 displayed on the recording interface 35 includes content: Please long press a certain position on the picture to start recording, and a voice mark will be generated at that position.
  • the recording interface including the target picture can be displayed.
  • the voice signal is continuously collected.
  • the collected voice signal is stored as an audio file, and in step S240, a voice mark associated with the audio file is added to the target picture .
  • the above-mentioned recording start instruction and recording end instruction may respectively correspond to multiple operation modes.
  • the recording start instruction may correspond to a long-press operation on the target picture
  • the recording end instruction corresponds to a cancel pressing operation on the target picture.
  • the recording interface includes a recording icon, wherein the recording start instruction may correspond to: a click operation on the recording icon in the first state, and, in response to this click operation, the recording icon is switched and displayed as the first state.
  • the second state correspondingly, where the recording end instruction may correspond to: a click operation on the recording icon in the second state, and, in response to this click operation, the recording icon is switched and displayed as the first state.
  • the recording start instruction corresponds to a click operation on the right mouse button
  • the recording end instruction corresponds to a click operation on the right mouse button again.
  • the recording start instruction corresponds to: a trigger instruction for the recording start icon in the recording interface
  • the recording end instruction corresponds to: a trigger instruction for the recording end icon in the recording interface .
  • step S220 may include: continuously collecting voice signals in response to a recording start instruction issued based on the first position of the target picture in the recording interface.
  • the first position can be any position in the target picture.
  • step S240 may include: adding the voice mark at the first position of the target picture.
  • the recording interface in response to a long-press operation on the first position 52 in the target picture 51, the recording starts, and then, in response to the cancellation of the long-press operation, the recording ends, and in the first position A position 52 adds a voice mark 53 related to the recorded audio. In this way, it is possible to add a voice mark at the designated first position.
  • step S240 may include: adding the voice mark anywhere in the target picture.
  • any position may be a random position determined according to a random algorithm, or a fixed position defaulted by the system, such as the central area of the target picture.
  • the method may further include: in response to a movement instruction to the voice marker, moving the voice marker to a designated position in the target picture.
  • a target picture 61 is included in the recording interface.
  • the recording starts in response to the click instruction on the recording icon 62 in the first state, and then in response to the click on the recording icon 63 in the second state. Instruction, end the recording, and add a voice mark in the central area 64 of the picture. Further, in response to the movement instruction to the voice marker 64, the voice marker 64 is moved to the target area 65. In this way, the voice mark can be displayed at the designated position.
  • step S240 may further include: displaying a serial number in the voice mark.
  • the sequence number may be automatically generated, and may be specifically determined based on the number of prior voice tags added to the target picture.
  • the target picture already contains n (a natural number) voice markers the sequence number displayed on the newly-added voice marker is n+1.
  • the serial number can be customized by the user. According to an example, as shown in FIG. 7, the user changes the serial number displayed on the voice mark 71 from 1 to 2. In this way, it is convenient for the user to learn the recording content corresponding to the voice tag in sequence based on the sequence of the serial number identification.
  • step S240 may further include: displaying the topic text in the voice tag.
  • the above step S230 may further include: performing voice recognition on the collected voice signal to obtain the recognized text; and then determining the corresponding topic text according to the recognized text.
  • the speech recognition can be implemented by using existing technologies, such as a speech recognition model, etc., which will not be repeated here.
  • determining the corresponding topic text according to the recognized text may include: inputting the recognized text into a pre-trained abstract extraction model to obtain the corresponding summary text as the topic text.
  • determining the corresponding topic text according to the recognized text may include: inputting the recognized text into a pre-trained keyword extraction model to obtain the corresponding keywords as the topic text.
  • a pre-trained keyword extraction model to obtain the corresponding keywords as the topic text.
  • the method may further include: receiving a user-defined text input based on the voice mark, and displaying the customized text in the voice mark.
  • the method may further include: playing the audio file in response to a trigger instruction to the voice mark.
  • the trigger instruction may be a voice control instruction or a click instruction.
  • the audio file is played, and at the same time, the voice mark 91 is switched to be displayed as the voice mark 92 in the second state for prompting The user is playing the recorded audio.
  • the method may further include: after adding a voice mark associated with the audio file on the target picture, the method may further include: responding to triggering of the voice mark Instruction to display a menu bar.
  • the menu bar includes a play icon, and in response to a trigger instruction to the play icon, the audio file is played.
  • the menu bar includes a voice recognition icon, and in response to a trigger instruction to the voice recognition icon, voice recognition is performed on the audio file to obtain recognized text; in the recording interface The recognized text is displayed.
  • the method after displaying the recognized text in the recording interface, the method may further include: in response to a hiding instruction for the recognized text, hiding the recognized text in the recording interface text.
  • a menu bar 102 is displayed, which includes a voice recognition icon 103, and in response to a trigger instruction on the voice recognition icon 103, a recognized text 104 is displayed, and further , In response to a click instruction on the folding icon 105, hide the recognized text.
  • the method may further include: deleting the voice mark from the target picture in response to a deletion instruction for the voice mark.
  • a delete button is displayed in response to a long-press instruction on the voice mark, and the voice mark is deleted in response to a click instruction on the delete button. In this way, the voice mark can be deleted.
  • the user can easily and quickly add a voice tag to any position in the picture that needs to be marked, which greatly reduces the cost of picture editing and greatly improves users.
  • the user can easily and quickly add a voice tag to any position in the picture that needs to be marked, which greatly reduces the cost of picture editing and greatly improves users.
  • FIG. 11 shows a flowchart of a picture viewing method according to an embodiment.
  • the execution subject of the method can be any device, device, platform, server cluster with computing and processing capabilities, for example, a client (such as Image processing software), system software or client plug-in (such as the plug-in in the instant messaging client).
  • client such as Image processing software
  • system software or client plug-in (such as the plug-in in the instant messaging client).
  • client plug-in such as the plug-in in the instant messaging client.
  • the method includes the following steps:
  • step S1110 the picture with the voice mark is displayed. It needs to be understood that the voice therein is added with the method described in the foregoing embodiment.
  • Step S1120 in response to the trigger instruction to the voice mark, play the audio file associated with the voice mark.
  • step S1110 a picture with a voice mark is displayed.
  • this step may include: displaying the picture in the chat window.
  • the picture in a scenario of instant messaging (such as online social networking or online customer service consultation), the picture can be received from other instant messaging clients and displayed in the chat window.
  • this step may include: loading the picture in a web page.
  • the webpage in response to an instruction to open the webpage, enter the webpage and load the picture in the webpage.
  • the webpage may be a product detail page in an e-commerce platform, and correspondingly, the main object in the picture may be a product.
  • the product detail page shown therein includes a product picture 121 with voice tags.
  • the webpage may be a teaching guidance website, and correspondingly, the main object in the picture may be a test paper.
  • step S1120 in response to a trigger instruction to the voice mark, the audio file associated with the voice mark is played.
  • this step may include: based on the order in which the multiple voice tags are added to the picture, sequentially playing the voice tags corresponding to the multiple voice tags. Multiple audio files.
  • this step may include: playing the multiple voices in sequence based on the sequence of the serial numbers Mark the corresponding multiple audio files.
  • step S1120 can also be replaced with: in response to the trigger instruction to the voice mark, display the menu bar, which includes a voice recognition icon, and in response to the trigger instruction to the voice recognition icon, convert the audio file corresponding to the voice mark It is the speech recognition text, and the speech recognition text is displayed.
  • the use of the picture viewing method disclosed in the embodiments of this specification allows users to intuitively and conveniently obtain the content they need to know based on the picture, thereby improving user experience.
  • FIG. 13 shows a flowchart of a method for voice tagging a video according to an embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, Client (such as video processing software), system software or client plug-in (such as plug-in in instant messaging client).
  • Client such as video processing software
  • system software such as client plug-in
  • client plug-in such as plug-in in instant messaging client
  • Step S1310 display the recording interface including the first video, the first video including the first video frame; step S1320, in response to the selection instruction of the first video frame, determine the first video frame as the target video Frame; step S1330, in response to the recording start instruction issued based on the recording interface, continue to collect voice signals; step S1340, in response to the recording end instruction issued based on the recording interface, store the collected voice signals as an audio file; Step S1350: Add a voice mark associated with the audio file on the target video frame.
  • step S1310 a recording interface including a first video is displayed, and the first video includes a first video frame.
  • the method may further include: in response to the trigger instruction for the first video, displaying a voice mark icon. Accordingly, this step may include: responding to the voice mark icon To jump to the recording interface.
  • the trigger instruction for the first video may include a viewing instruction, or an editing instruction, or a sending instruction.
  • this step may include: in response to an import instruction for the first video issued based on the recording interface, displaying the first video in the recording interface.
  • the recording interface may be an editing interface in video processing software.
  • the import instruction may be a voice control instruction or a click instruction, such as a click instruction to an import icon.
  • the user's instruction to open the video processing APP enter the above-mentioned recording interface, in response to the click instruction on the video import icon in the recording interface, display the existing video in the terminal, and then respond to the user's
  • the selection instruction and the confirmation import instruction of a certain video are displayed on the above-mentioned recording interface.
  • step S1320 in response to the selection instruction of the first video frame, the first video frame is determined as the target video frame.
  • this step may include: in response to a jump instruction issued for the first video, determining the first video frame that is jumped to be displayed as the target video frame.
  • the jump instruction may correspond to: a drag operation on the video progress bar.
  • the jump instruction may correspond to: an input operation on the playback time point of the first video frame.
  • this step may include: in response to a pause instruction issued for the first video, determining the first video frame that is paused to be displayed as the target video frame.
  • the method before receiving the pause instruction, the method may further include: playing the first video.
  • the first video frame 142 displayed by the jump is determined as the target video frame. In this way, the target video frame can be determined.
  • step S1330 in response to the recording start instruction issued based on the recording interface, continue to collect the voice signal; step S1340, in response to the recording end instruction issued based on the recording interface, store the collected voice signal as an audio file ; Step S1350, adding a voice mark associated with the audio file on the target video frame.
  • step S1330 may include: continuously collecting voice signals in response to a recording start instruction issued based on the first position of the target video frame in the recording interface; accordingly, step S1350 may include: The voice mark is added to the first position of the frame.
  • step S1330 to step S1350 please refer to the description of step S220 to step S240 in the foregoing embodiment.
  • the method may further include: adding a time point mark to the progress bar of the first video, the time point mark corresponding to the playback time of the target video frame point.
  • the shape of the time point mark is not limited, and it may be a circle, a rectangle, a triangle, or the like. In this way, it can assist the user to quickly locate the video frame to be voice marked.
  • the method may further include: in response to the input control indicator moving to the time point mark, displaying all The number of voice tags and/or subject text included in the target video frame.
  • the input control indicator may be a cursor or an indicator generated by the user through a touch operation on the screen.
  • the number of voice marks and/or subject text can be displayed in a pop-up window or bubble.
  • the number of voice markers included in the corresponding video frame and the topic text corresponding to each voice marker are displayed in the chat bubble 152, such as the voice marker Quantity: 3.
  • the subject texts are: city gates, city walls and ancient bells. In this way, voice marking and verification of video frames in the video can be realized.
  • the method for voice tagging a video disclosed in the embodiments of this specification can enable users to add voice tags to any position in any video frame of the video conveniently and quickly, so as to enrich the way of editing and processing the video. And, it saves the time needed to intercept and edit video frames, etc., so as to improve user experience.
  • FIG. 16 shows a flowchart of a video viewing method according to an embodiment.
  • the execution subject of the method can be any device, device, platform, or device cluster with computing and processing capabilities, for example, it can be a client (such as Video processing software), system software or client plug-in (such as the plug-in in the instant messaging client).
  • client such as Video processing software
  • system software such as the plug-in in the instant messaging client.
  • client plug-in such as the plug-in in the instant messaging client
  • Step S1610 Display a video with a voice mark, and the video includes a first video frame with a first voice mark. It should be understood that the voice mark is added by the method of voice mark on the video described in the foregoing embodiment.
  • Step S1620 in response to the trigger instruction to the first voice mark, play the audio file associated with the first voice mark.
  • step S1610 a video with a voice mark is displayed, which includes a first video frame with a first voice mark.
  • this step may include: displaying the video in a chat window.
  • the video in the context of instant messaging (such as online social networking or online customer service consultation), the video can be received from other instant messaging clients and displayed in the chat window.
  • this step may include: loading the video in a web page.
  • in response to an instruction to open the webpage enter the webpage, and load the video in the webpage.
  • the webpage may be a product detail page in an e-commerce platform.
  • the webpage may be a teaching guidance website.
  • the webpage may be a playback page provided by a video playback platform.
  • step S1620 in response to the trigger instruction to the first voice mark, the audio file associated with the first voice mark is played.
  • the associated audio file is played.
  • step S1620 can also refer to the aforementioned related description of step 1120 and so on.
  • the first video frame includes several voice marks including the first voice mark
  • the progress bar of the video displays a time point mark at the playback time point of the first video frame
  • the method may further include: in response to the input control indicator moving to the time point mark, displaying the number of marks of the several voice marks and/or several themes corresponding to the several voice marks text.
  • adopting the video viewing method disclosed in the embodiments of this specification allows the user to intuitively and conveniently obtain the specific description content for some video frames, thereby improving the user experience.
  • the method for adding voice marks to multimedia media (such as pictures and videos, which may actually include electronic documents, etc.) and viewing multimedia media with voice marks is mainly described. Furthermore, the inventor proposes that it can also be extended to the category of text markup. Specifically, adding text markup can be realized by combining the recording function and speech recognition technology.
  • FIG. 18 shows a flowchart of a method for text labeling a picture according to an embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, it may be a client (such as image processing software), system software or client plug-in (such as the plug-in in the instant messaging client).
  • client Such as image processing software
  • system software or client plug-in (such as the plug-in in the instant messaging client).
  • the method includes the following steps:
  • Step S1810 displaying the editing interface for the target picture; step S1820, continuously collecting voice signals in response to the recording start instruction issued based on the editing interface; step S1830, responding to the recording end instruction issued based on the editing interface, Perform voice recognition on the received voice signal to obtain recognized text; step S1840, add a text mark associated with the recognized text on the target picture.
  • step S1810 an editing interface for the target picture is displayed.
  • the target picture can be imported into the picture processing or picture editing software, and then the editing interface for the target picture is displayed.
  • the target picture may be selected in the gallery first, and then the edit icon may be clicked to enter the editing interface containing the target picture.
  • the editing interface here can also provide other image processing functions, such as graffiti.
  • step S1820 in response to the recording start instruction issued based on the editing interface, continue to collect the voice signal; step S1830, in response to the recording end instruction issued based on the editing interface, perform voice recognition on the collected voice signal , Obtain the recognized text; step S1840, add a text mark associated with the recognized text on the target picture.
  • step S1820 may include: in response to the recording start instruction issued based on the first position of the target picture in the editing interface, continuously collecting voice signals.
  • step S1840 may also include: Add the text mark to the first position of the target picture.
  • step S1840 may further include: displaying the recognized text in the text mark.
  • step S1840 may further include: folding and displaying the recognized text in the text mark.
  • the method may further include: in response to a trigger instruction for the text mark, expanding and displaying the recognized text in the text mark.
  • step S1840 may further include: displaying a serial number in the text mark, the serial number being determined based on the number of previous text marks added to the target picture.
  • step S1830 may further include: determining the subject text corresponding to the recognized text, and accordingly, step S1840 may further include: displaying the subject text in the text mark.
  • the method may further include: displaying the recognized text in response to a trigger instruction for marking the text.
  • the serial number 3 is displayed in the newly added text mark 191, and in response to the click instruction on the text mark 191, the corresponding recognized text 192 is displayed.
  • FIG. 20 shows a flowchart of a picture viewing method according to an embodiment.
  • the device, equipment, platform, and server cluster with processing capabilities may be a client (such as image processing software), system software, or a client plug-in (such as a plug-in in an instant messaging client).
  • client such as image processing software
  • system software such as system software
  • client plug-in such as a plug-in in an instant messaging client
  • Step S2010 displaying a picture with a text mark, requires understanding, where the text mark is added based on the method described in the foregoing embodiment.
  • Step S2020 in response to the trigger instruction for the text mark, display the recognized text associated with the text mark.
  • step S2010 a picture with a text mark is displayed.
  • this step may further include: displaying a serial number in the text mark.
  • this step may further include: displaying the subject text corresponding to the recognized text in the text mark.
  • this step may further include: folding and displaying the recognized text in the text mark.
  • step S2020 in response to the trigger instruction to the text mark, the recognized text associated with the text mark is displayed.
  • this step may include: expanding and displaying the recognized text in the text mark.
  • the method may further include: in response to the retracting instruction of the recognized text, restoring and displaying the text mark.
  • the corresponding recognized text 212 is expanded and displayed, and further, in response to the click instruction on the folding icon 213, the recognized text is folded show.
  • the use of the picture viewing method disclosed in the embodiments of this specification allows users to intuitively and conveniently obtain the content they need to know based on the picture, thereby improving user experience.
  • FIG. 22 shows a flowchart of a method for text tagging a video according to an embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, Client (such as video processing software), system software or client plug-in (such as plug-in in instant messaging client).
  • Client such as video processing software
  • system software such as client plug-in
  • client plug-in such as plug-in in instant messaging client
  • Step S2210 display an editing interface including a first video, the first video including a first video frame;
  • Step S2220 in response to a selection instruction for the first video frame, determine the first video frame as a target video Frame;
  • Step S2230 in response to the recording start instruction issued based on the editing interface, continue to collect voice signals;
  • Step S2240 in response to the recording end instruction issued based on the editing interface, perform voice recognition on the collected voice signals to obtain Recognize the text;
  • Step S2250 add a text mark associated with the recognized text on the target video frame.
  • step S2230 may include: in response to the recording start instruction issued based on the first position of the target video frame in the editing interface, continuously collecting voice signals, accordingly, step S2250 may The method includes: adding the text mark at the first position of the target video frame.
  • a text mark 231 is added to a certain video frame in the video, and in response to a click instruction on the text mark 231, the corresponding recognized text 232 is displayed.
  • step S2210 to step S2250, reference may also be made to the relevant description in the foregoing embodiment.
  • the method for text marking a video disclosed in the embodiments of this specification can enable users to add text markings to any position in any video frame of the video conveniently and quickly, so as to enrich the way of editing the video. And, it saves the time needed to intercept and edit video frames, etc., so as to improve user experience.
  • FIG. 24 shows a flowchart of a video viewing method according to an embodiment.
  • the execution subject of the method can be any device, device, platform, or device cluster with computing and processing capabilities, for example, it can be a client (such as Video processing software), system software or client plug-in (such as the plug-in in the instant messaging client).
  • client such as Video processing software
  • system software such as the plug-in in the instant messaging client.
  • client plug-in such as the plug-in in the instant messaging client
  • Step S2410 displaying a video with a text mark
  • the video includes a first video frame with a first text mark, which needs to be explained, wherein the text mark is added by the method in the foregoing embodiment
  • step S2420 in response to the The trigger instruction of the first text mark displays the recognized text associated with the text mark.
  • the corresponding recognized text 252 is displayed in a position other than the video frame in the video playback interface.
  • adopting the video viewing method disclosed in the embodiments of this specification allows the user to intuitively and conveniently obtain the specific description content for some video frames, thereby improving the user experience.
  • Fig. 26 shows a structure diagram of an apparatus for voice tagging a picture according to an embodiment.
  • the device 2600 includes:
  • the display unit 2610 displays the recording interface including the target picture; the collection unit 2620 continuously collects voice signals in response to the recording start instruction issued based on the recording interface; the storage unit 2630 responds to the recording end instruction issued based on the recording interface , Storing the collected voice signal as an audio file; adding unit 2640, adding a voice mark associated with the audio file on the target picture.
  • the display unit 2610 is specifically configured to display the target picture in the recording interface in response to an import instruction for the target picture issued based on the recording interface.
  • the device 2600 further includes: a trigger unit configured to display a voice mark icon in response to a trigger instruction for the target picture; wherein the display unit 2610 is specifically configured to: respond to triggering of the voice mark icon Instruction to jump to the recording interface.
  • a trigger unit configured to display a voice mark icon in response to a trigger instruction for the target picture
  • the display unit 2610 is specifically configured to: respond to triggering of the voice mark icon Instruction to jump to the recording interface.
  • the trigger unit is specifically configured to display the voice mark icon in response to a viewing instruction for the target picture; or, in response to an editing instruction for the target picture, display the voice The mark icon, or, in response to a screenshot instruction for capturing the target picture, display the voice mark icon; or, in response to a sending instruction for the target picture, display the voice mark icon.
  • the device 2600 further includes: a playback interface display unit configured to display a playback interface for the first video; wherein the trigger unit is specifically configured to: respond to a jump sent for the first video.
  • a transfer instruction determining the video frame displayed by the jump as the target picture, and displaying the voice mark icon on the playback interface; or, in response to a pause instruction issued for the first video, pause the displayed video
  • the frame is determined as the target picture, and the voice mark icon is displayed on the playback interface.
  • the display unit 2610 is further configured to display prompt information in the recording interface for prompting the user of an operation mode of adding a voice mark to the target picture.
  • the collecting unit 2620 is specifically configured to continuously collect voice signals in response to a recording start instruction issued based on the first position of the target picture in the recording interface; wherein the adding unit 2640 is specifically configured to: The voice tag is added to the first position of the picture.
  • the recording start instruction corresponds to: a long press operation on the target picture
  • the recording end instruction corresponds to: a cancel press operation on the target picture
  • the adding unit 2640 is further configured to: display a serial number in the voice mark, the serial number being determined based on the number of previous voice marks added to the target picture or inputted by a user. .
  • the storage unit 2630 is further configured to: perform voice recognition on the collected voice signals to obtain recognized text; determine the subject text corresponding to the recognized text; wherein the adding unit 2640 is further configured to: The subject text is displayed in the voice tag.
  • the storage unit 2630 is specifically configured to: input the recognized text into a pre-trained abstract extraction model to obtain the corresponding abstract text as the topic text; or, input the recognized text In the pre-trained keyword extraction model, the corresponding keywords are obtained as the topic text.
  • the device 2600 further includes: a receiving unit configured to receive a user-defined text input based on the voice mark; a first text display unit configured to display the customized text in the voice mark text.
  • the device 2600 further includes: a playing unit configured to play the audio file in response to a trigger instruction to the voice mark.
  • the device 2600 further includes: a moving unit configured to move the voice mark to a designated position in the target picture in response to a movement instruction to the voice mark.
  • the device 2600 further includes a deleting unit configured to delete the voice mark from the target picture in response to a deletion instruction for the voice mark.
  • the device 2600 further includes: a menu bar display unit configured to display a menu bar in response to a trigger instruction to the voice mark, the menu bar including a voice recognition icon; and a recognition unit configured to In response to the trigger instruction for the voice recognition icon, perform voice recognition on the audio file to obtain recognized text; the second text display unit is configured to display the recognized text on the recording interface.
  • a menu bar display unit configured to display a menu bar in response to a trigger instruction to the voice mark, the menu bar including a voice recognition icon
  • a recognition unit configured to In response to the trigger instruction for the voice recognition icon, perform voice recognition on the audio file to obtain recognized text
  • the second text display unit is configured to display the recognized text on the recording interface.
  • the device 2600 further includes a hiding unit configured to hide the recognized text in the recording interface in response to a hiding instruction for the recognized text.
  • Fig. 27 shows a structural diagram of a picture viewing device according to an embodiment. As shown in FIG. 27, the device 2700 includes:
  • the display unit 2710 is configured to display a picture with a voice mark, and the voice mark is added to the picture by the above-mentioned device 2600; the playback unit 2720 is configured to play the voice in response to a trigger instruction to the voice mark Mark the associated audio file.
  • the display unit 2710 is specifically configured to: display the picture in the chat window; or load the picture in a web page.
  • the picture includes the target product, and the webpage is a product detail page.
  • the voice markers are multiple voice markers
  • the playing unit 2720 is specifically configured to: based on the order in which the multiple voice markers are added to the picture, sequentially play the voice markers corresponding to the multiple voice markers. Multiple audio files.
  • the voice markers are multiple voice markers, wherein each voice marker displays a corresponding sequence number; wherein the playing unit 2720 is specifically configured to: based on the sequence of the sequence numbers, sequentially play the corresponding voice markers Multiple audio files.
  • Fig. 28 shows a structural diagram of an apparatus for voice tagging a video according to an embodiment.
  • the device 2800 includes:
  • the display unit 2810 is configured to display a recording interface including a first video, the first video includes a first video frame; the determining unit 2820 is configured to respond to a selection instruction of the first video frame, The video frame is determined as the target video frame; the collection unit 2830 is configured to continuously collect voice signals in response to the recording start instruction issued based on the recording interface; the storage unit 2840 is configured to respond to the recording end instruction issued based on the recording interface , Storing the collected voice signal as an audio file; the adding unit 2850 is configured to add a voice mark associated with the audio file on the target video frame.
  • the collecting unit 2830 is specifically configured to continuously collect voice signals in response to a recording start instruction issued based on the first position of the target video frame in the recording interface; the adding unit 2850 is specifically configured to: The voice mark is added to the first position of the video frame.
  • the determining unit 2820 is specifically configured to: in response to a jump instruction issued for the first video, determine the first video frame displayed by the jump as the target video frame; or, in response to The pause instruction issued for the first video determines the first video frame that is paused to be displayed as the target video frame.
  • the device 2800 further includes: a progress bar marking unit configured to add a time point mark to the progress bar of the first video, the time point mark corresponding to the playback time of the target video frame point.
  • the device 2800 further includes: a first display unit configured to display the number of voice markers included in the target video frame in response to the input control indicator moving to the time point marker .
  • the storage unit 2840 is further configured to: perform voice recognition on the collected voice signal to obtain a recognized text; determine the subject text corresponding to the recognized text; the device further includes: a second display unit , Configured to display the topic text in response to the input control indicator moving to the time point mark.
  • Fig. 29 shows a structural diagram of a video viewing device according to an embodiment. As shown in FIG. 29, the device 2900 includes:
  • the display unit 2910 is configured to display a video with a voice mark, the voice mark is added to the video through the above-mentioned device 2800, and the video includes a first video frame with a first voice mark; a playback unit 2920, It is configured to play an audio file associated with the first voice mark in response to a trigger instruction to the first voice mark.
  • the first video frame includes several voice marks including the first voice mark
  • the progress bar of the video displays a time point mark at the playback time point of the first video frame
  • the device further includes: a display unit configured to display the number of marks of the plurality of voice marks and/or the plurality of topic texts corresponding to the plurality of voice marks in response to the movement of the input control indicator to the time point mark.
  • Fig. 30 shows a structural diagram of an apparatus for text labeling a picture according to an embodiment. As shown in FIG. 30, the device 3000 further includes:
  • the display unit 3010 is configured to display an editing interface for the target picture; the collection unit 3020 is configured to continuously collect voice signals in response to a recording start instruction issued based on the editing interface; the recognition unit 3030 is configured to respond to the editing based on the editing interface.
  • the recording end instruction issued by the interface performs voice recognition on the collected voice signal to obtain recognized text; the adding unit 3040 is configured to add a text mark associated with the recognized text on the target picture.
  • the collecting unit 3020 is specifically configured to continuously collect voice signals in response to the recording start instruction issued based on the first position of the target picture in the editing interface; the adding unit 3040 is specifically configured to: Add the text mark at the first position.
  • the adding unit 3040 is further configured to display a serial number in the text mark, the serial number being determined based on the number of previous text marks added to the target picture.
  • the adding unit 3040 is further configured to: fold and display the recognized text in the text mark; the device further includes an unfolding unit configured to respond to a trigger instruction to the text mark, The recognized text is expanded and displayed in the text mark.
  • the device 3000 further includes: a determining unit configured to determine the topic text corresponding to the recognized text; wherein the adding unit 3040 is specifically configured to display the topic text in the text mark.
  • the device 3000 further includes a display unit configured to display the recognized text in response to a trigger instruction to mark the text.
  • Fig. 31 shows a structural diagram of a picture viewing device according to an embodiment. As shown in FIG. 31, the device 3100 includes:
  • the display unit 3110 is configured to display a picture with a text mark, and the text mark is added to the picture by the above device 3000; the display unit 3120 is configured to display the text in response to a trigger instruction to the text mark The recognized text associated with the tag.
  • the display unit 3120 is specifically configured to: display the serial number in the text mark; or, display the subject text corresponding to the recognized text in the text mark.
  • the display unit 3120 is specifically configured to: fold and display the recognized text in the text mark; the display unit 3120 is specifically configured to: expand and display the recognized text in the text mark .
  • Fig. 32 shows a structural diagram of an apparatus for text tagging a video according to an embodiment.
  • the device 3200 includes:
  • the display unit 3210 is configured to display an editing interface including a first video, the first video includes a first video frame; the determining unit 3220 is configured to respond to a selection instruction of the first video frame, The video frame is determined as the target video frame; the acquisition unit 3230 is configured to continuously collect voice signals in response to the recording start instruction issued based on the editing interface; the recognition unit 3240 is configured to respond to the recording end instruction issued based on the editing interface , Perform voice recognition on the collected voice signal to obtain recognized text; the adding unit 3250 is configured to add a text mark associated with the recognized text on the target video frame.
  • the collecting unit 3230 is configured to continuously collect voice signals in response to the recording start instruction issued based on the first position of the target video frame in the editing interface; wherein the adding unit 3250 is specifically configured to: The text mark is added to the first position of the video frame.
  • Fig. 33 shows a structural diagram of a video viewing device according to an embodiment. As shown in FIG. 33, the device 3300 includes:
  • the display unit 3310 is configured to display a video with a text mark, the voice mark is added to the video through the above-mentioned device 3200, and the video includes a first video frame with a first text mark; the display unit 3320, It is configured to display the recognized text associated with the text mark in response to a trigger instruction to the first text mark.
  • FIG. 34 shows a flowchart of a picture processing method according to an embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, an instant messaging client.
  • the method includes the following steps:
  • Step S3410 display the chat interface, and receive the target picture to be sent selected based on the chat interface;
  • Step S3420 in response to the editing instruction for the target picture, enter the picture editing interface, where a voice mark icon is displayed;
  • Step S3430 In response to the trigger instruction to the voice mark icon, enter the recording interface;
  • step S3440 store the voice signal collected based on the recording interface as an audio file, and add a voice mark associated with the audio file on the target picture .
  • step S3440 may include: displaying the chat nickname and/or chat avatar of the currently editing user in the voice tag.
  • a voice mark 3502 with the current user's avatar is displayed.
  • the method may further include: in response to the sending instruction to the target picture, sending the target picture with the voice mark.
  • chat software can be implemented in the chat software to add voice tags to pictures and send target pictures with voice tags.
  • FIG. 36 shows a flowchart of a picture processing method according to another embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, an instant messaging client.
  • the method includes the following steps:
  • Step S3610 display a chat interface, the chat window of the chat interface contains a target picture; step S3620, in response to a trigger instruction for the target picture, display a menu bar, the menu bar includes a voice mark icon; step S3630, In response to the trigger instruction to the voice mark icon, enter the recording interface; step S3640, store the voice signal collected based on the recording interface as an audio file, and add a voice mark associated with the audio file on the target picture .
  • the chat window of the chat interface shown in FIG. 3 contains a target picture 32, and in response to a trigger instruction for the target picture, a menu bar 33 is displayed, and the menu bar 33 includes a voice mark
  • the icon 34 enters the recording interface in response to the trigger instruction to the voice mark icon. Further, a voice mark can be added to the target picture.
  • the method may further include: receiving the target picture from the contact corresponding to the chat window, the target picture includes an existing voice tag, and the contact is displayed Chat nickname and/or chat avatar.
  • step S3640 may include: displaying the chat nickname and/or chat avatar of the currently editing user in the voice tag.
  • FIG. 37 shows a schematic diagram of a chat interface according to another embodiment.
  • the chat window of the interface includes a picture 3701 received from the current contact, with a voice mark 3702 displaying the current contact's avatar.
  • the window also includes a picture 3703 sent after the current user edits the picture 3701. Compared with the picture 3701, there is also a voice mark 3704 on it, in which the current user's avatar is displayed.
  • the method may further include: in response to the exit instruction for the recording interface, updating and displaying the original target picture as the target with the voice mark in the chat window picture. Further, in a specific embodiment, a prompt message may also be displayed in the chat window to inform the contacts of all parties that the target picture has been modified.
  • FIG. 38 shows a schematic diagram of a chat interface according to another embodiment.
  • the picture 3801 in the chat window is updated and displayed as a picture 3803 with a voice mark 3802, and a prompt message 3804 is displayed, that is, "Youtiao" has been displayed. Add a voice tag to the picture "1.jpg".
  • FIG. 39 shows a flowchart of a picture processing method according to another embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, an instant messaging client.
  • the method includes the following steps:
  • Step S3910 display a chat interface
  • the chat window of the chat interface contains a target picture with a first voice tag
  • the first voice tag is added by the current contact corresponding to the chat window
  • step S3920 in response to the The trigger instruction of the first voice mark displays a menu bar, and the menu bar includes a voice reply icon
  • step S3930 in response to the trigger instruction to the voice reply icon, enter the recording interface
  • step S3940 based on the The voice signal collected by the recording interface is stored as an audio file, and a second voice mark associated with the audio file is added in the area adjacent to the first voice mark in the target picture; or, based on the recording interface collection
  • the voice signal of is added to the audio file corresponding to the first voice mark.
  • FIG. 40 shows a schematic diagram of a chat interface according to another embodiment.
  • the chat window contains a picture 4002 with a voice mark 4001.
  • a menu is displayed
  • the column which includes the voice reply icon 4003, enters the recording interface in response to the triggering instruction to the voice reply icon 4003, and further, for the user's input voice, a voice mark 4004 can be added near the voice mark 4001.
  • the user's input voice can be added to the audio file corresponding to the voice mark 4001, and the voice mark 4001 displaying serial number 1 can be updated and displayed as 4101 displaying serial number 2.
  • the meanings of serial numbers 1 and 2 are: The number of different users edited, or the total number of edits.
  • the method may further include: in response to the exit instruction for the recording interface, displaying a prompt message in the chat window to inform all contacts of the target picture
  • the voice tag in has been modified.
  • FIG. 41 shows a schematic diagram of a chat interface according to still another embodiment. As shown in FIG. 41, a prompt message 4102 is displayed in the chat window of the chat interface. Voice tag content, click to listen.
  • FIG. 42 shows a flowchart of a picture processing method according to still another embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, picture editing software.
  • the method includes the following steps:
  • Step S4210 display the picture editing interface containing the target picture, the function menu of this interface includes a voice mark icon; Step S4220, in response to the trigger instruction of the voice mark icon, enter the voice mark interface; Step S4230, based on the The input text received by the voice mark interface is converted into an audio file, and a voice mark associated with the audio file is added to the target picture.
  • FIG. 43 shows a flowchart of a picture processing method according to an embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, picture editing software.
  • the method includes the following steps:
  • Step S4310 display the chat interface containing the target picture; step S4320, in response to the trigger instruction for the target picture, display the menu bar, which includes the voice adding emoticon icon; step S4330, in response to the triggering of the voice adding emoticon icon Instruction, enter the voice adding emoticon interface; step S4340, converting the voice signal collected based on the voice adding emoticon interface into text, and generating an animated emoticon based on the text; step S4350, adding the animated emoticon to the target picture .
  • step S4340 may include: retrieving an original expression related to the text from an expression database; adding the text to the original expression to obtain the animated expression.
  • the picture a certain picture
  • the picture includes an animated expression 4401 added by the user through voice input (for example, the input content is hahaha).
  • users can generate animated expressions through voice input, and edit or reply to target pictures, thereby improving the fun of chatting between users and the convenience of communication.
  • FIG. 45 shows a flowchart of a picture processing method according to another embodiment.
  • the execution subject of the method may be any device, device, platform, or device cluster with computing and processing capabilities, for example, an e-commerce platform.
  • the method includes the following steps:
  • Step S4510 display the product information editing interface, which includes the target image for the target product; step S4520, in response to the trigger instruction to the target image, display a menu bar, the menu bar includes a voice mark icon; step S4530, response After triggering the voice mark icon, enter the recording interface; step S4540, store the voice signal collected based on the recording interface as an audio file, and add a voice mark associated with the audio file on the target picture.
  • FIG. 46 shows a schematic diagram of a product information editing interface according to an embodiment.
  • the product information editing interface 4601 includes a product image 4602.
  • the product image 4602 In response to a trigger instruction to the product image 4602, it may be displayed including voice
  • the menu bar of the mark icon 4603 can further be triggered by triggering the voice mark icon 4603 to enter the recording interface, and edit the voice mark for the product image 4602, including adding, deleting and modifying.
  • the method may further include: displaying the target picture in a product detail page for the target product.
  • the product detail page shown in FIG. 12 includes a target picture 121 with voice tags.
  • the target product can be displayed to the user more intuitively through the product introduction picture with voice mark in the product detail page.
  • Fig. 47 shows a flowchart of a picture processing method according to another embodiment.
  • the execution subject of the method can be any device, device, platform, or device cluster with computing and processing capabilities, for example, an e-commerce platform.
  • the method includes the following steps:
  • Step S4710 displaying the order evaluation interface for the first order, which includes adding a picture icon
  • Step S4720 receiving the selected target picture in response to the trigger instruction for the adding picture icon
  • Step S4730 responding to the target picture
  • the issued voice mark instruction enters the recording interface
  • step S4740 the voice signal collected based on the recording interface is stored as an audio file, and the voice mark associated with the audio file is added to the target electronic file.
  • Fig. 48 shows a flowchart of a picture processing method according to another embodiment.
  • the execution subject of the method can be any device, device, platform, or device cluster with computing and processing capabilities, for example, an e-commerce platform.
  • the method includes the following steps:
  • Step S4810 display the product evaluation interface for the target product, which includes the first user evaluation, and the first user evaluation includes the target picture with the first voice mark;
  • step S4820 in response to the trigger instruction for the target picture, Display a menu bar, the menu bar includes a voice reply icon;
  • step S4830 in response to the trigger instruction to the voice reply icon, enter the recording interface;
  • step S4840 store the voice signal collected based on the recording interface as an audio file , And in the area adjacent to the first voice mark in the target picture, add a second voice mark associated with the audio file; or, add a voice signal collected based on the recording interface to the first In the audio file corresponding to the voice tag.
  • the method may further include: in response to the exit instruction for the recording interface, displaying information about the first user’s evaluation on the product evaluation interface.
  • the notification message is used to notify that the first user evaluation has been replied.
  • the above can realize quick and targeted communication between the seller and the buyer, or between the buyer and the buyer based on the product picture in the evaluation scenario.
  • FIG. 49 shows a flowchart of a picture processing method according to still another embodiment, and the execution subject of the method may be a customer service platform. As shown in Figure 49, the method includes the following steps:
  • Step S4910 Receive a conversation message sent by a user, the conversation message includes a target picture with a first voice tag;
  • Step S4920 Obtain an audio file associated with the first voice tag, and perform voice recognition on the audio file , Obtain the recognized text;
  • step S4930 input the recognized text into the pre-trained user questioning prediction model, and output the corresponding user standard question;
  • step S4940 feed back the answer to the question corresponding to the user standard question to the user .
  • step S4940 may include: converting the answer to the question into answer audio, and adding a second voice mark associated with the answer audio to the target picture; adding the new The target picture of the second voice mark is sent to the user.
  • step S4940 may include: converting the answer to the question into answer audio; adding a second voice mark associated with the answer audio in an area adjacent to the first voice mark in the target picture Or, adding the reply audio to the audio file corresponding to the first voice tag; sending prompt information to the user to remind the user that the reply content has been added to the target picture.
  • the above can realize convenient communication between the user and the customer service in the customer service scenario.
  • FIG. 50 shows a flowchart of an electronic file processing method according to still another embodiment.
  • the execution subject of the method can be any device, equipment, platform, equipment cluster with computing and processing capabilities, for example, office software or auditing platform, etc. .
  • the method includes the following steps:
  • Step S5010 displaying a file processing interface for the target electronic file, the function menu bar of the file processing interface includes a voice mark icon;
  • Step S5020 in response to a trigger instruction to the voice mark icon, enter the recording interface;
  • Step S5030 The voice signal collected based on the recording interface is stored as an audio file, and a voice mark associated with the audio file is added to the target electronic file.
  • the target electronic document is an electronic contract
  • the document processing interface is a contract approval interface
  • the file format of the target electronic file is a word document, or a PDF document, or an excel form.
  • FIG. 51 shows a schematic diagram of an office software interface according to an embodiment.
  • the menu bar of the interface includes a voice mark icon 5101, in which a voice mark 5102 is displayed in a word document.
  • FIG. 52 shows a flowchart of a picture processing method according to an embodiment, and the execution subject of the method is a live broadcast platform. As shown in Figure 52, the method includes:
  • Step S5210 display the product information editing interface, which includes the target picture for the target product to be put on the shelf; step S5220, in response to the trigger instruction to the target picture, display a menu bar, the menu bar including a voice mark icon; step S5230, in response to the triggering instruction to the voice mark icon, enter the recording interface; step S5240, store the voice signal collected based on the recording interface as an audio file, and add a link associated with the audio file to the target picture Voice tag.
  • the method further includes: displaying a live broadcast interface, the live broadcast interface includes a commodity listing icon; in response to a trigger instruction to the commodity listing diagram, displaying a picture of the commodity to be put on the shelf, The target picture is included; in response to a selection instruction of the target picture, the target picture is displayed in the commodity display window of the live broadcast interface.
  • FIG. 53 shows a schematic diagram of a live broadcast interface according to an embodiment. In the commodity display window of the live broadcast interface, a target picture 5302 with a voice mark 5301 is displayed.
  • Fig. 54 shows a structural diagram of a picture processing apparatus according to an embodiment. As shown in FIG. 54, the device 5400 includes:
  • the display unit 5410 is configured to display a chat interface; the receiving unit 5420 is configured to receive a target picture to be sent selected based on the chat interface; the first interface switching unit 5430 is configured to respond to an editing instruction for the target picture, Enter the picture editing interface, where the voice mark icon is displayed; the second interface switching unit 5440 is configured to enter the recording interface in response to a trigger instruction to the voice mark icon; the marking unit 5450 is configured to collect data collected based on the recording interface
  • the voice signal is stored as an audio file, and a voice mark associated with the audio file is added to the target picture.
  • Fig. 55 shows a structural diagram of a picture processing apparatus according to another embodiment. As shown in FIG. 55, the device 5500 includes:
  • the interface display unit 5510 is configured to display a chat interface, and the chat window of the chat interface contains a target picture; the menu bar display unit 5520 is configured to display a menu bar in response to a trigger instruction for the target picture, the menu bar
  • the interface switching unit 5530 is configured to enter the recording interface in response to a trigger instruction to the voice mark icon; the marking unit 5540 is configured to store the voice signal collected based on the recording interface as an audio file, And add a voice mark associated with the audio file on the target picture.
  • Fig. 56 shows a structural diagram of a picture processing apparatus according to still another embodiment. As shown in FIG. 56, the device 5600 includes:
  • the interface display unit 5610 is configured to display a chat interface, the chat window of the chat interface contains a target picture with a first voice mark, and the first voice mark is added by the current contact corresponding to the chat window; menu bar
  • the display unit 5620 is configured to display a menu bar in response to a trigger instruction to the first voice mark, and the menu bar includes a voice reply icon;
  • the interface switching unit 5630 is configured to respond to triggering of the voice reply icon Instruction to enter the recording interface;
  • the marking unit 5640 is configured to store the voice signal collected based on the recording interface as an audio file, and add the associated audio in the area adjacent to the first voice mark in the target picture The second voice mark of the file; or, adding the voice signal collected based on the recording interface to the audio file corresponding to the first voice mark.
  • Fig. 57 shows a structural diagram of a picture processing apparatus according to still another embodiment.
  • the device 5700 includes:
  • the display unit 5710 is configured to display a picture editing interface containing a target picture, and the function menu of the picture editing interface includes a voice mark icon; the interface switching unit 5720 is configured to enter the voice in response to a trigger instruction to the voice mark icon Marking interface; marking unit 5730, configured to convert input text received based on the voice marking interface into an audio file, and add a voice mark associated with the audio file on the target picture.
  • Fig. 58 shows a structural diagram of a picture processing apparatus according to still another embodiment.
  • the device 5800 includes:
  • the interface display unit 5810 is configured to display a chat interface containing a target picture;
  • the menu bar display unit 5820 is configured to display a menu bar in response to a trigger instruction to the target picture, the menu bar including a voice adding emoticon icon;
  • interface The switching unit 5830 is configured to respond to a trigger instruction to add an emoticon icon to the voice to enter the voice adding emoticon interface;
  • the emoticon generating unit 5840 is configured to convert the voice signal collected based on the voice adding emoticon interface into text based on The text generates an animated expression;
  • the expression adding unit 5850 is configured to add the animated expression to the target picture.
  • Fig. 59 shows a structural diagram of a picture processing apparatus according to an embodiment.
  • the apparatus is integrated in an e-commerce platform, and the apparatus 5900 includes:
  • the interface display unit 5910 is configured to display a product information editing interface, which includes a target picture for the target product; the menu bar display unit 5920 is configured to display a menu bar in response to a trigger instruction to the target picture, the menu bar Including a voice mark icon; an interface switching unit 5930, configured to enter the recording interface in response to a trigger instruction to the voice mark icon; a marking unit 5940, configured to store the voice signal collected based on the recording interface as an audio file, and Adding a voice mark associated with the audio file on the target picture.
  • Fig. 60 shows a structural diagram of a picture processing apparatus according to still another embodiment. As shown in FIG. 60, the device 6000 includes:
  • the display unit 6010 is configured to display the order evaluation interface for the first order, which includes an add picture icon; the receiving unit 6020 is configured to receive the selected target picture in response to the trigger instruction for the add picture icon; the interface switching unit 6030 , Configured to enter the recording interface in response to the voice marking instruction issued to the target picture; marking unit 6040, configured to store the voice signal collected based on the recording interface as an audio file, and add it to the target electronic file Associate the voice tag of the audio file.
  • Fig. 61 shows a structural diagram of a picture processing apparatus according to another embodiment. As shown in FIG. 61, the device 6100 includes:
  • the interface display unit 6110 is configured to display a product evaluation interface for the target product, which includes a first user evaluation, and the first user evaluation includes a target picture with a first voice mark;
  • the menu bar display unit 6120 is configured to respond to A menu bar is displayed for the trigger instruction of the target picture, and the menu bar includes a voice reply icon;
  • the interface switching unit 6130 is configured to enter the recording interface in response to the trigger instruction of the voice reply icon;
  • the marking unit 6140 It is configured to store the voice signal collected based on the recording interface as an audio file, and add a second voice mark associated with the audio file in an area adjacent to the first voice mark in the target picture; or The voice signal collected based on the recording interface is added to the audio file corresponding to the first voice mark.
  • Fig. 62 shows a structural diagram of a picture processing apparatus according to still another embodiment.
  • the apparatus is integrated in a customer service platform, and the apparatus 6200 includes:
  • the receiving unit 6210 is configured to receive a conversation message sent by a user, and the conversation message includes a target picture with a first voice mark; the obtaining unit 6220 is configured to obtain an audio file associated with the first voice mark, and to The audio file performs voice recognition to obtain the recognized text; the prediction unit 6230 is configured to input the recognized text into the pre-trained user questioning prediction model, and output corresponding user standard questions; the feedback unit 6240 is configured to communicate with the The answer to the question corresponding to the user's standard question is fed back to the user.
  • Fig. 63 shows a structural diagram of an electronic file processing apparatus according to an embodiment.
  • the apparatus 6300 includes:
  • the display unit 6310 is configured to display a file processing interface for the target electronic file, and the function menu bar of the file processing interface includes a voice mark icon; the interface switching unit 6320 is configured to respond to a trigger instruction to the voice mark icon, Enter the recording interface; the marking unit 6330 is configured to store the voice signal collected based on the recording interface as an audio file, and add a voice mark associated with the audio file on the target electronic file.
  • Fig. 64 shows a structural diagram of a picture processing apparatus according to still another embodiment.
  • the apparatus is integrated in a live broadcast platform, and the apparatus 6400 includes:
  • the interface display unit 6410 is configured to display a product information editing interface, which includes a target picture for the target product to be put on the shelf;
  • the menu bar display unit 6420 is configured to display a menu bar in response to a trigger instruction to the target picture, the The menu bar includes a voice mark icon;
  • the interface switching unit 6430 is configured to enter the recording interface in response to a trigger instruction to the voice mark icon;
  • the marking unit 6440 is configured to store the voice signal collected based on the recording interface as audio File, and add a voice mark associated with the audio file on the target picture.
  • a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed in a computer, the computer is executed in conjunction with FIG. 2 or FIG. 11 or FIG. 13 or FIG. 16 or 18 or 20 or 22 or 24 or 34 or 36 or 39 or 42 or 45 or 47 or 48 or 49 or 50 or 52.
  • a computing device including a memory and a processor, the memory stores executable code, and when the processor executes the executable code, a combination of FIG. 2 or FIG. 11 or FIG. 13 or 16 or 18 or 20 or 22 or 24 or 34 or 36 or 39 or 42 or 45 or 47 or 48 or 49 or 50 or 52 .
  • the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof.
  • these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on the computer-readable medium.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Library & Information Science (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Health & Medical Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Television Signal Processing For Recording (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

一种对图片、视频进行语音标记的方法及装置。其中,对图片进行语音标记的方法包括:首先,显示包括目标图片的录音界面(S210);接着,响应于基于所述录音界面发出的录音开始指令,持续采集语音信号(S220);然后,响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件(S230);再接着,在所述目标图片上添加关联所述音频文件的语音标记(S240)。

Description

对图片、视频进行语音标记的方法及装置
本申请要求2020年03月11日递交的申请号为202010167913.8、发明名称为“对图片、视频进行语音标记的方法及装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本说明书一个或多个实施例涉及计算机技术领域,尤其涉及一种对图片进行语音标记的方法及装置、一种对视频进行语音标记的方法及装置、一种对图片进行文本标记的方法及装置、一种对视频进行文本标记的方法及装置、一种图片查看方法及装置、一种视频查看方法及装置。
背景技术
在许多场景下,需要对图片或视频中内容进行说明。比如说,为介绍某个网站的使用方式,需要对网页截图中的操作图标进行说明。又比如,为介绍某款产品,需结合对该产品的拍摄图片对产品的各组成部分进行说明。
然而,目前可供用户选择的对图片或视频中特定内容进行说明的方式,例如,用户使用图片编辑软件,通过添加文本框的方式补入说明性文字,比较单一,且操作成本较高。
因此,需要一种方案,可以提高需要基于图片或视频中的内容进行说明时的便捷度,从而提高用户体验。
发明内容
本说明书一个或多个实施例描述了一种对图片进行语音标记的方法,通过在图片上进行任何位置的录音,即语音标记功能,可以实现方便、快捷地对图片进行解释、说明。
根据第一方面,提供一种对图片进行语音标记的方法,该方法包括:显示包括目标图片的录音界面;响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;在所述目标图片上添加关联所述音频文件的语音标记。
根据第二方面,提供一种图片查看方法,该方法包括:显示带有语音标记的图片,所述语音标记通过第一方面提供的方法添加至所述图片中;响应于对所述语音标记的触 发指令,播放所述语音标记关联的音频文件。
根据第三方面,提供一种对视频进行语音标记的方法,该方法包括:显示包括第一视频的录音界面,所述第一视频包括第一视频帧;响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;在所述目标视频帧上添加关联所述音频文件的语音标记。
根据第四方面,提供一种视频查看方法,该方法包括:显示带有语音标记的视频,所述语音标记通过第三方面提供的方法添加至所述视频中,所述视频中包括带有第一语音标记的第一视频帧;响应于对所述第一语音标记的触发指令,播放所述第一语音标记关联的音频文件。
根据第五方面,提供一种对图片进行文本标记的方法,该方法包括:显示针对目标图片的编辑界面;响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;在所述目标图片上添加关联所述识别文本的文本标记。
根据第六方面,提供一种图片查看方法,该方法包括:显示带有文本标记的图片,所述文本标记通过第五方面提供的方法添加至所述图片中;响应于对所述文本标记的触发指令,展示所述文本标记关联的识别文本。
根据第七方面,提供一种对视频进行文本标记的方法,该方法包括:显示包括第一视频的编辑界面,所述第一视频包括第一视频帧;响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;在所述目标视频帧上添加关联所述识别文本的文本标记。
根据第八方面,提供一种视频查看方法,该方法包括:显示带有文本标记的视频,所述语音标记通过第七方面提供的方法添加至所述视频中,所述视频中包括带有第一文本标记的第一视频帧;响应于对所述第一文本标记的触发指令,展示所述文本标记关联的识别文本。
根据第九方面,提供一种对图片进行语音标记的装置,该装置包括:显示单元,配置为显示包括目标图片的录音界面;采集单元,配置为响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;存储单元,配置为响应于基于所述录音界面发出的 录音结束指令,将采集到的语音信号存储为音频文件;添加单元,配置为在所述目标图片上添加关联所述音频文件的语音标记。
根据第十方面,提供一种图片查看装置,该装置包括:显示单元,配置为显示带有语音标记的图片,所述语音标记通过第九方面提供的装置添加至所述图片中;播放单元,配置为响应于对所述语音标记的触发指令,播放所述语音标记关联的音频文件。
根据第十一方面,提供一种对视频进行语音标记的装置,该装置包括:显示单元,配置为显示包括第一视频的录音界面,所述第一视频包括第一视频帧;确定单元,配置为响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;采集单元,配置为响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;存储单元,配置为响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;添加单元,配置为在所述目标视频帧上添加关联所述音频文件的语音标记。
根据第十二方面,提供一种视频查看装置,该装置包括:显示单元,配置为显示带有语音标记的视频,所述语音标记通过第十一方面提供的装置添加至所述视频中,所述视频中包括带有第一语音标记的第一视频帧;播放单元,配置为响应于对所述第一语音标记的触发指令,播放所述第一语音标记关联的音频文件。
根据第十三方面,提供一种对图片进行文本标记的装置,该装置包括:显示单元,配置为显示针对目标图片的编辑界面;采集单元,配置为响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;识别单元,配置为响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;添加单元,配置为在所述目标图片上添加关联所述识别文本的文本标记。
根据第十四方面,提供一种图片查看装置,该装置包括:显示单元,配置为显示带有文本标记的图片,所述文本标记通过第十三方面提供的装置添加至所述图片中;展示单元,配置为响应于对所述文本标记的触发指令,展示所述文本标记关联的识别文本。
根据第十五方面,提供一种对视频进行文本标记的装置,该装置包括:显示单元,配置为显示包括第一视频的编辑界面,所述第一视频包括第一视频帧;确定单元,配置为响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;采集单元,配置为响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;识别单元,配置为响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;添加单元,配置为在所述目标视频帧上添加关联所述识别文本的文本标记。
根据第十六方面,提供一种视频查看装置,该装置包括:显示单元,配置为显示带有文本标记的视频,所述语音标记通过第十五方面提供的装置添加至所述视频中,所述视频中包括带有第一文本标记的第一视频帧;展示单元,配置为响应于对所述第一文本标记的触发指令,展示所述文本标记关联的识别文本。
根据第十七方面,提供一种图片处理方法,包括:显示聊天界面,并接收基于所述聊天界面选取的待发送的目标图片;响应于针对所述目标图片的编辑指令,进入图片编辑界面,其中显示语音标记图标;响应于对所述语音标记图标的触发指令,进入录音界面;将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第十八方面,提供一种图片处理方法,包括:显示聊天界面,所述聊天界面的聊天窗口中包含目标图片;响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;响应于对所述语音标记图标的触发指令,进入录音界面;将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第十九方面,提供一种图片处理方法,包括:显示聊天界面,所述聊天界面的聊天窗口中包含带有第一语音标记的目标图片,所述第一语音标记由所述聊天窗口对应的当前联系人添加;响应于对所述第一语音标记的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;响应于对所述语音回复图标的触发指令,进入录音界面;将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
根据第二十方面,提供一种图片处理方法,包括:显示包含目标图片的图片编辑界面,所述图片编辑界面的功能菜单中包括语音标记图标;响应于对所述语音标记图标的触发指令,进入语音标记界面;将基于所述语音标记界面接收的输入文本转化为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第二十一方面,提供一种图片处理方法,包括:显示包含目标图片的聊天界面;响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音添加表情图标;响应于对所述语音添加表情图标的触发指令,进入语音添加表情界面;将基于所述语音添加表情界面采集的语音信号转化为文字,并基于所述文字生成动画表情;在所述目标图片中添加所述动画表情。
根据第二十二方面,提供一种图片处理方法,所述方法的执行主体为电商平台,所述方法包括:显示商品信息编辑界面,其中包括针对目标商品的目标图片;响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;响应于对所述语音标记图标的触发指令,进入录音界面;将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第二十三方面,提供一种图片处理方法,包括:显示针对第一订单的订单评价界面,其中包括添加图片图标;响应于针对所述添加图片图标的触发指令,接收选取的目标图片;响应于对所述目标图片发出的语音标记指令,进入录音界面;将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
根据第二十四方面,提供一种图片处理方法,包括:显示针对目标商品的商品评价界面,其中包括第一用户评价,所述第一用户评价中包括带第一语音标记的目标图片;响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;响应于对所述语音回复图标的触发指令,进入录音界面;将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
根据第二十五方面,提供一种图片处理方法,所述方法的执行主体为客服平台,所述方法包括:接收用户发送的会话消息,所述会话消息中包括带有第一语音标记的目标图片;获取与所述第一语音标记关联的音频文件,对所述音频文件进行语音识别,得到识别文本;将所述识别文本输入预先训练的用户标问预测模型中,输出对应的用户标准问题;将与所述用户标准问题对应的问题答案反馈给所述用户。
根据第二十六方面,提供一种电子文件的处理方法,包括:显示针对目标电子文件的文件处理界面,所述文件处理界面的功能菜单栏中包括语音标记图标;响应于对所述语音标记图标的触发指令,进入录音界面;将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
根据第二十七方面,提供一种图片处理方法,所述方法的执行主体为直播平台,所述方法包括:显示商品信息编辑界面,其中包括针对待上架的目标商品的目标图片;响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;响应于对所述语音标记图标的触发指令,进入录音界面;将基于所述录音界面采集的语音信 号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第二十八方面,提供一种图片处理装置,包括:显示单元,配置为显示聊天界面;接收单元,配置为收基于所述聊天界面选取的待发送的目标图片;第一界面切换单元,配置为响应于针对所述目标图片的编辑指令,进入图片编辑界面,其中显示语音标记图标;第二界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第二十九方面,提供一种图片处理装置,包括:界面显示单元,配置为显示聊天界面,所述聊天界面的聊天窗口中包含目标图片;菜单栏显示单元,配置为响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第三十方面,提供一种图片处理装置,包括:界面显示单元,配置为显示聊天界面,所述聊天界面的聊天窗口中包含带有第一语音标记的目标图片,所述第一语音标记由所述聊天窗口对应的当前联系人添加;菜单栏显示单元,配置为响应于对所述第一语音标记的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;界面切换单元,配置为响应于对所述语音回复图标的触发指令,进入录音界面;标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
根据第三十一方面,提供一种图片处理装置,所述装置包括:显示单元,配置为显示包含目标图片的图片编辑界面,所述图片编辑界面的功能菜单中包括语音标记图标;界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入语音标记界面;标记单元,配置为将基于所述语音标记界面接收的输入文本转化为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第三十二方面,提供一种图片处理装置,包括:界面显示单元,配置为显示包含目标图片的聊天界面;菜单栏显示单元,配置为响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音添加表情图标;界面切换单元,配置为响应于对所述语音添加表情图标的触发指令,进入语音添加表情界面;表情生成单元,配置为将基 于所述语音添加表情界面采集的语音信号转化为文字,并基于所述文字生成动画表情;表情添加单元,配置为在所述目标图片中添加所述动画表情。
根据第三十三方面,提供一种图片处理装置,所述装置集成于电商平台,所述装置包括:界面显示单元,配置为显示商品信息编辑界面,其中包括针对目标商品的目标图片;菜单栏显示单元,配置为响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第三十四方面,提供一种图片处理装置,包括:显示单元,配置为显示针对第一订单的订单评价界面,其中包括添加图片图标;接收单元,配置为响应于针对所述添加图片图标的触发指令,接收选取的目标图片;界面切换单元,配置为响应于对所述目标图片发出的语音标记指令,进入录音界面;标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
根据第三十五方面,提供一种图片处理装置,包括:界面显示单元,配置为显示针对目标商品的商品评价界面,其中包括第一用户评价,所述第一用户评价中包括带第一语音标记的目标图片;菜单栏显示单元,配置为响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;界面切换单元,配置为响应于对所述语音回复图标的触发指令,进入录音界面;标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
根据第三十六方面,提供一种图片处理装置,所述装置集成于客服平台,所述装置包括:接收单元,配置为接收用户发送的会话消息,所述会话消息中包括带有第一语音标记的目标图片;获取单元,配置为获取与所述第一语音标记关联的音频文件,对所述音频文件进行语音识别,得到识别文本;预测单元,配置为将所述识别文本输入预先训练的用户标问预测模型中,输出对应的用户标准问题;反馈单元,配置为将与所述用户标准问题对应的问题答案反馈给所述用户。
根据第三十七方面,提供一种电子文件的处理装置,包括:显示单元,配置为显示针对目标电子文件的文件处理界面,所述文件处理界面的功能菜单栏中包括语音标记图 标;界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
根据第三十八方面,提供一种图片处理装置,所述装置集成于直播平台,所述装置包括:界面显示单元,配置为显示商品信息编辑界面,其中包括针对待上架的目标商品的目标图片;菜单栏显示单元,配置为响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据第三十九方面,提供了一种计算机可读存储介质,其上存储有计算机程序,当所述计算机程序在计算机中执行时,令计算机执行上述第一方面至第八方面、第十七方面至第二十七方面中任一方面的方法。
根据第四十方面,提供了一种计算设备,包括存储器和处理器,所述存储器中存储有可执行代码,所述处理器执行所述可执行代码时,实现第一方面至第八方面、第十七方面至第二十七方面中任一方面的方法。
综上,在本说明书实施例披露的上述对图片或视频进行语音标记的方法及装置中,通过在图片或视频上进行任何位置的录音,即语音标记功能,可以实现方便、快捷地对图片或视频进行解释、说明。
附图说明
为了更清楚地说明本发明实施例的技术方案,下面将对实施例描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本发明的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其它的附图。
图1示出根据一个实施例的语音标记功能的应用场景示意图;
图2示出根据一个实施例的对图片进行语音标记的方法流程图;
图3示出根据一个实施例的切换至录音界面的切换示意图;
图4示出根据另一个实施例的切换至录音界面的切换示意图;
图5示出根据一个实施例的使用语音标记功能时的界面切换示意图;
图6示出根据另一个实施例的使用语音标记功能时的界面切换示意图;
图7示出根据一个实施例的语音标记图标中序号修改的界面切换示意图;
图8示出根据一个实施例的语音标记图标中包含主题文本的界面示意图;
图9示出根据一个实施例的播放音频文件时的界面切换示意图;
图10示出根据一个实施例的语音识别相关的界面切换示意图;
图11示出根据一个实施例的图片查看方法流程图;
图12示出根据一个实施例的包含带语音图标的商品图片的商品详情界面;
图13示出根据一个实施例的对视频进行语音标记的方法流程图;
图14示出根据一个实施例的针对视频的语音标记界面示意图;
图15示出根据另一个实施例的针对视频的语音标记界面示意图;
图16示出根据一个实施例的视频查看方法流程图;
图17示出根据一个实施例的查看带语音标记的视频的界面切换示意图;
图18示出根据一个实施例的对图片进行文本标记的方法流程图;
图19示出根据一个实施例的对图片进行文本标记时的界面切换示意图;
图20示出根据一个实施例的图片查看方法流程图;
图21示出根据一个实施例的查看图片中文本标记的界面切换示意图;
图22示出根据一个实施例的对视频进行文本标记的方法流程图;
图23示出根据一个实施例的对视频进行文本标记时的界面切换示意图;
图24示出根据一个实施例的视频查看方法流程图;
图25示出根据一个实施例的查看视频中文本标记的界面切换示意图;
图26示出根据一个实施例的对图片进行语音标记的装置结构图;
图27示出根据一个实施例的图片查看装置结构图;
图28示出根据一个实施例的对视频进行语音标记的装置结构图;
图29示出根据一个实施例的视频查看装置结构图;
图30示出根据一个实施例的对图片进行文本标记的装置结构图;
图31示出根据一个实施例的图片查看装置结构图;
图32示出根据一个实施例的对视频进行文本标记的装置结构图;
图33示出根据一个实施例的视频查看装置结构图;
图34示出根据一个实施例的图片处理方法流程图;
图35示出根据一个实施例的聊天界面示意图;
图36示出根据另一个实施例的图片处理方法流程图;
图37示出根据另一个实施例的聊天界面示意图;
图38示出根据又一个实施例的聊天界面示意图;
图39示出根据又一个实施例的图片处理方法流程图;
图40示出根据还一个实施例的聊天界面示意图;
图41示出根据再一个实施例的聊天界面示意图;
图42示出根据再一个实施例的图片处理方法流程图;
图43示出根据一个实施例的图片处理方法流程图;
图44示出根据一个实施例的聊天界面示意图;
图45示出根据另一个实施例的图片处理方法流程图;
图46示出根据一个实施例的商品信息编辑界面示意图;
图47示出根据又一个实施例的图片处理方法流程图;
图48示出根据还一个实施例的图片处理方法流程图;
图49示出根据再一个实施例的图片处理方法流程图;
图50示出根据一个实施例的电子文件的处理方法流程图;
图51示出根据一个实施例的办公软件界面示意图;
图52示出根据一个实施例的图片处理方法流程图;
图53示出根据一个实施例的直播界面示意图;
图54示出根据一个实施例的图片处理装置结构图;
图55示出根据另一个实施例的图片处理装置结构图;
图56示出根据又一个实施例的图片处理装置结构图;
图57示出根据再一个实施例的图片处理装置结构图;
图58示出根据还一个实施例的图片处理装置结构图;
图59示出根据一个实施例的图片处理装置结构图;
图60示出根据又一个实施例的图片处理装置结构图;
图61示出根据另一个实施例的图片处理装置结构图;
图62示出根据再一个实施例的图片处理装置结构图;
图63示出根据一个实施例的电子文件的处理装置结构图;
图64示出根据还一个实施例的图片处理装置结构图。
具体实施方式
下面结合附图,对本说明书提供的方案进行描述。
如前所述,在许多场景下,需要对图片或视频中的特定内容进行说明。举例来说,若需要基于图片做说明(如教父母操作淘宝某个页面),目前的聊天工具只能用图片加文字的形式说明,图片上编辑文字的操作成本很高,而且大段的文字容易对图片中的原始内容进行遮挡。又举例来说,在商品详情页,若需要对图片中的商品进行说明,通常是先放图片,然后在图片下方用文字进行说明,此种图文结合的方式,难免不够直观,需要用户自主对齐图片内容和文字内容。
基于此,发明人提出设计一种语音标记功能,通过将语音录音功能与图片标记功能相结合,实现在图片、视频等电子媒介上,带指向性地说明内容。在一个实施例中,图1示出根据一个实施例的语音标记功能的应用场景示意图,如图1所示,首先,用户A在使用客户端A(如即时通讯客户端)的过程中,可以通过语音标记功能在图片的若干位置添加语音标记,例如,可以通过在图片中进行长按操作(如图1中的手指进行按压)触发录音,再通过取消按压操作结束录音,并在长按操作所对应的位置添加语音标记,如分别添加语音标记10和语音标记11;接着,响应于对发送图标12的点击指令,将带语音标记的图片发送至客户端B;然后,用户B在通过客户端B查看该图片的过程中,可以通过点击其上的语音标记,听取对应的录音文件,例如,通过点击语音标记10,可以听取对应的录音,同时语音标记10的显示状态切换为语音标记13,用于提示用户当前正在播放对应的录音文件。
如此,对于为图片添加语音标记的用户而言,其可以方便、快捷地对图片中任意位置的内容做出带指向性的说明;对于查看带语音标记的图片的用户而言,可以直观、便捷地获取其针对该图片需要了解的内容。因此,发明人提出的语音标记功能,可以充分提高各方用户的体验。
下面结合具体的实施例,对上述语音标记功能的实施方式进行说明。
具体地,图2示出根据一个实施例的对图片进行语音标记的方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,可以为客户端(如图片处理软件)、系统软件或客户端插件(如即时通讯客户端中的插件)。如图2所示,所述方法包括以下步骤:
步骤S210,显示包括目标图片的录音界面;步骤S220,响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;步骤S230,响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;步骤S240,在所述目标图片上添加 关联所述音频文件的语音标记。
以上步骤具体如下:
首先,在步骤S210,显示包括目标图片的录音界面。
在一个实施例中,本步骤可以包括:响应于基于录音界面发出的针对所述目标图片的导入指令,在录音界面中显示所述目标图片。在一个具体的实施例中,其中录音界面可以为图片处理软件中的编辑界面,或电商平台中的商品信息编辑界面,又或即时通讯客户端中针对待发送图片的编辑界面。在一个具体的实施例中,其中导入指令可以为声控指令或点击指令,如对导入图标的点击指令。在一个具体的例子中,响应于用户对图片处理APP的打开指令,进入上述录音界面,响应于对录音界面中图片导入图标的点击指令,显示图库中的图片,再响应于用户对其中某张图片的选取指令和确认导入指令,将该某张图片显示在上述录音界面。
在一个实施例中,在本步骤之前,所述方法还可以包括:响应于针对目标图片的触发指令,显示语音标记图标;相应地,本步骤可以包括:响应于对所述语音标记图标的触发指令,跳转至所述录音界面。
在一个具体的实施例中,其中针对目标图片的触发指令可以包括查看指令、或编辑指令、或发送指令。在另一个具体的实施例中,其中针对目标图片的触发指令可以包括用于截取所述目标图片的截屏指令,也就是说,触发指令是截屏指令,截屏得到的图片为所述目标图片。根据一个具体的例子,图3中示出聊天界面31,响应于对其中图片32的进一步操作指令(可以对应于长按操作),显示操作菜单栏33,其中包括语音标记图标34。相应地,响应于对语音标记图标34的点击指令,跳转至录音界面35。
在另一个具体的实施例中,在上述响应于针对目标图片的触发指令,显示语音标记图标之前,所述方法还可以包括:显示针对第一视频的播放界面;进一步地,在一个更具体的实施例中,响应于针对目标图片的触发指令,显示语音标记图标,可以包括:响应于针对所述第一视频发出的跳转指令,将跳转显示的视频帧确定为所述目标图片,并在所述播放界面显示所述语音标记图标。在另一个更具体的实施例中,响应于针对目标图片的触发指令,显示语音标记图标,可以包括:响应于针对所述第一视频发出的暂停指令,将暂停显示的视频帧确定为所述目标图片,并在所述播放界面显示所述语音标记图标。
根据一个具体的例子,如图4所示,其中显示针对第一视频的播放界面41,响应于对第一视频发出的暂停指令,在播放界面中显示语音标记图标42,接着,响应于对语音 标记图标42的点击指令,跳转至录音界面43(或称语音标记界面43),其中包括目标图片44,为暂停播放界面中的视频帧。如此,可以实现响应于对语音标记图标的触发指令,跳转至上述录音界面。
另一方面,在一个实施例中,本步骤中还可以包括:在所述录音界面中显示提示信息,用于向用户提示为所述目标图片添加语音标记的操作方式。在一个例子中,如图3所示,其中录音界面35所显示的提示信息36包括内容:请长按图片上某个位置开始录音,将在该位置生成语音标记。
以上,可以显示包括目标图片的录音界面。接着,在步骤S220,响应于基于所述录音界面发出的录音开始指令,持续采集语音信号。并且,在步骤S230,响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件,以及在步骤S240,在所述目标图片上添加关联所述音频文件的语音标记。
在一个实施例中,上述录音开始指令和录音结束指令可以分别对应于多种操作方式。在一个具体的实施例中,其中录音开始指令可以对应于:对所述目标图片的长按操作,并且,其中录音结束指令对应于:对所述目标图片的取消按压操作。在另一个具体的实施例中,录音界面中包括录音图标,其中录音开始指令可以对应于:对第一状态的录音图标的点击操作,并且,响应于此点击操作,将录音图标切换显示为第二状态,相应地,其中录音结束指令可以对应于:对第二状态的录音图标的点击操作,并且,响应于此点击操作,将录音图标切换显示为第一状态。在又一个具体的实施例中,所述录音开始指令对应于:对鼠标右键的点击操作,所述录音结束指令对应于:对鼠标右键的再次点击操作。在还一个具体的实施例中,所述录音开始指令对应于:对所述录音界面中录音开始图标的触发指令,所述录音结束指令对应于:对所述录音界面中录音结束图标的触发指令。
进一步地,在一个实施例中,步骤S220中可以包括:响应于基于所述录音界面中目标图片的第一位置发出的录音开始指令,持续采集语音信号。需要说明,其中第一位置可以为目标图片中任意的一个位置。相应地,在步骤S240可以包括:在所述目标图片的第一位置添加所述语音标记。
根据一个例子,如图5所示,在录音界面中,响应于对目标图片51中第一位置52发出的长按操作,开始录音,接着,响应于取消长按操作,结束录音,并在第一位置52添加与录制音频相关的语音标记53。如此,可以实现在指定的第一位置添加语音标记。
在另一个实施例中,步骤S240中可以包括:在目标图片中的任意位置添加所述语音 标记。在一个具体的实施例中,其中任意位置可以是根据随机算法确定的随机位置,也可以是系统默认的固定位置,如目标图片的中心区域。进一步地,在一个实施例中,在步骤S240之后,所述方法还可以包括:响应于对所述语音标记的移动指令,将所述语音标记移动至位于所述目标图片中的指定位置。
根据一个例子,如图6所示,在录音界面中包括目标图片61,响应于对第一状态的录音图标62的点击指令,开始录音,接着,响应于对第二状态的录音图标63的点击指令,结束录音,并在图片的中心区域64添加语音标记。进一步地,响应于对语音标记64的移动指令,将语音标记64移动至目标区域65。如此,可以实现将语音标记显示在指定位置。
需要说明,通过重复执行上述步骤S220-步骤S240,可以为目标图片添加多个语音标记,位于目标图片中的不同位置。
另一方面,在一个实施例中,步骤S240中还可以包括:在所述语音标记中显示序号。在一个具体的实施例中,其中序号可以是自动生成的,具体可以基于对所述目标图片添加的在先的语音标记的数量而确定。在一个更具体的实施例中,假定目标图片中已经包含n(为自然数)个语音标记,则对于新增的语音标记,其上显示的序号为n+1。在另一个具体的实施例中,其中序号可以是由用户自定义输入的。根据一个例子,如图7所示,用户将语音标记71上显示的序号由1修改为2。如此,可以方便用户基于序号标识的顺序,依次获悉语音标记对应的录制内容。
在一个实施例中,步骤S240中还可以包括:在所述语音标记中显示主题文本。对于上述主题文本的确定,在一个具体的实施例中,上述步骤S230中还可以包括:对采集到的语音信号进行语音识别,得到识别文本;再根据识别文本确定对应的主题文本。需要理解,其中语音识别可以采用现有技术实现,如语音识别模型等,在此不作赘述。在一个更具体的实施例中,其中根据识别文本确定对应的主题文本可以包括:将所述识别文本输入预先训练的摘要抽取模型中,得到对应的摘要文本,作为所述主题文本。在另一个更具体的实施例中,其中根据识别文本确定对应的主题文本可以包括:将所述识别文本输入预先训练的关键词提取模型中,得到对应的关键词,作为所述主题文本。根据一个例子,假定识别文本包括“这件衣服的衣领采用圆领设计,可以使人脖子显得更长”,将其输入关键词提取模型,可以得到对应的关键词为“衣领”,并将其作为主题文本,可参见图8,其中语音标记81中显示主题文本“衣领”。
在另一个实施例中,在步骤S240之后,所述方法还可以包括:接收用户基于所述语 音标记输入的自定义文本,并在该语音标记中显示该自定义文本。
又一方面,在一个实施例中,在步骤S240之后,所述方法还可以包括:响应于对所述语音标记的触发指令,播放所述音频文件。在一个具体的实施例中,其中,触发指令可以为声控指令或点击指令。在一个例子中,如图9所示,响应于对第一状态的语音标记91的点击指令,播放所述音频文件,同时,语音标记91切换显示为第二状态的语音标记92,用于提示用户正在播放录制音频。
在一个实施例中,在步骤S240之后,所述方法还可以包括:在所述目标图片上添加关联所述音频文件的语音标记之后,所述方法还包括:响应于对所述语音标记的触发指令,显示菜单栏,进一步地,在一个具体的实施例中,所述菜单栏中包括播放图标,响应于对播放图标的触发指令,播放所述音频文件。在另一个具体的实施例中,所述菜单栏中包括语音识别图标,响应于对所述语音识别图标的触发指令,对所述音频文件进行语音识别,得到识别文本;在所述录音界面中显示所述识别文本。在一个更具体的实施例中,在所述录音界面中显示所述识别文本之后,所述方法还可以包括:响应于对所述识别文本的隐藏指令,在所述录音界面中隐藏所述识别文本。
根据一个例子,如图10所示,响应于对语音标记101的点击指令,显示菜单栏102,其中包括语音识别图标103,响应于对语音识别图标103的触发指令,显示识别文本104,进一步地,响应于对折叠图标105的点击指令,隐藏所述识别文本。
再一方面,在一个实施例中,在步骤S240之后,所述方法还可以包括:响应于对所述语音标记的删除指令,从所述目标图片中删除所述语音标记。在一个具体的实施例中,响应于对语音标记的长按指令,显示删除按钮,响应于对删除按钮的点击指令,删除所述语音标记。如此,可以实现对语音标记的删除。
综上,在本说明书实施例披露的对图片进行语音标记的方法中,用户可以方便、快捷地在图片中任意的、需要进行标识的位置添加语音标记,如此大幅降低图片编辑成本,大大提高用户体验。
根据另一方面实施例,本说明书实施例还披露一种图片查看方法。具体地,图11示出根据一个实施例的图片查看方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、服务器集群,例如,可以为客户端(如图片处理软件)、系统软件或客户端插件(如即时通讯客户端中的插件)。如图11所示,所述方法包括以下步骤:
步骤S1110,显示带有语音标记的图片。需要理解,其中的语音标记前述实施例中描述的方法而添加。步骤S1120,响应于对所述语音标记的触发指令,播放所述语音标记关联的音频文件。
以上步骤具体如下:
首先,在步骤S1110,显示带有语音标记的图片。
在一个实施例中,本步骤可以包括:在聊天窗口中显示所述图片。在一个具体的实施例中,在即时通讯(如在线社交或在线客服咨询)的场景中,可以从其他即时通讯客户端接收所述图片,并在聊天窗口中对其进行显示。
在另一个实施例中,本步骤可以包括:在网页中加载所述图片。在一个具体的实施例中,响应于对网页的打开指令,进入所述网页,并在网页中加载所述图片。在一个具体的实施例中,其中网页可以为电商平台中的商品详情页,相应地,所述图片中的主体对象可以为商品。在一个例子中,如图12所示,其中示出的商品详情页包括带语音标记的商品图片121。在另一个具体的实施例中,其中网页可以为教学辅导网站,相应地,所述图片中的主体对象可以为试卷。
在一个实施例中,图片中的语音标记可以为一个或多个。
以上,可以显示带有语音标记的图片。进一步地,在步骤S1120,响应于对所述语音标记的触发指令,播放所述语音标记关联的音频文件。
在一个实施例中,图片中的语音标记可以为多个,相应地,本步骤中可以包括:基于多个语音标记被添加至所述图片中的顺序,依次播放所述多个语音标记对应的多个音频文件。
在另一个实施例中,图片中的语音标记可以为多个,其中各个语音标记中显示对应的序号,相应地,本步骤中可以包括:基于所述序号的顺序,依次播放所述多个语音标记对应的多个音频文件。
另一方面,上述步骤S1120还可以替换为:响应于对语音标记的触发指令,显示菜单栏,其中包括语音识别图标,响应于对语音识别图标的触发指令,将该语音标记对应的音频文件转换为语音识别文本,并对该语音识别文本进行显示。
综上,采用本说明书实施例披露的图片查看方法,使用户可以直观、便捷地获取其基于该图片需要了解的内容,进而提高用户体验。
根据再一方面的实施例,本说明书还披露一种对视频进行语音标记的方法。具体地, 图13示出根据一个实施例的对视频进行语音标记的方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,可以为客户端(如视频处理软件)、系统软件或客户端插件(如即时通讯客户端中的插件)。如图13所示,所述方法包括以下步骤:
步骤S1310,显示包括第一视频的录音界面,所述第一视频包括第一视频帧;步骤S1320,响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;步骤S1330,响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;步骤S1340,响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;步骤S1350,在所述目标视频帧上添加关联所述音频文件的语音标记。
以上步骤具体如下:
首先,在步骤S1310,显示包括第一视频的录音界面,所述第一视频包括第一视频帧。
在一个实施例中,在本步骤之前,所述方法还可以包括:响应于针对第一视频的触发指令,显示语音标记图标,相应地,本步骤中可以包括:响应于对所述语音标记图标的触发指令,跳转至所述录音界面。在一个具体的实施例中,其中针对第一视频的触发指令可以包括查看指令、或编辑指令、或发送指令。
在一个实施例中,本步骤可以包括:响应于基于录音界面发出的针对所述第一视频的导入指令,在录音界面中显示所述第一视频。在一个具体的实施例中,其中录音界面可以为视频处理软件中的编辑界面。在一个具体的实施例中,其中导入指令可以为声控指令或点击指令,如对导入图标的点击指令。在一个具体的例子中,响应于用户对视频处理APP的打开指令,进入上述录音界面,响应于对录音界面中视频导入图标的点击指令,显示终端中已有的视频,再响应于用户对其中某个视频的选取指令和确认导入指令,将该某个视频显示在上述录音界面。
以上,可以显示包括第一视频的录音界面。接着,在步骤S1320,响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧。
在一个实施例中,本步骤可以包括:响应于针对所述第一视频发出的跳转指令,将跳转显示的所述第一视频帧确定为所述目标视频帧。在一个具体的实施例中,其中跳转指令可以对应于:对视频进度条的拖动操作。在另一个具体的实施例中,其中跳转指令可以对应于:对第一视频帧的播放时间点的输入操作。
在另一个实施例中,本步骤可以包括:响应于针对所述第一视频发出的暂停指令, 将暂停显示的所述第一视频帧确定为所述目标视频帧。在一个具体的实施例中,在接收所述暂停指令之前,所述方法还可以包括:播放所述第一视频。
根据一个例子,如图14所示,响应于对进度条141的拖动操作,将跳转显示的第一视频帧142确定为目标视频帧。如此,可以确定目标视频帧。
然后,在步骤S1330,响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;步骤S1340,响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;步骤S1350,在所述目标视频帧上添加关联所述音频文件的语音标记。
在一个实施例中,步骤S1330可以包括:响应于基于所述录音界面中目标视频帧的第一位置发出的录音开始指令,持续采集语音信号;相应地,步骤S1350可以包括:在所述目标视频帧的第一位置添加所述语音标记。
需要说明的是,其中目标视频帧本质上也是图片,因此,对步骤S1330-步骤S1350的描述,可以参见前述实施例中对步骤S220-步骤S240的描述。
此外,在一个实施例中,在步骤S1350之后,所述方法还可以包括:在所述第一视频的进度条中添加时间点标记,所述时间点标记对应于所述目标视频帧的播放时间点。在一个具体的实施例中,对其中时间点标记的形状不作限定,可以为圆形、矩形或三角形等。如此,可以辅助用户快速定位到待语音标记的视频帧。进一步地,在一个具体的实施例中,在所述第一视频的进度条中添加时间点标记之后,所述方法还可以包括:响应于输入控件指示符移动至所述时间点标记,展示所述目标视频帧中所包括的语音标记的数量和/或主题文本。在一个更具体的实施例中,其中输入控件指示符可以为光标或者用户通过对屏幕的触摸操作产生的指示符。在一个更具体的实施例中,可以在弹窗或者说气泡中对语音标记的数量和/或主题文本进行显示。
根据一个例子,如图15所示,响应于将光标移动至时间点标记151,在聊天气泡152中显示对应视频帧中所包括语音标记的数量和各个语音标记对应的主题文本,如语音标记的数量:3,主题文本分别为:城门、城墙和古钟。如此,可以实现对视频中视频帧的语音标记和核查。
综上,通过本说明书实施例披露的对视频进行语音标记的方法,可以使得用户能够方便、快捷地对视频中任一视频帧中的任意位置添加语音标记,以丰富对视频进行编辑处理的方式,并且,节省原有的需要截取视频帧、编辑视频帧等所需要消耗的时间,从而提高用户体验。
根据再一方面的实施例,本说明书还披露一种视频查看方法。具体地,图16示出根据一个实施例的视频查看方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,可以为客户端(如视频处理软件)、系统软件或客户端插件(如即时通讯客户端中的插件)。如图16所示,所述方法包括以下步骤:
步骤S1610,显示带有语音标记的视频,所述视频中包括带有第一语音标记的第一视频帧。需要理解,其中的语音标记通过前述实施例中描述的对视频进行语音标记的方法而添加。步骤S1620,响应于对所述第一语音标记的触发指令,播放所述第一语音标记关联的音频文件。
以上步骤具体如下:
首先,在步骤S1610,显示带有语音标记的视频,其中包括带有第一语音标记的第一视频帧。
在一个实施例中,本步骤可以包括:在聊天窗口中显示所述视频。在一个具体的实施例中,在即时通讯(如在线社交或在线客服咨询)的场景中,可以从其他即时通讯客户端接收所述视频,并在聊天窗口中对其进行显示。
在另一个实施例中,本步骤可以包括:在网页中加载所述视频。在一个具体的实施例中,响应于对网页的打开指令,进入所述网页,并在网页中加载所述视频。在一个具体的实施例中,其中网页可以为电商平台中的商品详情页。在另一个具体的实施例中,其中网页可以为教学辅导网站。在又一个具体的实施例中,其中网页可以为视频播放平台提供的播放页面。
在一个实施例中,上述视频中带有语音标记的视频帧可以为一个或多个。进一步地,对于这些视频帧中的任一视频帧,所带有语音标记的数量可以为一个或多个。
接着,在步骤S1620,响应于对所述第一语音标记的触发指令,播放所述第一语音标记关联的音频文件。在一个例子中,如图17所示,响应于将光标移动至语音标记171的操作,播放其关联的音频文件。
需要说明的是,对步骤S1620的描述还可以参见前述对步骤1120等的相关描述。
此外,在一个实施例中,第一视频帧中包括含所述第一语音标记在内的若干语音标记,所述视频的进度条在所述第一视频帧的播放时间点显示有时间点标记。相应地,在步骤S1620之后,所述方法还可以包括:响应于输入控件指示符移动至所述时间点标记, 展示所述若干语音标记的标记数量和/或所述若干语音标记对应的若干主题文本。对此可以参见前述实施例中的相关描述,不作赘述。
综上,采用本说明书实施例披露的视频查看方法,使用户可以直观、便捷地获取针对部分视频帧做出的特定说明内容,进而提高用户体验。
以上,主要对在多媒体媒介(如图片、视频,实际还可以包括电子文稿等)上添加语音标记,以及对带有语音标记的多媒体媒介进行查看的方法进行说明。进一步地,发明人提出,还可以拓展到文本标记的范畴,具体地,添加文本标记可以结合录音功能和语音识别技术实现。
图18示出根据一个实施例的对图片进行文本标记的方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,可以为客户端(如图片处理软件)、系统软件或客户端插件(如即时通讯客户端中的插件)。如图18所示,所述方法包括以下步骤:
步骤S1810,显示针对目标图片的编辑界面;步骤S1820,响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;步骤S1830,响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;步骤S1840,在所述目标图片上添加关联所述识别文本的文本标记。
以上步骤具体如下:
首先,在步骤S1810,显示针对目标图片的编辑界面。
在一个实施例中,可以在图片处理或图片编辑软件中导入目标图片,进而显示针对目标图片的编辑界面。在另一个实施例中,可以先在图库中选取目标图片,再点击编辑图标,进而进入包含该目标图片的编辑界面。需要理解,此处编辑界面除提供添加语音标记的功能外,还可以提供其他图片处理功能,如涂鸦等。
以上,可以显示针对目标图片的编辑界面。接着,可以在步骤S1820,响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;步骤S1830,响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;步骤S1840,在所述目标图片上添加关联所述识别文本的文本标记。
在一个实施例中,步骤S1820中可以包括:响应于基于所述编辑界面中目标图片的第一位置发出的录音开始指令,持续采集语音信号,相应地,在步骤S1840中还可以包括:在所述目标图片的第一位置添加所述文本标记。
在一个实施例中,步骤S1840中还可以包括:在所述文本标记中显示所述识别文本。
在一个实施例中,步骤S1840中还可以包括:在所述文本标记中对所述识别文本进行折叠显示。相应地,在步骤S1840之后,所述方法还可以包括:响应于对所述文本标记的触发指令,在所述文本标记中对所述识别文本进行展开显示。
另一方面,在一个实施例中,步骤S1840中还可以包括:在所述文本标记中显示序号,所述序号基于对所述目标图片添加的在先的文本标记的数量而确定。在另一个实施例中,步骤S1830中还可以包括:确定所述识别文本对应的主题文本,相应地,在步骤S1840中还可以包括:在所述文本标记中显示所述主题文本。
进一步地,在步骤S1840之后,所述方法还可以包括:响应于对所述文本标记的触发指令,展示所述识别文本。
根据一个具体的例子,如图19所示,在新添加的文本标记191中显示序号3,响应于对文本标记191的点击指令,展示对应的识别文本192。
综上,在本说明书实施例披露的对图片进行文本标记的方法中,用户可以方便、快捷地在图片中任意的、需要进行标识的位置添加文本标记,如此大幅降低图片编辑成本,大大提高用户体验。
根据另一方面的实施例,本说明书实施例还披露一种图片查看方法,具体地,图20示出根据一个实施例的图片查看方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、服务器集群,例如,可以为客户端(如图片处理软件)、系统软件或客户端插件(如即时通讯客户端中的插件)。如图20所示,所述方法包括以下步骤:
步骤S2010,显示带有文本标记的图片,需要理解,其中文本标记是基于前述实施例中描述的方法添加的。步骤S2020,响应于对所述文本标记的触发指令,展示所述文本标记关联的识别文本。
以上步骤具体如下:
首先,在步骤S2010,显示带有文本标记的图片。在一个实施例中,本步骤还可以包括:在所述文本标记中显示序号。在另一个实施例中,本步骤还可以包括:在所述文本标记中显示所述识别文本对应的主题文本。在又一个实施例中,本步骤还可以包括:在所述文本标记中对识别文本进行折叠显示。
接着,在步骤S2020,响应于对所述文本标记的触发指令,展示所述文本标记关联 的识别文本。在一个实施例中,本步骤可以包括:在所述文本标记中对所述识别文本进行展开显示。
需要说明,在步骤S2020之后,所述方法还可以包括:响应于对识别文本的收起指令,对文本标记进行还原显示。
根据一个具体的例子,如图21所示,响应于对文本标记211的点击指令,对对应的识别文本212进行展开显示,进一步地,响应于对折叠图标213的点击指令,对识别文本进行折叠显示。
综上,采用本说明书实施例披露的图片查看方法,使用户可以直观、便捷地获取其基于该图片需要了解的内容,进而提高用户体验。
根据再一方面的实施例,本说明书还披露一种对视频进行文本标记的方法。具体地,图22示出根据一个实施例的对视频进行文本标记的方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,可以为客户端(如视频处理软件)、系统软件或客户端插件(如即时通讯客户端中的插件)。如图22所示,所述方法包括以下步骤:
步骤S2210,显示包括第一视频的编辑界面,所述第一视频包括第一视频帧;步骤S2220,响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;步骤S2230,响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;步骤S2240,响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;步骤S2250,在所述目标视频帧上添加关联所述识别文本的文本标记。
针对以上步骤,在一个实施例中,步骤S2230中可以包括:响应于基于所述编辑界面中目标视频帧的第一位置发出的录音开始指令,持续采集语音信号,相应地,在步骤S2250中可以包括:在所述目标视频帧的第一位置添加所述文本标记。
根据一个具体的例子,如图23所示,在视频中的某个视频帧中添加文本标记231,响应于对文本标记231的点击指令,展示对应的识别文本232。
需要说明,对步骤S2210-步骤S2250的介绍,还可以参见前述实施例中的相关描述。
综上,通过本说明书实施例披露的对视频进行文本标记的方法,可以使得用户能够方便、快捷地对视频中任一视频帧中的任意位置添加文本标记,以丰富对视频进行编辑处理的方式,并且,节省原有的需要截取视频帧、编辑视频帧等所需要消耗的时间,从 而提高用户体验。
根据再一方面的实施例,本说明书还披露一种视频查看方法。具体地,图24示出根据一个实施例的视频查看方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,可以为客户端(如视频处理软件)、系统软件或客户端插件(如即时通讯客户端中的插件)。如图24所示,所述方法包括以下步骤:
步骤S2410,显示带有文本标记的视频,所述视频中包括带有第一文本标记的第一视频帧,需要说明,其中文本标记通过前述实施例中的方法添加;步骤S2420,响应于对所述第一文本标记的触发指令,展示所述文本标记关联的识别文本。
根据一个例子,如图25所示,响应于将光标移动至视频帧中文本标记251的指令,在视频播放界面中,除视频帧以外的位置显示对应的识别文本252。
综上,采用本说明书实施例披露的视频查看方法,使用户可以直观、便捷地获取针对部分视频帧做出的特定说明内容,进而提高用户体验。
与上述方法相对应的,本说明书实施例还披露多种装置。具体如下:
图26示出根据一个实施例的对图片进行语音标记的装置结构图。如图26所示,所述装置2600包括:
显示单元2610,显示包括目标图片的录音界面;采集单元2620,响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;存储单元2630,响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;添加单元2640,在所述目标图片上添加关联所述音频文件的语音标记。
在一个实施例中,其中显示单元2610具体配置为:响应于基于录音界面发出的针对所述目标图片的导入指令,在录音界面中显示所述目标图片。
在一个实施例中,所述装置2600还包括:触发单元,配置为响应于针对目标图片的触发指令,显示语音标记图标;其中显示单元2610具体配置为:响应于对所述语音标记图标的触发指令,跳转至所述录音界面。
在一个具体的实施例中,其中触发单元具体配置为:响应于针对所述目标图片的查看指令,显示所述语音标记图标;或,响应于针对所述目标图片的编辑指令,显示所述语音标记图标,或,响应于用于截取所述目标图片的截屏指令,显示所述语音标记图标; 或,响应于针对所述目标图片的发送指令,显示所述语音标记图标。
在另一个具体的实施例中,所述装置2600还包括:播放界面显示单元,配置为显示针对第一视频的播放界面;其中触发单元具体配置为:响应于针对所述第一视频发出的跳转指令,将跳转显示的视频帧确定为所述目标图片,并在所述播放界面显示所述语音标记图标;或,响应于针对所述第一视频发出的暂停指令,将暂停显示的视频帧确定为所述目标图片,并在所述播放界面显示所述语音标记图标。
在一个实施例中,显示单元2610还配置为:在所述录音界面中显示提示信息,用于向用户提示为所述目标图片添加语音标记的操作方式。
在一个实施例中,采集单元2620具体配置为:响应于基于所述录音界面中目标图片的第一位置发出的录音开始指令,持续采集语音信号;其中添加单元2640具体配置为:在所述目标图片的第一位置添加所述语音标记。
在一个实施例中,所述录音开始指令对应于:对所述目标图片的长按操作,所述录音结束指令对应于:对所述目标图片的取消按压操作。
在一个实施例中,所述添加单元2640还配置为:在所述语音标记中显示序号,所述序号基于对所述目标图片添加的在先的语音标记的数量而确定或由用户自定义输入。
在一个实施例中,所述存储单元2630还配置为:对采集到的语音信号进行语音识别,得到识别文本;确定所述识别文本对应的主题文本;其中添加单元2640还配置为:在所述语音标记中显示所述主题文本。
在一个具体的实施例中,其中存储单元2630具体配置为:将所述识别文本输入预先训练的摘要抽取模型中,得到对应的摘要文本,作为所述主题文本;或,将所述识别文本输入预先训练的关键词提取模型中,得到对应的关键词,作为所述主题文本。
在一个实施例中,所述装置2600还包括:接收单元,配置为接收用户基于所述语音标记输入的自定义文本;第一文本显示单元,配置为在所述语音标记中显示所述自定义文本。
在一个实施例中,所述装置2600还包括:播放单元,配置为响应于对所述语音标记的触发指令,播放所述音频文件。
在一个实施例中,所述装置2600还包括:移动单元,配置为响应于对所述语音标记的移动指令,将所述语音标记移动至位于所述目标图片中的指定位置。
在一个实施例中,所述装置2600还包括:删除单元,配置为响应于对所述语音标记的删除指令,从所述目标图片中删除所述语音标记。
在一个实施例中,所述装置2600还包括:菜单栏显示单元,配置为响应于对所述语音标记的触发指令,显示菜单栏,所述菜单栏中包括语音识别图标;识别单元,配置为响应于对所述语音识别图标的触发指令,对所述音频文件进行语音识别,得到识别文本;第二文本显示单元,配置为在所述录音界面中显示所述识别文本。
在一个具体的实施例中,所述装置2600还包括:隐藏单元,配置为响应于对所述识别文本的隐藏指令,在所述录音界面中隐藏所述识别文本。
图27示出根据一个实施例的图片查看装置结构图。如图27所示,所述装置2700包括:
显示单元2710,配置为显示带有语音标记的图片,所述语音标记通过上述装置2600添加至所述图片中;播放单元2720,配置为响应于对所述语音标记的触发指令,播放所述语音标记关联的音频文件。
在一个实施例中,其中显示单元2710具体配置为:在聊天窗口中显示所述图片;或,在网页中加载所述图片。
在一个具体的实施例中,所述图片中包括目标商品,所述网页为商品详情页。
在一个实施例中,所述语音标记为多个语音标记,其中播放单元2720具体配置为:基于所述多个语音标记被添加至所述图片中的顺序,依次播放所述多个语音标记对应的多个音频文件。
在一个实施例中,所述语音标记为多个语音标记,其中各个语音标记中显示对应的序号;其中播放单元2720具体配置为:基于所述序号的顺序,依次播放所述多个语音标记对应的多个音频文件。
图28示出根据一个实施例的对视频进行语音标记的装置结构图。如图28所示,所述装置2800包括:
显示单元2810,配置为显示包括第一视频的录音界面,所述第一视频包括第一视频帧;确定单元2820,配置为响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;采集单元2830,配置为响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;存储单元2840,配置为响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;添加单元2850,配置为在所述目标视频帧上添加关联所述音频文件的语音标记。
在一个实施例中,采集单元2830具体配置为:响应于基于所述录音界面中目标视频帧的第一位置发出的录音开始指令,持续采集语音信号;添加单元2850具体配置为:在所述目标视频帧的第一位置添加所述语音标记。
在一个实施例中,确定单元2820具体配置为:响应于针对所述第一视频发出的跳转指令,将跳转显示的所述第一视频帧确定为所述目标视频帧;或,响应于针对所述第一视频发出的暂停指令,将暂停显示的所述第一视频帧确定为所述目标视频帧。
在一个实施例中,所述装置2800还包括:进度条标记单元,配置为在所述第一视频的进度条中添加时间点标记,所述时间点标记对应于所述目标视频帧的播放时间点。
在一个具体的实施例中,所述装置2800还包括:第一展示单元,配置为响应于输入控件指示符移动至所述时间点标记,展示所述目标视频帧中所包括的语音标记的数量。
在另一个具体的实施例中,存储单元2840还配置为:对采集到的语音信号进行语音识别,得到识别文本;确定所述识别文本对应的主题文本;所述装置还包括:第二展示单元,配置为响应于输入控件指示符移动至所述时间点标记,展示所述主题文本。
图29示出根据一个实施例的视频查看装置结构图。如图29所示,所述装置2900包括:
显示单元2910,配置为显示带有语音标记的视频,所述语音标记通过上述装置2800添加至所述视频中,所述视频中包括带有第一语音标记的第一视频帧;播放单元2920,配置为响应于对所述第一语音标记的触发指令,播放所述第一语音标记关联的音频文件。
在一个实施例中,所述第一视频帧中包括含所述第一语音标记在内的若干语音标记,所述视频的进度条在所述第一视频帧的播放时间点显示有时间点标记;所述装置还包括:展示单元,配置为响应于输入控件指示符移动至所述时间点标记,展示所述若干语音标记的标记数量和/或所述若干语音标记对应的若干主题文本。
图30示出根据一个实施例的对图片进行文本标记的装置结构图。如图30所示,所述装置3000还包括:
显示单元3010,配置为显示针对目标图片的编辑界面;采集单元3020,配置为响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;识别单元3030,配置为响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;添加单元3040,配置为在所述目标图片上添加关联所述识别文本的文本标 记。
在一个实施例中,采集单元3020具体配置为:响应于基于所述编辑界面中目标图片的第一位置发出的录音开始指令,持续采集语音信号;添加单元3040具体配置为:在所述目标图片的第一位置添加所述文本标记。
在一个实施例中,添加单元3040还配置为:在所述文本标记中显示序号,所述序号基于对所述目标图片添加的在先的文本标记的数量而确定。
在一个实施例中,添加单元3040还配置为:在所述文本标记中对所述识别文本进行折叠显示;所述装置还包括展开单元,配置为响应于对所述文本标记的触发指令,在所述文本标记中对所述识别文本进行展开显示。
在一个实施例中,所述装置3000还包括:确定单元,配置为确定所述识别文本对应的主题文本;其中添加单元3040具体配置为:在所述文本标记中显示所述主题文本。
在一个实施例中,所述装置3000还包括展示单元,配置为响应于对所述文本标记的触发指令,展示所述识别文本。
图31示出根据一个实施例的图片查看装置结构图。如图31所示,所述装置3100包括:
显示单元3110,配置为显示带有文本标记的图片,所述文本标记通过上述装置3000添加至所述图片中;展示单元3120,配置为响应于对所述文本标记的触发指令,展示所述文本标记关联的识别文本。
在一个实施例中,其中显示单元3120具体配置为:在所述文本标记中显示序号;或,在所述文本标记中显示所述识别文本对应的主题文本。
在一个实施例中,其中显示单元3120具体配置为:在所述文本标记中对所述识别文本进行折叠显示;展示单元3120具体配置为:在所述文本标记中对所述识别文本进行展开显示。
图32示出根据一个实施例的对视频进行文本标记的装置结构图。如图32所示,所述装置3200包括:
显示单元3210,配置为显示包括第一视频的编辑界面,所述第一视频包括第一视频帧;确定单元3220,配置为响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;采集单元3230,配置为响应于基于所述编辑界面发出的录音开始指令, 持续采集语音信号;识别单元3240,配置为响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;添加单元3250,配置为在所述目标视频帧上添加关联所述识别文本的文本标记。
在一个实施例中,采集单元3230,配置为响应于基于所述编辑界面中目标视频帧的第一位置发出的录音开始指令,持续采集语音信号;其中添加单元3250具体配置为:在所述目标视频帧的第一位置添加所述文本标记。
图33示出根据一个实施例的视频查看装置结构图。如图33所示,所述装置3300包括:
显示单元3310,配置为显示带有文本标记的视频,所述语音标记通过上述装置3200添加至所述视频中,所述视频中包括带有第一文本标记的第一视频帧;展示单元3320,配置为响应于对所述第一文本标记的触发指令,展示所述文本标记关联的识别文本。
需要说明的是,上述语音标记添加功能可以应用于多个场景中。下面结合具体的应用场景,对添加语音标记的方法进行进一步说明。具体如下:
图34示出根据一个实施例的图片处理方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,即时通讯客户端。如图34所示,所述方法包括以下步骤:
步骤S3410,显示聊天界面,并接收基于所述聊天界面选取的待发送的目标图片;步骤S3420,响应于针对所述目标图片的编辑指令,进入图片编辑界面,其中显示语音标记图标;步骤S3430,响应于对所述语音标记图标的触发指令,进入录音界面;步骤S3440,将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
在一个实施例中,步骤S3440中可以包括:在所述语音标记中显示当前编辑用户的聊天昵称和/或聊天头像。在一个例子中,图35中聊天窗口显示的当前用户发送的目标图片3501中,显示带有当前用户头像的语音标记3502。
在一个实施例中,在执行步骤S3440之后,所述方法还可以包括:响应于对目标图片的发送指令,发送带有语音标记的目标图片。
由上,可以实现在聊天软件中,给图片添加语音标记,并发送带有语音标记的目标图片。
图36示出根据另一个实施例的图片处理方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,即时通讯客户端。如图36所示,所述方法包括以下步骤:
步骤S3610,显示聊天界面,所述聊天界面的聊天窗口中包含目标图片;步骤S3620,响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;步骤S3630,响应于对所述语音标记图标的触发指令,进入录音界面;步骤S3640,将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
针对以上步骤,在一个例子中,图3示出的聊天界面的聊天窗口中包含目标图片32,响应于针对所述目标图片的触发指令,显示菜单栏33,所述菜单栏33中包括语音标记图标34,响应于对所述语音标记图标的触发指令,进入录音界面。进一步地,可以为该目标图片添加语音标记。
在一个实施例中,在步骤S3610之前,所述方法还可以包括:从所述聊天窗口对应的联系人接收所述目标图片,所述目标图片中包括已有语音标记,其中显示所述联系人的聊天昵称和/或聊天头像。相应地,在步骤S3640中可以包括:在所述语音标记中显示当前编辑用户的聊天昵称和/或聊天头像。
在一个例子中,图37示出根据另一个实施例的聊天界面示意图,界面的聊天窗口中包括从当前联系人接收的图片3701,其上带有语音标记3702中显示当前联系人头像,该聊天窗口中还包括当前用户对图片3701进行编辑后发送的图片3703,相较于图片3701,其上还带有语音标记3704,其中显示当前用户的头像。
在一个实施例中,在执行步骤S3640之后,所述方法还可以包括:响应于针对所述录音界面的退出指令,在所述聊天窗口将原始的目标图片更新显示为带所述语音标记的目标图片。进一步地,在一个具体的实施例中,还可以在聊天窗口中显示提示信息,以告知各方联系人所述目标图片已发生修改。
在一个例子中,图38示出根据又一个实施例的聊天界面示意图,聊天窗口中的图片3801更新显示为带有语音标记3802的图片3803,并且显示提示信息3804,即,“小油条”已在图片“1.jpg”中添加语音标记。
由上,可以实现在即时通讯场景中,不同用户各自为同一图片添加语音标记,如此可以提高用户之间的沟通效率。
图39示出根据又一个实施例的图片处理方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,即时通讯客户端。如图39所示,所述方法包括以下步骤:
步骤S3910,显示聊天界面,所述聊天界面的聊天窗口中包含带有第一语音标记的目标图片,所述第一语音标记由所述聊天窗口对应的当前联系人添加;步骤S3920,响应于对所述第一语音标记的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;步骤S3930,响应于对所述语音回复图标的触发指令,进入录音界面;步骤S3940,将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
在一个例子中,图40示出根据还一个实施例的聊天界面示意图,如图40所示,聊天窗口中包含带有语音标记4001的图片4002,响应于对语音标记4001的点击指令,显示菜单栏,其中包括语音回复图标4003,响应于对语音回复图标4003的触发指令,进入录音界面,进一步地,针对用户输入语音,可以在语音标记4001的附近添加语音标记4004。如图41所示,可以将用户输入语音添加至语音标记4001对应的音频文件中,并将显示序号1的语音标记4001更新显示为显示序号2的4101,其中序号1和2的含义分别为:编辑的不同用户数,或者编辑的总次数。
在一个实施例中,在执行步骤S3940之后,所述方法还可以包括:响应于针对所述录音界面的退出指令,在所述聊天窗口中显示提示信息,以告知各方联系人所述目标图片中的语音标记已发生修改。在一个例子中,图41示出根据再一个实施例的聊天界面示意图,如图41所示,聊天界面的聊天窗口中显示提示信息4102,“小油条”已更新图片“1.jpg”中的语音标记内容,点击可收听。
由上,可以实现在即时通讯场景中,对语音标记的直接回复,此种回复带有更强的指向性,方便用户直观、快捷地针对图片指定位置的内容进行沟通。
图42示出根据再一个实施例的图片处理方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,即图片编辑软件。如图42所示,所述方法包括以下步骤:
步骤S4210,显示包含目标图片的图片编辑界面,此界面的功能菜单中包括语音标记图标;步骤S4220,响应于对所述语音标记图标的触发指令,进入语音标记界面;步骤 S4230,将基于所述语音标记界面接收的输入文本转化为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
通过以上步骤,可以实现通过输入文本的方式,为图片添加语音标记,从而丰富用户添加语音标记的交互方式,提高用户体验。
图43示出根据一个实施例的图片处理方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,即图片编辑软件。如图43所示,所述方法包括以下步骤:
步骤S4310,显示包含目标图片的聊天界面;步骤S4320,响应于对所述目标图片的触发指令,显示菜单栏,其中包括语音添加表情图标;步骤S4330,响应于对所述语音添加表情图标的触发指令,进入语音添加表情界面;步骤S4340,将基于所述语音添加表情界面采集的语音信号转化为文字,并基于所述文字生成动画表情;步骤S4350,在所述目标图片中添加所述动画表情。
针对以上步骤,在一个实施例中,步骤S4340中可以包括:从表情库中检索与所述文字相关的原始表情;在所述原始表情中添加所述文字,得到所述动画表情。根据一个例子,如图44所示,其中图片(某图)包括用户通过语音输入(如,输入内容为哈哈哈)添加的动画表情4401。
采用以上方法,在即时通讯场景下,使得用户可以实现通过语音输入,生成动画表情,对目标图片进行编辑或回复,从而提高用户之间聊天的趣味性和沟通的便捷性。
图45示出根据另一个实施例的图片处理方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,电商平台。如图45所示,所述方法包括以下步骤:
步骤S4510,显示商品信息编辑界面,其中包括针对目标商品的目标图片;步骤S4520,响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;步骤S4530,响应于对所述语音标记图标的触发指令,进入录音界面;步骤S4540,将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据一个例子,图46示出根据一个实施例的商品信息编辑界面示意图,如图46所示,商品信息编辑界面4601中包括商品图片4602,响应于对商品图片4602的触发指令, 可以显示包括语音标记图标4603的菜单栏,进一步地,可以通过触发语音标记图标4603进入录音界面,为商品图片4602编辑语音标记,包括增加、删除和修改。
在一个实施例中,在步骤S4540之后,所述方法还可以包括:在针对所述目标商品的商品详情页中,显示所述目标图片。在一个例子中,图12示出的商品详情页中包括带语音标记的目标图片121。
以上,可以实现在商品详情页中,通过带语音标记的商品介绍图片,更加直观地向用户展示目标商品。
图47示出根据又一个实施例的图片处理方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,电商平台。如图47所示,所述方法包括以下步骤:
步骤S4710,显示针对第一订单的订单评价界面,其中包括添加图片图标;步骤S4720,响应于针对所述添加图片图标的触发指令,接收选取的目标图片;步骤S4730,响应于对所述目标图片发出的语音标记指令,进入录音界面;步骤S4740,将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
由上,可以使用户可以更加有针对性地对产品做出评价。
图48示出根据还一个实施例的图片处理方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,电商平台。如图48所示,所述方法包括以下步骤:
步骤S4810,显示针对目标商品的商品评价界面,其中包括第一用户评价,所述第一用户评价中包括带第一语音标记的目标图片;步骤S4820,响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;步骤S4830,响应于对所述语音回复图标的触发指令,进入录音界面;步骤S4840,将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
针对以上步骤,在一个实施例中,在执行步骤S4840之后,所述方法还可以包括:响应于针对所述录音界面的退出指令,在所述商品评价界面中显示针对所述第一用户评 价的通知消息,用于通知所述第一用户评价已被回复。
以上,可以实现在评价场景下,卖家和买家,或者买家与买家之间基于商品图片的快捷地、有针对性地沟通。
图49示出根据再一个实施例的图片处理方法流程图,所述方法的执行主体可以为客服平台。如图49所示,所述方法包括以下步骤:
步骤S4910,接收用户发送的会话消息,所述会话消息中包括带有第一语音标记的目标图片;步骤S4920,获取与所述第一语音标记关联的音频文件,对所述音频文件进行语音识别,得到识别文本;步骤S4930,将所述识别文本输入预先训练的用户标问预测模型中,输出对应的用户标准问题;步骤S4940,将与所述用户标准问题对应的问题答案反馈给所述用户。
针对以上步骤,在一个实施例中,步骤S4940可以包括:将所述问题答案转化成答复音频,并在所述目标图片中添加与所述答复音频关联的第二语音标记;将新添加所述第二语音标记的目标图片发送给所述用户。
在另一个实施例中,步骤S4940可以包括:将所述问题答案转化成答复音频;在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述答复音频的第二语音标记;或者,将所述答复音频,添加至所述第一语音标记对应的音频文件中;向用户发送提示信息,以提示用户所述目标图片中已添加答复内容。
以上,可以实现客服场景下,用户和客服之间的便捷沟通。
图50示出根据还一个实施例的电子文件的处理方法流程图,所述方法的执行主体可以为任何具有计算、处理能力的装置、设备、平台、设备集群,例如,办公软件或审核平台等。如图50所示,所述方法包括以下步骤:
步骤S5010,显示针对目标电子文件的文件处理界面,所述文件处理界面的功能菜单栏中包括语音标记图标;步骤S5020,响应于对所述语音标记图标的触发指令,进入录音界面;步骤S5030,将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
针对以上步骤,在一个实施例中,所述目标电子文件为电子合同,所述文件处理界面为合同审批界面。在一个实施例中,所述目标电子文件的文件格式为word文档、或PDF文档、或excel表格。在一个例子中,图51示出根据一个实施例的办公软件界面示意图,界面菜单栏中包括语音标记图标5101,其中展示word文档中带有语音标记5102。
以上,可以实现对文件进行语音标记。
图52示出根据一个实施例的图片处理方法流程图,所述方法的执行主体为直播平台。如图52所示,所述方法包括:
步骤S5210,显示商品信息编辑界面,其中包括针对待上架的目标商品的目标图片;步骤S5220,响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;步骤S5230,响应于对所述语音标记图标的触发指令,进入录音界面;步骤S5240,将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
在一个实施例中,在步骤S5240之后,所述方法还包括:显示直播界面,所述直播界面中包括商品上架图标;响应于对所述商品上架图的触发指令,显示待上架商品的图片,其中包括所述目标图片;响应于对所述目标图片的选取指令,在所述直播界面的商品展示窗口中展示所述目标图片。在一个例子中,图53示出根据一个实施例的直播界面示意图,该直播界面的商品展示窗口中对带有语音标记5301的目标图片5302进行展示。
以上,可以实现在直播界面中展示带语音标记的商品图片,以方便收看直播的用户更加快捷地了解目标商品。
与上述方法相对应的,本说明书实施例还披露多种处理转置。具体如下:
图54示出根据一个实施例的图片处理装置结构图。如图54所示,所述装置5400包括:
显示单元5410,配置为显示聊天界面;接收单元5420,配置为收基于所述聊天界面选取的待发送的目标图片;第一界面切换单元5430,配置为响应于针对所述目标图片的编辑指令,进入图片编辑界面,其中显示语音标记图标;第二界面切换单元5440,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元5450,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
图55示出根据另一个实施例的图片处理装置结构图。如图55所示,所述装置5500包括:
界面显示单元5510,配置为显示聊天界面,所述聊天界面的聊天窗口中包含目标图 片;菜单栏显示单元5520,配置为响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;界面切换单元5530,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元5540,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
图56示出根据又一个实施例的图片处理装置结构图。如图56所示,所述装置5600包括:
界面显示单元5610,配置为显示聊天界面,所述聊天界面的聊天窗口中包含带有第一语音标记的目标图片,所述第一语音标记由所述聊天窗口对应的当前联系人添加;菜单栏显示单元5620,配置为响应于对所述第一语音标记的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;界面切换单元5630,配置为响应于对所述语音回复图标的触发指令,进入录音界面;标记单元5640,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
图57示出根据再一个实施例的图片处理装置结构图。如图57所示,所述装置5700包括:
显示单元5710,配置为显示包含目标图片的图片编辑界面,所述图片编辑界面的功能菜单中包括语音标记图标;界面切换单元5720,配置为响应于对所述语音标记图标的触发指令,进入语音标记界面;标记单元5730,配置为将基于所述语音标记界面接收的输入文本转化为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
图58示出根据还一个实施例的图片处理装置结构图。如图58所示,所述装置5800包括:
界面显示单元5810,配置为显示包含目标图片的聊天界面;菜单栏显示单元5820,配置为响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音添加表情图标;界面切换单元5830,配置为响应于对所述语音添加表情图标的触发指令,进入语音添加表情界面;表情生成单元5840,配置为将基于所述语音添加表情界面采集的语音信号转化为文字,并基于所述文字生成动画表情;表情添加单元5850,配置为在所述 目标图片中添加所述动画表情。
图59示出根据一个实施例的图片处理装置结构图,所述装置集成于电商平台,所述装置5900包括:
界面显示单元5910,配置为显示商品信息编辑界面,其中包括针对目标商品的目标图片;菜单栏显示单元5920,配置为响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;界面切换单元5930,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元5940,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
图60示出根据又一个实施例的图片处理装置结构图。如图60所示,所述装置6000包括:
显示单元6010,配置为显示针对第一订单的订单评价界面,其中包括添加图片图标;接收单元6020,配置为响应于针对所述添加图片图标的触发指令,接收选取的目标图片;界面切换单元6030,配置为响应于对所述目标图片发出的语音标记指令,进入录音界面;标记单元6040,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
图61示出根据另一个实施例的图片处理装置结构图。如图61所示,所述装置6100包括:
界面显示单元6110,配置为显示针对目标商品的商品评价界面,其中包括第一用户评价,所述第一用户评价中包括带第一语音标记的目标图片;菜单栏显示单元6120,配置为响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;界面切换单元6130,配置为响应于对所述语音回复图标的触发指令,进入录音界面;标记单元6140,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
图62示出根据再一个实施例的图片处理装置结构图,所述装置集成于客服平台,所述装置6200包括:
接收单元6210,配置为接收用户发送的会话消息,所述会话消息中包括带有第一语 音标记的目标图片;获取单元6220,配置为获取与所述第一语音标记关联的音频文件,对所述音频文件进行语音识别,得到识别文本;预测单元6230,配置为将所述识别文本输入预先训练的用户标问预测模型中,输出对应的用户标准问题;反馈单元6240,配置为将与所述用户标准问题对应的问题答案反馈给所述用户。
图63示出根据一个实施例的电子文件的处理装置结构图,所述装置6300包括:
显示单元6310,配置为显示针对目标电子文件的文件处理界面,所述文件处理界面的功能菜单栏中包括语音标记图标;界面切换单元6320,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元6330,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
图64示出根据还一个实施例的图片处理装置结构图,所述装置集成于直播平台,所述装置6400包括:
界面显示单元6410,配置为显示商品信息编辑界面,其中包括针对待上架的目标商品的目标图片;菜单栏显示单元6420,配置为响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;界面切换单元6430,配置为响应于对所述语音标记图标的触发指令,进入录音界面;标记单元6440,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
根据又一方面的实施例,还提供一种计算机可读存储介质,其上存储有计算机程序,当所述计算机程序在计算机中执行时,令计算机执行结合图2或图11或图13或图16或图18或图20或图22或图24或图34或图36或图39或图42或图45或图47或图48或图49或图50或图52所描述的方法。
根据再一方面的实施例,还提供一种计算设备,包括存储器和处理器,所述存储器中存储有可执行代码,处理器执行所述可执行代码时,实现结合图2或图11或图13或图16或图18或图20或图22或图24或图34或图36或图39或图42或图45或图47或图48或图49或图50或图52所描述的方法。
本领域技术人员应该可以意识到,在上述一个或多个示例中,本发明所描述的功能可以用硬件、软件、固件或它们的任意组合来实现。当使用软件实现时,可以将这些功能存储在计算机可读介质中或者作为计算机可读介质上的一个或多个指令或代码进行传 输。
以上所述的具体实施方式,对本发明的目的、技术方案和有益效果进行了进一步详细说明,所应理解的是,以上所述仅为本发明的具体实施方式而已,并不用于限定本发明的保护范围,凡在本发明的技术方案的基础之上,所做的任何修改、等同替换、改进等,均应包括在本发明的保护范围之内。

Claims (84)

  1. 一种对图片进行语音标记的方法,包括:
    显示包括目标图片的录音界面;
    响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;
    响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;
    在所述目标图片上添加关联所述音频文件的语音标记。
  2. 根据权利要求1所述的方法,其中,在显示包括目标图片的录音界面之前,所述方法还包括:
    响应于针对目标图片的触发指令,显示语音标记图标;
    其中显示包括目标图片的录音界面,包括:
    响应于对所述语音标记图标的触发指令,跳转至所述录音界面。
  3. 根据权利要求2所述的方法,其中,响应于针对目标图片的触发指令,显示语音标记图标,包括:
    响应于针对所述目标图片的查看指令,显示所述语音标记图标;或,
    响应于针对所述目标图片的编辑指令,显示所述语音标记图标,或,
    响应于用于截取所述目标图片的截屏指令,显示所述语音标记图标;或,
    响应于针对所述目标图片的发送指令,显示所述语音标记图标。
  4. 根据权利要求2所述的方法,其中,在响应于针对目标图片的触发指令,显示语音标记图标之前,所述方法还包括:
    显示针对第一视频的播放界面;
    其中响应于针对目标图片的触发指令,显示语音标记图标,包括:
    响应于针对所述第一视频发出的跳转指令,将跳转显示的视频帧确定为所述目标图片,并在所述播放界面显示所述语音标记图标;或,
    响应于针对所述第一视频发出的暂停指令,将暂停显示的视频帧确定为所述目标图片,并在所述播放界面显示所述语音标记图标。
  5. 根据权利要求1所述的方法,其中,响应于基于所述录音界面发出的录音开始指令,持续采集语音信号,包括:
    响应于基于所述录音界面中目标图片的第一位置发出的录音开始指令,持续采集语音信号;
    其中,在所述目标图片上添加关联所述音频文件的语音标记,包括:
    在所述目标图片的第一位置添加所述语音标记。
  6. 根据权利要求1所述的方法,其中,所述录音开始指令对应于:对所述目标图片的长按操作,所述录音结束指令对应于:对所述目标图片的取消按压操作;或,
    所述录音开始指令对应于:对鼠标右键的点击操作,所述录音结束指令对应于:对鼠标右键的再次点击操作;或,
    所述录音开始指令对应于:对所述录音界面中录音开始图标的触发指令,所述录音结束指令对应于:对所述录音界面中录音结束图标的触发指令。
  7. 根据权利要求1所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记,还包括:
    在所述语音标记中显示序号,所述序号基于对所述目标图片添加的在先的语音标记的数量而确定或由用户自定义输入。
  8. 根据权利要求1所述的方法,其中,响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件,还包括:
    对采集到的语音信号进行语音识别,得到识别文本;
    确定所述识别文本对应的主题文本;
    其中在所述目标图片上添加关联所述音频文件的语音标记,还包括:
    在所述语音标记中显示所述主题文本。
  9. 根据权利要求8所述的方法,其中,确定所述识别文本对应的主题文本,包括:
    将所述识别文本输入预先训练的摘要抽取模型中,得到对应的摘要文本,作为所述主题文本;或,
    将所述识别文本输入预先训练的关键词提取模型中,得到对应的关键词,作为所述主题文本。
  10. 根据权利要求1所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记之后,所述方法还包括:
    接收用户基于所述语音标记输入的自定义文本;
    在所述语音标记中显示所述自定义文本。
  11. 根据权利要求1所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记之后,所述方法还包括:
    响应于对所述语音标记的触发指令,播放所述音频文件。
  12. 根据权利要求1所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记之后,所述方法还包括:
    响应于对所述语音标记的移动指令,将所述语音标记移动至位于所述目标图片中的指定位置。
  13. 根据权利要求1所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记之后,所述方法还包括:
    响应于对所述语音标记的删除指令,从所述目标图片中删除所述语音标记。
  14. 根据权利要求1所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记之后,所述方法还包括:
    响应于对所述语音标记的触发指令,显示菜单栏,所述菜单栏中包括语音识别图标;
    响应于对所述语音识别图标的触发指令,对所述音频文件进行语音识别,得到识别文本;
    在所述录音界面中显示所述识别文本。
  15. 一种图片查看方法,包括:
    显示带有语音标记的图片,所述语音标记通过权利要求1所述的方法添加至所述图片中;
    响应于对所述语音标记的触发指令,播放所述语音标记关联的音频文件。
  16. 根据权利要求15所述的方法,其中,显示带有语音标记的图片,包括:
    在聊天窗口中显示所述图片;或,
    在网页中加载所述图片。
  17. 根据权利要求16所述的方法,其中,所述图片中包括目标商品,所述网页为商品详情页。
  18. 根据权利要求15所述的方法,其中,所述语音标记为多个语音标记,其中播放所述语音标记关联的音频文件,包括:
    基于所述多个语音标记被添加至所述图片中的顺序,依次播放所述多个语音标记对应的多个音频文件。
  19. 根据权利要求15所述的方法,其中,所述语音标记为多个语音标记,其中各个语音标记中显示对应的序号;
    其中播放所述语音标记关联的音频文件,包括:
    基于所述序号的顺序,依次播放所述多个语音标记对应的多个音频文件。
  20. 一种对视频进行语音标记的方法,包括:
    显示包括第一视频的录音界面,所述第一视频包括第一视频帧;
    响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;
    响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;
    响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;
    在所述目标视频帧上添加关联所述音频文件的语音标记。
  21. 根据权利要求20所述的方法,其中,响应于基于所述录音界面发出的录音开始指令,持续采集语音信号,包括:
    响应于基于所述录音界面中目标视频帧的第一位置发出的录音开始指令,持续采集语音信号;
    其中,在所述目标视频帧上添加关联所述音频文件的语音标记,包括:
    在所述目标视频帧的第一位置添加所述语音标记。
  22. 根据权利要求20所述的方法,其中,响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧,包括:
    响应于针对所述第一视频发出的跳转指令,将跳转显示的所述第一视频帧确定为所述目标视频帧;或,
    响应于针对所述第一视频发出的暂停指令,将暂停显示的所述第一视频帧确定为所述目标视频帧。
  23. 根据权利要求20所述的方法,其中,在响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件之后,所述方法还包括:
    在所述第一视频的进度条中添加时间点标记,所述时间点标记对应于所述目标视频帧的播放时间点。
  24. 根据权利要求23所述的方法,其中,在所述第一视频的进度条中添加时间点标记之后,所述方法还包括:
    响应于输入控件指示符移动至所述时间点标记,展示所述目标视频帧中所包括的语音标记的数量。
  25. 根据权利要求23所述的方法,其中,响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件,包括:
    对采集到的语音信号进行语音识别,得到识别文本;
    确定所述识别文本对应的主题文本;
    其中在所述第一视频的进度条中添加时间点标记之后,所述方法还包括:
    响应于输入控件指示符移动至所述时间点标记,展示所述主题文本。
  26. 一种视频查看方法,包括:
    显示带有语音标记的视频,所述语音标记通过权利要求20所述的方法添加至所述视频中,所述视频中包括带有第一语音标记的第一视频帧;
    响应于对所述第一语音标记的触发指令,播放所述第一语音标记关联的音频文件。
  27. 根据权利要求26所述的方法,其中,所述第一视频帧中包括含所述第一语音标记在内的若干语音标记,所述视频的进度条在所述第一视频帧的播放时间点显示有时间点标记;
    其中在显示带有语音标记的视频之后,所述方法还包括:
    响应于输入控件指示符移动至所述时间点标记,展示所述若干语音标记的标记数量和/或所述若干语音标记对应的若干主题文本。
  28. 一种对图片进行文本标记的方法,包括:
    显示针对目标图片的编辑界面;
    响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;
    响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;
    在所述目标图片上添加关联所述识别文本的文本标记。
  29. 根据权利要求28所述的方法,其中,响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号,包括:
    响应于基于所述编辑界面中目标图片的第一位置发出的录音开始指令,持续采集语音信号;
    其中,在所述目标图片上添加关联所述识别文本的文本标记,包括:
    在所述目标图片的第一位置添加所述文本标记。
  30. 根据权利要求28所述的方法,其中,在所述目标图片上添加关联所述识别文本的文本标记,还包括:
    在所述文本标记中显示序号,所述序号基于对所述目标图片添加的在先的文本标记的数量而确定。
  31. 根据权利要求28所述的方法,其中,在所述目标图片上添加关联所述识别文本 的文本标记,还包括:
    在所述文本标记中对所述识别文本进行折叠显示;
    其中,在所述目标图片上添加关联所述识别文本的文本标记之后,所述方法还包括:
    响应于对所述文本标记的触发指令,在所述文本标记中对所述识别文本进行展开显示。
  32. 根据权利要求28所述的方法,其中,在得到识别文本之后,所述方法还包括:
    确定所述识别文本对应的主题文本;
    其中,在所述目标图片上添加关联所述识别文本的文本标记,包括:
    在所述文本标记中显示所述主题文本。
  33. 根据权利要求28所述的方法,其中,在所述目标图片上添加关联所述识别文本的文本标记之后,所述方法还包括:
    响应于对所述文本标记的触发指令,展示所述识别文本。
  34. 一种图片查看方法,包括:
    显示带有文本标记的图片,所述文本标记通过权利要求31所述的方法添加至所述图片中;
    响应于对所述文本标记的触发指令,展示所述文本标记关联的识别文本。
  35. 根据权利要求34所述的方法,其中,显示带有文本标记的图片,还包括:
    在所述文本标记中显示序号;或,
    在所述文本标记中显示所述识别文本对应的主题文本。
  36. 根据权利要求34所述的方法,其中,显示带有文本标记的图片,还包括:
    在所述文本标记中对所述识别文本进行折叠显示;
    其中,展示所述文本标记关联的识别文本,包括:
    在所述文本标记中对所述识别文本进行展开显示。
  37. 一种对视频进行文本标记的方法,包括:
    显示包括第一视频的编辑界面,所述第一视频包括第一视频帧;
    响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;
    响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;
    响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;
    在所述目标视频帧上添加关联所述识别文本的文本标记。
  38. 根据权利要求37所述的方法,其中,响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号,包括:
    响应于基于所述编辑界面中目标视频帧的第一位置发出的录音开始指令,持续采集语音信号;
    其中,在所述目标视频帧上添加关联所述识别文本的文本标记,包括:
    在所述目标视频帧的第一位置添加所述文本标记。
  39. 一种视频查看方法,包括:
    显示带有文本标记的视频,所述语音标记通过权利要求37所述的方法添加至所述视频中,所述视频中包括带有第一文本标记的第一视频帧;
    响应于对所述第一文本标记的触发指令,展示所述文本标记关联的识别文本。
  40. 一种对图片进行语音标记的装置,包括:
    显示单元,配置为显示包括目标图片的录音界面;
    采集单元,配置为响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;
    存储单元,配置为响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;
    添加单元,配置为在所述目标图片上添加关联所述音频文件的语音标记。
  41. 一种图片查看装置,包括:
    显示单元,配置为显示带有语音标记的图片,所述语音标记通过权利要求40所述的装置添加至所述图片中;
    播放单元,配置为响应于对所述语音标记的触发指令,播放所述语音标记关联的音频文件。
  42. 一种对视频进行语音标记的装置,包括:
    显示单元,配置为显示包括第一视频的录音界面,所述第一视频包括第一视频帧;
    确定单元,配置为响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;
    采集单元,配置为响应于基于所述录音界面发出的录音开始指令,持续采集语音信号;
    存储单元,配置为响应于基于所述录音界面发出的录音结束指令,将采集到的语音信号存储为音频文件;
    添加单元,配置为在所述目标视频帧上添加关联所述音频文件的语音标记。
  43. 一种视频查看装置,包括:
    显示单元,配置为显示带有语音标记的视频,所述语音标记通过权利要求42所述的装置添加至所述视频中,所述视频中包括带有第一语音标记的第一视频帧;
    播放单元,配置为响应于对所述第一语音标记的触发指令,播放所述第一语音标记关联的音频文件。
  44. 一种对图片进行文本标记的装置,包括:
    显示单元,配置为显示针对目标图片的编辑界面;
    采集单元,配置为响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;
    识别单元,配置为响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;
    添加单元,配置为在所述目标图片上添加关联所述识别文本的文本标记。
  45. 一种图片查看装置,包括:
    显示单元,配置为显示带有文本标记的图片,所述文本标记通过权利要求44所述的装置添加至所述图片中;
    展示单元,配置为响应于对所述文本标记的触发指令,展示所述文本标记关联的识别文本。
  46. 一种对视频进行文本标记的装置,包括:
    显示单元,配置为显示包括第一视频的编辑界面,所述第一视频包括第一视频帧;
    确定单元,配置为响应于对所述第一视频帧的选取指令,将所述第一视频帧确定为目标视频帧;
    采集单元,配置为响应于基于所述编辑界面发出的录音开始指令,持续采集语音信号;
    识别单元,配置为响应于基于所述编辑界面发出的录音结束指令,对采集到的语音信号进行语音识别,得到识别文本;
    添加单元,配置为在所述目标视频帧上添加关联所述识别文本的文本标记。
  47. 一种视频查看装置,包括:
    显示单元,配置为显示带有文本标记的视频,所述语音标记通过权利要求46所述的装置添加至所述视频中,所述视频中包括带有第一文本标记的第一视频帧;
    展示单元,配置为响应于对所述第一文本标记的触发指令,展示所述文本标记关联的识别文本。
  48. 一种图片处理方法,包括:
    显示聊天界面,并接收基于所述聊天界面选取的待发送的目标图片;
    响应于针对所述目标图片的编辑指令,进入图片编辑界面,其中显示语音标记图标;
    响应于对所述语音标记图标的触发指令,进入录音界面;
    将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
  49. 根据权利要求48所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记,包括:
    在所述语音标记中显示当前编辑用户的聊天昵称和/或聊天头像。
  50. 一种图片处理方法,包括:
    显示聊天界面,所述聊天界面的聊天窗口中包含目标图片;
    响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;
    响应于对所述语音标记图标的触发指令,进入录音界面;
    将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
  51. 根据权利要求50所述的方法,其中,在显示聊天界面之前,所述方法还包括:
    从所述聊天窗口对应的联系人接收所述目标图片,所述目标图片中包括已有语音标记,其中显示所述联系人的聊天昵称和/或聊天头像;
    其中,在所述目标图片上添加关联所述音频文件的语音标记,包括:
    在所述语音标记中显示当前编辑用户的聊天昵称和/或聊天头像。
  52. 根据权利要求50所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记之后,所述方法还包括:
    响应于针对所述录音界面的退出指令,在所述聊天窗口将原始的目标图片更新显示为带所述语音标记的目标图片。
  53. 根据权利要求50所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记之后,所述方法还包括:
    在所述聊天窗口中显示提示信息,以告知各方联系人所述目标图片已发生修改。
  54. 一种图片处理方法,包括:
    显示聊天界面,所述聊天界面的聊天窗口中包含带有第一语音标记的目标图片,所述第一语音标记由所述聊天窗口对应的当前联系人添加;
    响应于对所述第一语音标记的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;
    响应于对所述语音回复图标的触发指令,进入录音界面;
    将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
  55. 根据权利要求54所述的方法,其中,所述方法还包括:
    响应于针对所述录音界面的退出指令,在所述聊天窗口中显示提示信息,以告知各方联系人所述目标图片中的语音标记已发生修改。
  56. 一种图片处理方法,所述方法包括:
    显示包含目标图片的图片编辑界面,所述图片编辑界面的功能菜单中包括语音标记图标;
    响应于对所述语音标记图标的触发指令,进入语音标记界面;
    将基于所述语音标记界面接收的输入文本转化为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
  57. 一种图片处理方法,包括:
    显示包含目标图片的聊天界面;
    响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音添加表情图标;
    响应于对所述语音添加表情图标的触发指令,进入语音添加表情界面;
    将基于所述语音添加表情界面采集的语音信号转化为文字,并基于所述文字生成动画表情;
    在所述目标图片中添加所述动画表情。
  58. 根据权利要求57所述的方法,其中,基于所述文字生成动画表情,包括:
    从表情库中检索与所述文字相关的原始表情;
    在所述原始表情中添加所述文字,得到所述动画表情。
  59. 一种图片处理方法,所述方法的执行主体为电商平台,所述方法包括:
    显示商品信息编辑界面,其中包括针对目标商品的目标图片;
    响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;
    响应于对所述语音标记图标的触发指令,进入录音界面;
    将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
  60. 根据权利要求59所述的方法,其中,在所述目标图片上添加关联所述音频文件的语音标记之后,所述方法还包括:
    在针对所述目标商品的商品详情页中,显示所述目标图片。
  61. 一种图片处理方法,包括:
    显示针对第一订单的订单评价界面,其中包括添加图片图标;
    响应于针对所述添加图片图标的触发指令,接收选取的目标图片;
    响应于对所述目标图片发出的语音标记指令,进入录音界面;
    将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
  62. 一种图片处理方法,包括:
    显示针对目标商品的商品评价界面,其中包括第一用户评价,所述第一用户评价中包括带第一语音标记的目标图片;
    响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;
    响应于对所述语音回复图标的触发指令,进入录音界面;
    将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
  63. 根据权利要求62所述的方法,其中,所述方法还包括:
    响应于针对所述录音界面的退出指令,在所述商品评价界面中显示针对所述第一用户评价的通知消息,用于通知所述第一用户评价已被回复。
  64. 一种图片处理方法,所述方法的执行主体为客服平台,所述方法包括:
    接收用户发送的会话消息,所述会话消息中包括带有第一语音标记的目标图片;
    获取与所述第一语音标记关联的音频文件,对所述音频文件进行语音识别,得到识别文本;
    将所述识别文本输入预先训练的用户标问预测模型中,输出对应的用户标准问题;
    将与所述用户标准问题对应的问题答案反馈给所述用户。
  65. 根据权利要求64所述的方法,其中,将与所述用户标准问题对应的问题答案反馈给所述用户,包括:
    将所述问题答案转化成答复音频,并在所述目标图片中添加与所述答复音频关联的第二语音标记;
    将新添加所述第二语音标记的目标图片发送给所述用户。
  66. 根据权利要求64所述的方法,其中,将与所述用户标准问题对应的问题答案反馈给所述用户,包括:
    将所述问题答案转化成答复音频;
    在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述答复音频的第二语音标记;或者,将所述答复音频,添加至所述第一语音标记对应的音频文件中;
    向用户发送提示信息,以提示用户所述目标图片中已添加答复内容。
  67. 一种电子文件的处理方法,包括:
    显示针对目标电子文件的文件处理界面,所述文件处理界面的功能菜单栏中包括语音标记图标;
    响应于对所述语音标记图标的触发指令,进入录音界面;
    将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
  68. 根据权利要求67所述的方法,其中,所述目标电子文件为电子合同,所述文件处理界面为合同审批界面。
  69. 根据权利要求67所述的方法,其中,所述目标电子文件的文件格式为word文档、或PDF文档、或excel表格。
  70. 一种图片处理方法,所述方法的执行主体为直播平台,所述方法包括:
    显示商品信息编辑界面,其中包括针对待上架的目标商品的目标图片;
    响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;
    响应于对所述语音标记图标的触发指令,进入录音界面;
    将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
  71. 根据权利要求70所述的方法,其中,在所述目标图片上添加关联所述音频文件 的语音标记之后,所述方法还包括:
    显示直播界面,所述直播界面中包括商品上架图标;
    响应于对所述商品上架图的触发指令,显示待上架商品的图片,其中包括所述目标图片;
    响应于对所述目标图片的选取指令,在所述直播界面的商品展示窗口中展示所述目标图片。
  72. 一种图片处理装置,包括:
    显示单元,配置为显示聊天界面;
    接收单元,配置为收基于所述聊天界面选取的待发送的目标图片;
    第一界面切换单元,配置为响应于针对所述目标图片的编辑指令,进入图片编辑界面,其中显示语音标记图标;
    第二界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;
    标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
  73. 一种图片处理装置,包括:
    界面显示单元,配置为显示聊天界面,所述聊天界面的聊天窗口中包含目标图片;
    菜单栏显示单元,配置为响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;
    界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;
    标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
  74. 一种图片处理装置,包括:
    界面显示单元,配置为显示聊天界面,所述聊天界面的聊天窗口中包含带有第一语音标记的目标图片,所述第一语音标记由所述聊天窗口对应的当前联系人添加;
    菜单栏显示单元,配置为响应于对所述第一语音标记的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;
    界面切换单元,配置为响应于对所述语音回复图标的触发指令,进入录音界面;
    标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件 中。
  75. 一种图片处理装置,所述装置包括:
    显示单元,配置为显示包含目标图片的图片编辑界面,所述图片编辑界面的功能菜单中包括语音标记图标;
    界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入语音标记界面;
    标记单元,配置为将基于所述语音标记界面接收的输入文本转化为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
  76. 一种图片处理装置,包括:
    界面显示单元,配置为显示包含目标图片的聊天界面;
    菜单栏显示单元,配置为响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音添加表情图标;
    界面切换单元,配置为响应于对所述语音添加表情图标的触发指令,进入语音添加表情界面;
    表情生成单元,配置为将基于所述语音添加表情界面采集的语音信号转化为文字,并基于所述文字生成动画表情;
    表情添加单元,配置为在所述目标图片中添加所述动画表情。
  77. 一种图片处理装置,所述装置集成于电商平台,所述装置包括:
    界面显示单元,配置为显示商品信息编辑界面,其中包括针对目标商品的目标图片;
    菜单栏显示单元,配置为响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;
    界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;
    标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片上添加关联所述音频文件的语音标记。
  78. 一种图片处理装置,包括:
    显示单元,配置为显示针对第一订单的订单评价界面,其中包括添加图片图标;
    接收单元,配置为响应于针对所述添加图片图标的触发指令,接收选取的目标图片;
    界面切换单元,配置为响应于对所述目标图片发出的语音标记指令,进入录音界面;
    标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
  79. 一种图片处理装置,包括:
    界面显示单元,配置为显示针对目标商品的商品评价界面,其中包括第一用户评价,所述第一用户评价中包括带第一语音标记的目标图片;
    菜单栏显示单元,配置为响应于针对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音回复图标;
    界面切换单元,配置为响应于对所述语音回复图标的触发指令,进入录音界面;
    标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标图片中邻近所述第一语音标记的区域内,添加关联所述音频文件的第二语音标记;或者,将基于所述录音界面采集的语音信号,添加至所述第一语音标记对应的音频文件中。
  80. 一种图片处理装置,所述装置集成于客服平台,所述装置包括:
    接收单元,配置为接收用户发送的会话消息,所述会话消息中包括带有第一语音标记的目标图片;
    获取单元,配置为获取与所述第一语音标记关联的音频文件,对所述音频文件进行语音识别,得到识别文本;
    预测单元,配置为将所述识别文本输入预先训练的用户标问预测模型中,输出对应的用户标准问题;
    反馈单元,配置为将与所述用户标准问题对应的问题答案反馈给所述用户。
  81. 一种电子文件的处理装置,包括:
    显示单元,配置为显示针对目标电子文件的文件处理界面,所述文件处理界面的功能菜单栏中包括语音标记图标;
    界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;
    标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述目标电子文件上添加关联所述音频文件的语音标记。
  82. 一种图片处理装置,所述装置集成于直播平台,所述装置包括:
    界面显示单元,配置为显示商品信息编辑界面,其中包括针对待上架的目标商品的目标图片;
    菜单栏显示单元,配置为响应于对所述目标图片的触发指令,显示菜单栏,所述菜单栏中包括语音标记图标;
    界面切换单元,配置为响应于对所述语音标记图标的触发指令,进入录音界面;
    标记单元,配置为将基于所述录音界面采集的语音信号存储为音频文件,并在所述 目标图片上添加关联所述音频文件的语音标记。
  83. 一种计算机可读存储介质,其上存储有计算机程序,其中,当所述计算机程序在计算机中执行时,令计算机执行权利要求1-39、48-71中任一项的所述的方法。
  84. 一种计算设备,包括存储器和处理器,其中,所述存储器中存储有可执行代码,所述处理器执行所述可执行代码时,实现权利要求1-39、48-71中任一项所述的方法。
PCT/CN2021/080145 2020-03-11 2021-03-11 对图片、视频进行语音标记的方法及装置 Ceased WO2021180155A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202010167913.8A CN113392272A (zh) 2020-03-11 2020-03-11 对图片、视频进行语音标记的方法及装置
CN202010167913.8 2020-03-11

Publications (1)

Publication Number Publication Date
WO2021180155A1 true WO2021180155A1 (zh) 2021-09-16

Family

ID=77615418

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/080145 Ceased WO2021180155A1 (zh) 2020-03-11 2021-03-11 对图片、视频进行语音标记的方法及装置

Country Status (2)

Country Link
CN (1) CN113392272A (zh)
WO (1) WO2021180155A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113886636A (zh) * 2021-09-27 2022-01-04 统信软件技术有限公司 一种影像标记方法、影像标记展示方法及移动终端
WO2024051612A1 (zh) * 2022-09-09 2024-03-14 抖音视界有限公司 视频内容预览交互方法、装置、电子设备及存储介质

Families Citing this family (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113988023A (zh) * 2021-10-27 2022-01-28 展讯通信(天津)有限公司 一种多媒体文件的记事方法、装置、存储介质和终端设备
TWI820677B (zh) * 2022-04-18 2023-11-01 開曼群島商粉迷科技股份有限公司 提供適地性內容連結圖像的方法、系統與電腦可讀取記錄媒體
CN114979054B (zh) * 2022-05-13 2024-06-18 维沃移动通信有限公司 视频生成方法、装置、电子设备及可读存储介质
CN115102917A (zh) * 2022-06-28 2022-09-23 维沃移动通信有限公司 消息发送方法、消息处理方法及装置
CN120751205A (zh) * 2025-06-17 2025-10-03 北京达佳互联信息技术有限公司 多媒体资源的标记方法、多媒体资源的显示方法及装置

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2014186662A1 (en) * 2013-05-17 2014-11-20 Aereo, Inc. Method and system for displaying speech to text converted audio with streaming video content data
CN104333817A (zh) * 2014-11-07 2015-02-04 重庆晋才富熙科技有限公司 一种进行快速视频标记的方法
CN104469544A (zh) * 2014-11-07 2015-03-25 重庆晋才富熙科技有限公司 一种基于语音技术的视频标记方法
CN108092873A (zh) * 2017-10-27 2018-05-29 颜厥护 一种即时通讯方法及系统
CN110215707A (zh) * 2019-07-12 2019-09-10 网易(杭州)网络有限公司 游戏中语音交互的方法及装置、电子设备、存储介质
CN110381382A (zh) * 2019-07-23 2019-10-25 腾讯科技(深圳)有限公司 视频笔记生成方法、装置、存储介质和计算机设备

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103530320A (zh) * 2013-09-18 2014-01-22 中兴通讯股份有限公司 多媒体文件处理方法、装置及终端
CN106250361A (zh) * 2016-08-02 2016-12-21 乐视控股(北京)有限公司 基于文字编辑的数据处理方法和装置
CN109189365A (zh) * 2018-08-17 2019-01-11 平安普惠企业管理有限公司 一种语音识别方法、存储介质和终端设备

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2014186662A1 (en) * 2013-05-17 2014-11-20 Aereo, Inc. Method and system for displaying speech to text converted audio with streaming video content data
CN104333817A (zh) * 2014-11-07 2015-02-04 重庆晋才富熙科技有限公司 一种进行快速视频标记的方法
CN104469544A (zh) * 2014-11-07 2015-03-25 重庆晋才富熙科技有限公司 一种基于语音技术的视频标记方法
CN108092873A (zh) * 2017-10-27 2018-05-29 颜厥护 一种即时通讯方法及系统
CN110215707A (zh) * 2019-07-12 2019-09-10 网易(杭州)网络有限公司 游戏中语音交互的方法及装置、电子设备、存储介质
CN110381382A (zh) * 2019-07-23 2019-10-25 腾讯科技(深圳)有限公司 视频笔记生成方法、装置、存储介质和计算机设备

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113886636A (zh) * 2021-09-27 2022-01-04 统信软件技术有限公司 一种影像标记方法、影像标记展示方法及移动终端
WO2024051612A1 (zh) * 2022-09-09 2024-03-14 抖音视界有限公司 视频内容预览交互方法、装置、电子设备及存储介质
US12542951B2 (en) 2022-09-09 2026-02-03 Douyin Vision Co., Ltd. Video content preview interactive method and apparatus, electronic device, and storage medium

Also Published As

Publication number Publication date
CN113392272A (zh) 2021-09-14

Similar Documents

Publication Publication Date Title
WO2021180155A1 (zh) 对图片、视频进行语音标记的方法及装置
CN102662919B (zh) 对内容片段设置书签
CN105635764B (zh) 视频直播中播放推送信息的方法和装置
CN111008520A (zh) 一种批注方法、装置、终端设备及存储介质
CN103530320A (zh) 多媒体文件处理方法、装置及终端
CN115563320A (zh) 信息回复方法、装置、电子设备、计算机存储介质和产品
CN113536172B (zh) 一种百科信息展示的方法、装置及计算机存储介质
CN113395605B (zh) 视频笔记生成方法及装置
CN111314204A (zh) 一种互动方法、装置、终端和存储介质
CN112329403A (zh) 一种直播文档处理方法和装置
CN119182980B (zh) 处理媒体资源的方法以及装置
JP5475259B2 (ja) テキスト情報共有方法、サーバ装置及びクライアント装置
CN117742538A (zh) 消息显示方法、装置、电子设备和可读存储介质
CN112004031A (zh) 视频生成方法、装置及设备
CN115174506A (zh) 会话信息处理方法、装置、可读存储介质和计算机设备
CN112073738B (zh) 一种信息的处理方法和装置
CN114139525A (zh) 数据处理方法、装置、电子设备及计算机存储介质
CN112783592A (zh) 信息发布方法、装置、设备和存储介质
WO2018169711A1 (en) Systems and methods for multi-user word processing
CN113296855B (zh) 会话内容收藏处理方法、装置、计算机设备和存储介质
CN111625740B (zh) 图像显示方法、图像显示装置和电子设备
CN110275742A (zh) 信息处理方法、信息显示方法、装置、终端及服务器
CN119835487A (zh) 内容发布方法及相关产品
CN116455852B (zh) 消息处理方法、装置、设备及存储介质
CN113821131B (zh) 多媒体信息处理方法、装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21766959

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 21766959

Country of ref document: EP

Kind code of ref document: A1