WO2023025005A1 - 音频数据播放方法与装置 - Google Patents

音频数据播放方法与装置 Download PDF

Info

Publication number
WO2023025005A1
WO2023025005A1 PCT/CN2022/113074 CN2022113074W WO2023025005A1 WO 2023025005 A1 WO2023025005 A1 WO 2023025005A1 CN 2022113074 W CN2022113074 W CN 2022113074W WO 2023025005 A1 WO2023025005 A1 WO 2023025005A1
Authority
WO
WIPO (PCT)
Prior art keywords
audio data
input
objects
picture
playing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/113074
Other languages
English (en)
French (fr)
Inventor
喻超宁
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Vivo Mobile Communication Co Ltd
Original Assignee
Vivo Mobile Communication Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Vivo Mobile Communication Co Ltd filed Critical Vivo Mobile Communication Co Ltd
Publication of WO2023025005A1 publication Critical patent/WO2023025005A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/16Sound input; Sound output
    • G06F3/165Management of the audio stream, e.g. setting of volume, audio stream path
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/40Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
    • G06F16/43Querying
    • G06F16/432Query formulation
    • G06F16/434Query formulation using image data, e.g. images, photos, pictures taken by a user
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/40Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
    • G06F16/48Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • G06F16/483Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning

Definitions

  • the present application belongs to the technical field of communication, and in particular relates to a method and device for playing audio data.
  • the electronic device can recognize the content of the area clicked by the user, and play the recognition result in the form of audio to inform the user of the content of the clicked area.
  • the purpose of the embodiment of the present application is to provide a method and device for playing audio data, which can solve the problem that the prior art recognizes and plays the partial content of the picture, and it is difficult to convey the overall content of the picture to the user.
  • the embodiment of the present application provides a method for playing audio data, the method comprising:
  • the first input to the first object in the picture is received, wherein the picture includes N second objects, the first object is any second object in the N second objects, and N is greater than an integer of 1;
  • the first audio data associated with each second object is played, wherein the playback parameters of the first audio data associated with each second object are related to the distance from each second object to the first object.
  • the embodiment of the present application provides an audio data playback device, the device comprising:
  • the first receiving module is configured to receive a first input to the first object in the picture when the picture is displayed, wherein the picture includes N second objects, and the first object is any one of the N second objects
  • the second object, N is an integer greater than 1;
  • the first playback module is used to play the first audio data associated with each second object in response to the first input, wherein the playback parameters of the first audio data associated with each second object are related to each second object to the first audio data The distance of an object is related.
  • the embodiment of the present application provides an electronic device, the electronic device includes a processor, a memory, and a program or instruction stored in the memory and operable on the processor.
  • the program or instruction is executed by the processor, the The steps of the method of the first aspect.
  • the embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method in the first aspect are implemented.
  • the embodiment of the present application provides a chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the method in the first aspect.
  • an embodiment of the present application provides a computer program product, including a computer program tangibly contained on a computer-readable medium, where the computer program includes program code for executing the method as described in the first aspect.
  • the embodiment of the present application provides an electronic device configured to execute the steps of the method described in the first aspect.
  • the audio data playback method in the case of displaying a picture, receives a first input to a first object in the picture, the picture includes a plurality of second objects, and the first object is the plurality of second objects Any second object in; in response to the first input, play the first audio data associated with each second object, the playback parameters of the first audio data associated with each second object are the same as each second object to the first The distance of the object is related.
  • the embodiment of the present application helps to convey the overall content of the picture to the user and improve user experience.
  • Fig. 1 is a schematic flow chart of the audio data playback method provided by the embodiment of the present application.
  • Fig. 2 is an example figure of the picture in the embodiment of the present application.
  • Fig. 3 is a schematic flow chart of an audio data playback method in a specific application example
  • Fig. 4 is a schematic structural diagram of an audio data playback device provided by an embodiment of the present application.
  • FIG. 5 is a schematic structural diagram of an electronic device provided in an embodiment of the present application.
  • FIG. 6 is a schematic diagram of a hardware structure of an electronic device provided by an embodiment of the present application.
  • the audio data playback method includes:
  • Step 101 in the case of displaying a picture, receiving a first input of a first object in the picture, wherein the picture includes N second objects, and the first object is any second object in the N second objects, N is an integer greater than 1;
  • Step 102 in response to the first input, play the first audio data associated with each second object, wherein the playback parameters of the first audio data associated with each second object and the distance from each second object to the first object relevant.
  • the audio data playing method provided in the embodiment of the present application may be applied to electronic devices.
  • the electronic device may be a mobile terminal or a personal computer, etc., which is not specifically limited here.
  • the electronic device can display a picture, and the picture can include multiple second objects.
  • the above picture may be a landscape picture taken at the seaside, and correspondingly, the picture may include second objects such as the sea, trees, and birds.
  • the picture may be a picture taken at a station, and correspondingly, the picture may include second objects such as person A, person B, and a vehicle.
  • the second object in the picture can be obtained by recognizing the picture.
  • each second object can be recognized from the picture based on the deep learning model in advance. In some other examples, the second objects may also be identified from the pictures in advance through manual identification.
  • the picture can be sent to the server in advance, and the server uses a deep learning model to identify the picture, and can send the relevant identification result to the electronic device.
  • These recognition results may include the aforementioned N second objects.
  • the electronic device may directly use a deep learning model to recognize the picture to obtain N second objects.
  • the electronic device may receive a first input for a first object, that is, a first input for any second object.
  • the first input may correspond to an input in the form of a single click, a multi-click, or a long press, which is not specifically limited here.
  • the position of each second object in the picture can be obtained at the same time, specifically, the position of the image area corresponding to each second object can be obtained in the picture.
  • the electronic device may acquire the input position of the first input in the picture, and based on the input position and the position of each second object in the picture, the electronic device determines the above-mentioned first input from the picture. an object.
  • the electronic device may also recognize the input position and the image area within a preset distance range in real time according to the input position of the first input in the picture, and then determine the first object.
  • the electronic device may play the first audio data associated with each second object in response to the first input.
  • each second object can be identified by the deep learning model.
  • each type of second object may be associated with an identifier, and the identifier may reflect the classification and recognition results of the second object to a certain extent.
  • an identifier "person" may be associated, and the identifier may be expressed in a textual manner.
  • the above-mentioned identification may also be expressed by means of numbering or the like.
  • Each type of second object can be associated with corresponding first audio data.
  • the association relationship between the second object and the first audio data can be reflected in the association relationship between the second object and the first audio data middle.
  • the electronic device can query the first audio data associated with each second object in a preset audio database according to the identifier of each second object, so as to perform the first audio data associated with each second object play.
  • the above-mentioned server queries the first audio data associated with each second object from the audio data block after completing the identification of the second object, and sends each first audio data and the association relationship of the second audio data object is sent to the electronic device for playing by the electronic device.
  • the following mainly takes the sound content emitted when the first audio data is played as an example to describe the first audio data instead.
  • the first audio data may be simple descriptors of the associated second object.
  • the associated first audio data may be "person"; for a second object such as a dog, the associated first audio data may be "dog”; for a second object such as the sea Object, the associated first audio data may be "sea”.
  • the first audio data may also be the sound often made by the associated second object in an actual environment.
  • the associated first audio data may be "hello"; for a second object such as a puppy, the associated first audio data may be "woof woof”;
  • the associated first audio data may be " ⁇ " (the sound of ocean waves).
  • the first audio data associated with each second object in the N second objects can be played, and at the same time, the first audio data associated with each second object
  • the playback parameter may be related to the distance between each second object and the first object.
  • the playback parameters of the first audio data may include playback volume
  • the playback volume of the first audio data associated with each second object may be negatively correlated with the distance between each second object and the first object.
  • the playback parameters of the first audio data may include playback speed, and the playback frequency of the first audio data associated with each second object may be negatively correlated with the distance between each second object and the first object.
  • the above are some examples of the types of playback parameters and the correlation between the playback parameters and the distance.
  • the types of playback parameters and the correlation can be set as required.
  • the playback parameters including the playback volume are mainly used as an example for description below.
  • the distance between it and the first object may be the distance in the image coordinate system, or the distance in a coordinate system such as the earth coordinate system.
  • the coordinates of any second object can be the coordinates of the midpoint of the image region corresponding to the second object in the image coordinate system; the distance between the first object and the second object can be The calculation is performed by the coordinates of the first object and the coordinates of the second object.
  • a straight line may be recognized from the picture, and the straight line may be used as a reference line to determine the distance between the first object and the second object.
  • straight lines can usually be the boundary of roads, the boundary between the water surface and the sky, and so on. Using the straight line as the reference line, the approximate distance between the first object and the second object in the earth coordinate system can be obtained.
  • the first audio data associated with the first object can be played at a higher volume, for example, at 50% of the maximum volume of the electronic device.
  • the N second objects include the first object, the second object A, and the second object B, wherein the distance between the second object A and the first object is smaller than the distance between the second object B and the first object. Then the first audio data associated with the second object A can be played at a maximum volume of 40%, and the first audio data associated with the second object B can be played at a maximum volume of 30%.
  • the playback parameters of the first audio data associated with each second object include a playback speed
  • the relative position of each second object may also be reflected by the playback speed
  • the audio data playback method in the case of displaying a picture, receives a first input to a first object in the picture, the picture includes a plurality of second objects, and the first object is the plurality of second objects Any second object in; in response to the first input, play the first audio data associated with each second object, the playback parameters of the first audio data associated with each second object are the same as each second object to the first The distance of the object is related.
  • the embodiments of the present application help to convey the overall content of the picture to the user and improve user experience.
  • the user can understand the content of the picture more accurately by performing the first input on the picture. Therefore, the audio data playback method provided by the embodiment of the present application can also effectively improve the convenience of operation.
  • the electronic device in response to the first input, may first play the first audio data associated with the first object, and then play the first audio data associated with the second object other than the first object after a preset time interval. audio data.
  • the playback order of the first audio data associated with the second objects other than the first object may also be related to the distance from each second object to the first object, the farther the second object is from the first object, The playing sequence of the associated first audio data is about later.
  • the playing tone color of the first audio data associated with each second object matches the color of the image area corresponding to each second object.
  • the image area of each second object in the picture can also be determined.
  • the color of the pixels in each image area is generally known, and the color of the image area can be determined according to the color of each pixel in the image area.
  • the color of a pixel can be represented by an RGB value or a grayscale value.
  • RGB value a value represented by an RGB value
  • the following mainly takes the color of a pixel represented by an RGB value as an example for illustration.
  • the second object may include multiple pixels in the image area in the picture, and the color of the image area may correspond to the mode or average of the RGB values of the pixels in the image area.
  • the playing tone may match the coolness or warmth of the color or the color system.
  • the playback timbre of the first audio data associated with each second object may be determined according to the preset matching relationship.
  • the warmth or coolness of a color can be calculated by the following formula:
  • CW is the value used to measure the degree of warmth of the color
  • r is the value of the red channel in the RGB value
  • g is the value of the green channel in the RGB value
  • b is the value of the blue channel in the RGB value.
  • the value ranges of r, g and b are usually 0-255.
  • the CW when the CW is greater than or equal to 192, it may be considered a cool color, and when the CW is less than 192, it may be considered a warm color.
  • cool colors you can match duller timbres; for warmer colors, you can match lighter timbres.
  • the value range of CW (0 ⁇ 253) can be divided into multiple intervals according to the preset step size, and the timbre can be matched according to the trend of lightness to dullness for each interval according to the order of CW from small to large .
  • the color of the second object of the sea is blue, the corresponding CW is lower, and the matching tone is lighter. Therefore, when playing the first audio data corresponding to the sea , which can be played back with a lighter tone.
  • the color of the second object, the sea is close to black or dark green, the corresponding CW is relatively high, and the matching tone is relatively dull.
  • the matching relationship between the timbre and the color can also be reflected in the matching relationship between the timbre and the color system to which the color belongs, that is, for different color systems, different timbres can be matched.
  • the process of determining the playing tone of the first audio data of each second object according to the color system of the color of the image area corresponding to each second object will not be described in detail here.
  • the playing tone of the first audio data associated with each second object may be determined by the server after identifying each second object in the picture according to the color of the image area corresponding to each second object.
  • the subsequent server may send the association relationship between the playing tone and the second object to the electronic device, so that the electronic device selects an appropriate playing tone to play the first audio data associated with each second object.
  • the electronic device when it receives the first input, it can also recognize the picture to obtain N second objects, and according to the color of the image area corresponding to each second object, and the preset playing tone and color The matching relationship between each second object is determined to determine the playing timbre corresponding to each second object, and the first audio data associated with the second object is played with the determined playing timbre.
  • the timing of determining the playing timbre there is no specific limitation on the timing of determining the playing timbre, as long as the playing timbre matches the color of the image area corresponding to each second object when playing the first audio data associated with each second object.
  • the playing tone color matches the color of the image area corresponding to the second object, which helps to convey the content of the picture to the user more effectively, and improves the user's listening experience of the picture content.
  • the corresponding image area in the picture may be divided into multiple sub-image areas by other second objects.
  • sub-image areas with the same or similar colors can be merged into image areas corresponding to the same second object, and then according to the degree of warmness or color of the merged image areas, system to determine the matching playback tone.
  • the audio data playing method may further include:
  • the object type of the first object is a preset object type
  • the second audio data includes at least one of the following:
  • the first object may be any one of the N second objects, and the determination of the first object may be related to the user's input. For example, if the user's first input is a click input on the second object of the sea, the clicked sea can be determined as the first object; if the user's second input is a long click on the second object of the person Press enter, then the first object can be updated to the long-pressed character.
  • the specific manner of the second input at this time may not be limited to the long press input, but may also be other preset gesture input and the like.
  • the second input may correspond to gesture input of drawing a question mark or ticking.
  • the electronic device may play the second audio data associated with the first object in response to the second input.
  • the preset object type may be a character
  • the electronic device may play the second audio data associated with the first object in the voice of the character.
  • the preset object type can also be a parrot, a trumpet or other object types, which can be set according to actual needs.
  • the preset object type may correspond to an object that can issue prompting language in an actual environment.
  • the preset object type may also be an object type such as an animal or a plant, and the corresponding object may issue a prompting language in an anthropomorphic manner.
  • the following description will mainly take the preset object type as a person as an example.
  • the second audio data may include audio data for prompting a state of at least one second object.
  • the second audio data may correspond to audio data for an introduction of the state.
  • the second audio data may be "I am character A reading a newspaper", "behind me is the blue sea” and so on.
  • the state of the second object may refer to the behavior state, color state, etc. of the second object, which may be set according to actual needs, and may not be specifically limited here.
  • the distance between at least one third object and itself may be introduced in a first-person tone.
  • the second audio data may be correspondingly used to prompt the distance between the at least one third object and the first object.
  • the second audio data may be "person B is about two finger widths to my right", "a tree is about one finger width to my left", and so on.
  • the second audio data can also be used to prompt the aforementioned status and distance at the same time.
  • the second audio data may be "person B is at a distance of about two finger widths to my left, and he is making a phone call”.
  • the second input to the first object is received, and when the object type of the first object is a preset object type, in response to the second input, the second audio data associated with the first object is played, It is helpful to convey information such as the status or distance of each second object in the picture to the user, so that the user can better understand the content expressed in the picture.
  • the method further includes:
  • the second audio data associated with the updated first object is played.
  • the picture may include two second objects, character A and character B, and the user may click (corresponding to the first input) the image area corresponding to character A.
  • the electronic device may play the first object associated with character A. Audio data, such as "person” or "hello".
  • the electronic device can play the second audio data associated with character A, for example, "A distance of about two finger widths on my right side Here is Person B".
  • the electronic device may re-determine the first object according to the end position of the third input.
  • the first object can be updated as character B, and at this time, the second audio data associated with character B can be played, for example, "I am character B, I am on the phone, and about a finger's width to my right is Person C".
  • the specific manner of the third input may also be long-press input, etc., which may not be specifically limited here.
  • the user can determine the image area corresponding to the character B according to the prompt of the second audio data associated with the character A, and then perform a long press input on the image area corresponding to the character B.
  • the first object is updated according to the end position of the third input, and the second audio data associated with the updated first object is played, which helps to guide the user to obtain more detailed Information such as the status or position of each second object in the picture facilitates the user's understanding of the overall content of the picture and improves user experience.
  • the user may swipe a finger from the image area where character A is located to the image area where character B is located.
  • the second audio data associated with character A can be kept playing, and as the distance between the user's finger and the image area of character A increases, the second audio data The playback volume of data can be reduced.
  • the first object can be updated from character A to character B, and then the second audio data associated with character B is played, and, as the user's finger reaches the image of character B As the distance of the area decreases, the playback volume of the second audio data can be increased.
  • the content and playback parameters (such as playback volume and playback speed) of the second audio data can be determined according to the real-time input position corresponding to the third input, so that the user can determine the input position and playback speed of the third input.
  • the distance relationship between each second object is acquired in real time, so as to better guide the user to acquire the content expressed in the picture.
  • the electronic device may count P second objects associated with all the second audio data played during the third input period, and the number Q of second objects whose object type is a preset object type in the picture.
  • Q is a positive integer
  • P is a positive integer less than or equal to Q.
  • the electronic device can output the value of Q-P.
  • the value of Q-P can be regarded as the second object without introduction of information such as status or distance.
  • the value of Q-P it can also be output in the form of audio playback. In this way, the user can fully understand the content shown in the picture.
  • the audio data playing method may further include at least one of the following:
  • the playback volume of the first audio data is adjusted according to the input parameter of the seventh input.
  • the first object may be any one of the N second objects, and the determination of the first object may be related to the user's input.
  • the first object may be the same second object or different second objects among the N second objects.
  • the pictures may include person A, person B, the sea, and grass.
  • the fourth input may be a double-click or multi-click input on the first object.
  • the user can double-click the sea in the picture, and the electronic device keeps playing the first audio data associated with the sea in response to the user's double-click input on the first object of the sea.
  • the first audio data associated with the sea can also include the sound of the sea breeze "whirring” and the calls of seabirds, and keep track of these first audio data. The data can be played.
  • the electronic device may stop playing the first audio data associated with the character A, the character B, and the grass respectively.
  • the electronic device may play the first audio data associated with each second object in response to the first input.
  • the first audio data associated with character A may include "Hello", “I am a character” and “How can I help you", and these first audio data can be played in turn or randomly at preset intervals.
  • the electronic device no longer plays the first audio data associated with character A.
  • the fourth input is a double-click input to character A or character B
  • the first audio data associated with character A and the first audio data associated with character B can be kept playing, and the playback is stopped.
  • the fifth input may be a pinch input.
  • the pinch input may specifically be a gesture input in which at least three fingers move closer to each other.
  • the electronic device When the electronic device detects the pinch input, it can determine the target image area matching the fifth input from the picture according to the end positions of at least three fingers.
  • the target image area may be an image area enclosed by a line connecting the contacts at the end positions of the three fingers.
  • the image area corresponding to each second object When judging whether the image area corresponding to each second object is within the target image area, it may be to determine whether the midpoint of the image area corresponding to each second object is within the target image area; Whether the image area corresponding to the object is wholly or partly within the target image area can be set according to actual needs.
  • the subsequent electronic device may stop playing the first audio data associated with the fourth object.
  • the electronic device can play the associated audio data for the second object that the user pays more attention to according to the input situation of the user.
  • the fifth input may also be other types of gesture input.
  • the fifth input may be an input along a closed track, and the target image area may be an area enclosed by the corresponding closed track.
  • the sixth input may be a back and forth sliding input.
  • the electronic device can make a "rustling" sound to imitate the sound of the grass being moved.
  • the associated first audio data can be the sound of "rustle” with a slower frequency; and the associated third audio data can be the sound of "rustle” with a faster frequency. sound.
  • the associated first audio data may be "Hello”
  • the associated third audio data may be "What's the matter?"
  • the user's sixth input to any second object may be regarded as an action interaction with the second object.
  • it when receiving the sixth input to the first object, it may respond to the sixth input by playing a preset sound that reflects the first object being interacted with, that is, playing the above-mentioned The third audio data associated with the first object, so that the user can obtain a better interactive experience.
  • the user can draw a circle on the electronic device with a small range, so as to adjust the playing volume of the first audio data.
  • the input circled with a smaller amplitude may be regarded as corresponding to the seventh input.
  • the determination of the smaller range may be based on the size of the area drawn by the user. For example, when the circled area is smaller than the preset area, it may be considered that the circle is drawn with a smaller range.
  • the input for drawing a circle may have corresponding input parameters, such as the direction and number of circles drawn.
  • the playback volume of each first audio data when the direction of the circle is clockwise, the playback volume of each first audio data can be increased; when the direction of the circle is counterclockwise, the playback volume of each first audio data can be decreased. And the degree of turning up or down of the above-mentioned playback volume can be determined by the number of turns.
  • the relative size relationship between the playback volumes of each first audio data can remain unchanged, that is to say, each second object association still exists
  • the playback volume of the first audio data is negatively correlated with the distance from each second object to the first object.
  • the electronic device can implement different audio data playback functions according to different gesture inputs of the user, greatly improving the user's operation convenience.
  • the audio data playing method may further include:
  • the distance between any two second objects is determined according to the positional relationship between the image area corresponding to each second object and the straight line.
  • a background image area in the picture such as the image area where the sky or the earth is located.
  • the image area where the second object identified as the sky or the earth is located may be directly determined as the background image area.
  • the image area where the sky or the earth is located may be divided into multiple sub-image areas, and the colors of these sub-image areas may be the same or similar, so these sub-image areas can be classified into The background image area.
  • the picture can be taken by a camera, and correspondingly, the content in the picture can be presented in the form of a perspective view. That is to say, the second object in the picture may be presented in the form of a near-large and far-small form.
  • the picture includes second objects such as the earth D1, the road D2, the tree D3, the sky D4, the person D5, and the vehicle D6.
  • the road D2 converges into a point TP at the far end, and the point TP can be a straight line The point of intersection between L1, the straight line L2 and the straight line L3.
  • the straight line L1 may be the dividing line between the land D1 and the sky D4; the straight line L2 and the straight line L3 are the dividing lines between the land D1 and the road D2.
  • the straight line L1 , the straight line L2 and the straight line L3 can all be determined from the background image area.
  • the method of determining the straight lines in the background image area can be obtained through techniques such as image segmentation or feature extraction. Specifically, it can be realized through existing technologies, and details will not be described here.
  • the distance between any two second objects may be determined according to the positional relationship between the straight line and the image area corresponding to each second object.
  • the midpoint of the image area corresponding to the second object may be used as the position of the second object in the picture.
  • the method of determining the image area corresponding to the second object it has been described in the above embodiment, and will not be repeated here.
  • the distance between the two second objects and the straight line L2 is relatively short, and the distance between the person D5 and the vehicle D6 can be determined using the straight line L2 as a reference.
  • the connecting line between the character D5 and the vehicle D6 can be decomposed into a sub-line segment parallel to the straight line L2 and a sub-line segment perpendicular to the straight line L2. According to the length of the two sub-line segments, the distance between the character D5 and the vehicle D6 can be roughly determined. The distances in the two directions can then be used to obtain the distance between the person D5 and the vehicle D6.
  • the distance between any two second objects in the earth coordinate system can be obtained relatively accurately. distance. Subsequently, when the first audio data associated with each second object is played according to the distance, the distance relationship between each second object and the first object can be conveyed to the user more accurately.
  • the audio data playing method may further include:
  • the server is used to identify the picture, and obtain the coordinates of the image area corresponding to each second object in the N second objects in the picture;
  • playing the first audio data associated with each second object includes:
  • the input parameters of the first input are sent to the server, and the server is used to generate audio playback rules according to the input parameters and the coordinates of the image area corresponding to each second object, and the audio playback rules include each second object The associated first audio data and its playback parameters;
  • the first audio data associated with each second object is played according to an audio playing rule.
  • the identification of the picture and the determination of each audio playback rule can be performed in the server. In this way, the requirements for the hardware configuration of the electronic device can be reduced, and the consumption of computing resources of the electronic device can be reduced.
  • an electronic device when it displays a picture, it can send the picture to the server, and the server can identify the picture, obtain each second object in the picture, and the coordinates of the image area corresponding to each second object in the picture .
  • Each second object may be represented by text or other forms of identification.
  • the server may associate and store the identifiers of each second object and the coordinates in the picture. To simplify the description, it can be considered that the server stores the identifier of each second object and the coordinates in the picture in the first mapping table.
  • the electronic device When the electronic device receives the first input, it can send the input parameters of the first input to the server.
  • the input parameter of the first input may include the position of the image area clicked by the user relative to the picture.
  • the server may determine the second object corresponding to the image area clicked by the user, that is, determine the above-mentioned first object.
  • the server when it establishes audio playback rules, it can mainly perform the following processes:
  • One is to determine the distance between each second object and the first object according to the coordinates of the image area corresponding to each second object in the first mapping table, so as to further determine the audio playback volume corresponding to each second object.
  • the distance here may be negatively correlated with the audio playback volume, that is, the larger the distance, the lower the audio playback volume, and vice versa.
  • the second is to query the first audio associated with each second object from the preset audio database according to the identification of each second object and the corresponding relationship between the object audio data (the corresponding relationship can be considered to be stored in the second mapping table). data.
  • the server may send the correspondence relationship of the second object—the first audio data—audio playback volume to the electronic device as an audio playback rule.
  • the electronic device can play the first audio data associated with each second object according to the audio playing rule.
  • the server can further determine the playing tone color of the first audio data associated with each second object according to the color of the image area corresponding to each second object, and add the playing tone color to the above-mentioned audio playing rule middle.
  • the server can also determine the audio playback speed corresponding to each second object according to the distance between the second object and the first object, and add the playback speed to the above-mentioned audio playback rule.
  • the audio data playing method can be applied to an electronic device, and the electronic device can perform data interaction with a server.
  • Audio data playback methods include:
  • Step 301 the server parses the content in the picture, and extracts the second object in the picture;
  • the picture parsed by the server may be sent to the server by the electronic device.
  • the server may use a deep learning model to analyze the image.
  • Step 302 the electronic device receives the user's first input, and sends the input parameters of the first input to the server;
  • the user can click a certain image area in the picture, and the electronic device can send the position information of the clicked image area relative to the whole picture as an input parameter to the server.
  • Step 303 the server acquires input parameters, determines the first object, and stores the second objects in a preset array in order from farthest to closest relative to the first object;
  • Step 304 according to the order of each second object in the preset array, the server assigns the audio playback volume to each second object according to the rule from small to large;
  • the server may also assign an audio playback speed or other types of playback parameters to the second object.
  • Step 305 calculating the degree CW of the color of the image area of each second object (hereinafter referred to as the CW of the second object);
  • CW is the value used to measure the degree of warmth of the color
  • r is the value of the red channel in the RGB value
  • g is the value of the green channel in the RGB value
  • b is the value of the blue channel in the RGB value.
  • the value ranges of r, g and b are usually 0-255.
  • Step 306 judge whether CW is greater than or equal to 192, if yes, execute step 307, if not, execute step 308;
  • Step 307 according to the difference between 255 and the CW of the second object, determine the dull degree of the audio playback timbre corresponding to the second object, and perform step 309;
  • Step 308 according to the difference between the CW of the second object and 0, determine the crispness of the audio playback tone corresponding to the second object, and execute step 309;
  • Step 309 Play the audio data associated with each second object according to the determined audio playback volume and audio playback tone color for each second object.
  • the audio data playback method provided by the embodiment of the present application can accurately convey the overall content of the picture to the user by determining the audio playback volume and audio playback tone color of each second object in the picture, and satisfy the needs of disabled users. Comprehension needs of picture content.
  • the audio data playback method provided in the embodiment of the present application may be executed by an audio data playback device, or a control module in the audio data playback device for executing the audio data playback method.
  • the audio data playing device provided in the embodiment of the present application is described by taking the audio data playing device executing the audio data playing method as an example.
  • the audio data playback device 400 provided by the embodiment of the present application includes:
  • the first receiving module 401 is configured to receive a first input to the first object in the picture when the picture is displayed, wherein the picture includes N second objects, and the first object is any of the N second objects A second object, N is an integer greater than 1;
  • the first playback module 402 is configured to play the first audio data associated with each second object in response to the first input, wherein the playback parameters of the first audio data associated with each second object are the same as each second object to The distance of the first object is related.
  • the playing tone color of the first audio data associated with each second object matches the color of the image area corresponding to each second object.
  • the audio data playback device 400 may also include:
  • a second receiving module configured to receive a second input to the first object
  • the second playing module is used to play the second audio data associated with the first object in response to the second input when the object type of the first object is a preset object type;
  • the second audio data includes at least one of the following:
  • the audio data playback device 400 may also include:
  • An update module configured to update the first object to a second object that is closest to the termination position of the third input and whose object type is a preset object type in response to the third input when the third input is received;
  • the third playing module is used to play the second audio data associated with the updated first object.
  • the audio data playback device 400 may also include at least one of the following:
  • the first stop playing module is configured to stop playing the first video associated with the second object whose object type is different from that of the first object in response to the fourth input when the fourth input to the first object is received. audio data;
  • the second stop playing module is used to determine the target image area matching the fifth input from the picture in response to the fifth input when receiving the fifth input to the picture, and stop playing the first picture associated with the fourth object.
  • the fourth object is a second object whose corresponding image area is located outside the target image area;
  • the fourth playing module is used to play the third audio data associated with the first object in response to the sixth input when the sixth input to the first object is received;
  • the adjustment module is configured to adjust the playing volume of the first audio data according to the input parameters of the seventh input in response to the seventh input when the seventh input is received.
  • the audio data playback device 400 may also include:
  • a first determination module configured to determine the background image area in the picture and the image area corresponding to each second object
  • the second determination module is used to determine straight lines from the background image area
  • the fourth determination module is configured to determine the distance between any two second objects according to the positional relationship between the image area corresponding to each second object and the straight line.
  • the audio data playback device 400 may also include:
  • the sending module is used to send the picture to the server, and the server is used to identify the picture and obtain the coordinates of the image area corresponding to each second object in the N second objects in the picture;
  • the first playing module 401 may include:
  • the sending unit is configured to send the input parameters of the first input to the server in response to the first input, and the server is used to generate audio playback rules according to the input parameters and the coordinates of the image area corresponding to each second object, and the audio playback rules include Associated first audio data and playback parameters of each second object;
  • a receiving unit configured to receive the audio playback rules sent by the server
  • the playing unit is used to play the first audio data associated with each second object according to the audio playing rule.
  • the audio data playback device receives a first input to the first object in the picture when the picture is displayed, and plays the first audio data associated with each second object in response to the first input,
  • the playback parameters of the first audio data associated with each second object are related to the distance from each second object to the first object. In this way, by processing the playback parameters of the first audio data associated with each second object, it is possible to compare Accurately convey the content of the image to the user.
  • the playback timbre of the first audio data associated with each second object matches the color of the image area corresponding to each second object, which further facilitates the user's understanding of the picture content.
  • the audio data playback device can also adjust the focus of the audio playback in response to relevant input from the user, so as to meet the user's demand for acquiring more concerned picture content and improve the user experience.
  • the audio data playback device in the embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal.
  • the device may be a mobile electronic device or a non-mobile electronic device.
  • the mobile electronic device may be a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle electronic device, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant).
  • non-mobile electronic devices can be servers, network attached storage (Network Attached Storage, NAS), personal computer (personal computer, PC), television (television, TV), teller machine or self-service machine, etc., this application Examples are not specifically limited.
  • Network Attached Storage NAS
  • personal computer personal computer, PC
  • television television
  • teller machine or self-service machine etc.
  • the audio data playback device in the embodiment of the present application may be a device with an operating system.
  • the operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in this embodiment of the present application.
  • the audio data playback device provided by the embodiment of the present application can realize various processes realized by the method embodiments in FIG. 1 to FIG. 3 , and details are not repeated here to avoid repetition.
  • the embodiment of the present application also provides an electronic device 500, including a processor 501, a memory 502, and a program or instruction stored in the memory 502 and operable on the processor 501.
  • the program when the instruction is executed by the processor 501, each process of the audio data playback method embodiment described above can be realized, and the same technical effect can be achieved. To avoid repetition, details are not repeated here.
  • the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
  • FIG. 6 is a schematic diagram of a hardware structure of an electronic device implementing an embodiment of the present application.
  • the electronic device 600 includes but is not limited to: a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609, and a processor 610, etc. part.
  • the electronic device 600 can also include a power supply (such as a battery) for supplying power to various components, and the power supply can be logically connected to the processor 610 through the power management system, so that the management of charging, discharging, and function can be realized through the power management system. Consumption management and other functions.
  • a power supply such as a battery
  • the structure of the electronic device shown in FIG. 6 does not constitute a limitation to the electronic device, and the electronic device may include more or fewer components than shown in the figure, or combine some components, or arrange different components, which will not be repeated here. .
  • the user input unit 607 is configured to receive a first input to the first object in the picture when the picture is displayed, wherein the picture includes N second objects, and the first object is one of the N second objects. Any second object, N is an integer greater than 1;
  • the audio output unit 603 is configured to play the first audio data associated with each second object, wherein the playing parameters of the first audio data associated with each second object are related to the distance from each second object to the first object.
  • the electronic device when displaying a picture, receives a first input to a first object in the picture, the picture includes a plurality of second objects, and the first object is one of the plurality of second objects Any second object; in response to the first input, play the first audio data associated with each second object, and the playback parameters of the first audio data associated with each second object are related to each second object to the first object Distance is relevant.
  • the embodiments of the present application help to convey the overall content of the picture to the user and improve user experience.
  • the playing timbre of the first audio data associated with each second object matches the color of the image area corresponding to each second object.
  • the user input unit 607 is further configured to receive a second input to the first object
  • the audio output unit 603 is further configured to play the second audio data associated with the first object in response to the second input when the object type of the first object is a preset object type;
  • the second audio data includes at least one of the following:
  • the processor 610 may be configured to update the first object to be the closest to the end position of the third input in response to the third input when the third input is received, and the object type is a preset object type second object;
  • the audio output unit 603 is further configured to play the second audio data associated with the updated first object.
  • the processor 610 may be configured to, in the case of receiving a fourth input to the first object, stop playing the video associated with the second object whose object type is different from that of the first object in response to the fourth input.
  • the processor 610 may be configured to, in the case of receiving a fifth input to the picture, in response to the fifth input, determine from the picture a target image area that matches the fifth input, and stop playing the video associated with the fourth object.
  • the first audio data, the fourth object is a second object whose corresponding image area is located outside the target image area;
  • the audio output unit 603 is further configured to play third audio data associated with the first object in response to the sixth input when receiving the sixth input to the first object;
  • the processor 610 may be configured to, in a case of receiving the seventh input, respond to the seventh input, and adjust the playing volume of the first audio data according to an input parameter of the seventh input.
  • the processor 610 may be configured to determine the background image area in the picture and the image area corresponding to each second object; determine straight lines from the background image area; The positional relationship of the lines, determining the distance between any two second objects.
  • the radio frequency unit 601 can be used to send the picture to the server, in response to the first input, send the input parameters of the first input to the server, and receive the audio playback rules sent by the server;
  • the server is used to identify the picture and obtain the coordinates of the image area corresponding to each second object among the N second objects in the picture, and the server is also used to generate an audio playback rule, the audio playback rule includes the associated first audio data of each second object and its playback parameters;
  • the audio output unit 603 is further configured to play the first audio data associated with each second object according to the audio playing rule.
  • the input unit 604 may include a graphics processor (Graphics Processing Unit, GPU) 6041 and a microphone 6042, and the graphics processor 6041 is used for the image capture device (such as the image data of the still picture or video obtained by the camera) for processing.
  • the display unit 606 may include a display panel 6061, and the display panel 6061 may be configured in the form of a liquid crystal display, an organic light emitting diode, or the like.
  • the user input unit 607 includes a touch panel 6071 and other input devices 6072 .
  • the touch panel 6071 is also called a touch screen.
  • the touch panel 6071 may include two parts, a touch detection device and a touch controller.
  • Other input devices 6072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, switch buttons, etc.), trackballs, mice, and joysticks, which will not be repeated here.
  • Memory 609 may be used to store software programs as well as various data, including but not limited to application programs and operating systems.
  • the processor 610 may integrate an application processor and a modem processor, wherein the application processor mainly processes operating systems, user interfaces, and application programs, and the modem processor mainly processes wireless communications. It can be understood that the foregoing modem processor may not be integrated into the processor 610 .
  • the embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, each process of the above audio data playback method embodiment is realized, and the same Technical effects, in order to avoid repetition, will not be repeated here.
  • the embodiment of the present application further provides a chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, the processor is used to run programs or instructions, and realize the various processes of the above audio data playback method embodiment, and can achieve the same To avoid repetition, the technical effects will not be repeated here.
  • chips mentioned in the embodiments of the present application may also be called system-on-chip, system-on-chip, system-on-a-chip, or system-on-a-chip.
  • the term “comprising”, “comprising” or any other variation thereof is intended to cover a non-exclusive inclusion such that a process, method, article or apparatus comprising a set of elements includes not only those elements, It also includes other elements not expressly listed, or elements inherent in the process, method, article, or device. Without further limitations, an element defined by the phrase “comprising a " does not preclude the presence of additional identical elements in the process, method, article, or apparatus comprising that element.
  • the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved. Functions are performed, for example, the described methods may be performed in an order different from that described, and various steps may also be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Data Mining & Analysis (AREA)
  • Software Systems (AREA)
  • Library & Information Science (AREA)
  • Databases & Information Systems (AREA)
  • Mathematical Physics (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Medical Informatics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Computing Systems (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • General Health & Medical Sciences (AREA)
  • Human Computer Interaction (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

一种音频数据播放方法与装置,属于通信技术领域。其中,音频数据播放方法,包括:在显示图片的情况下,接收对图片中的第一对象的第一输入,图片包括N个第二对象,第一对象为N个第二对象中的任一第二对象(101),N为大于1的整数;响应于第一输入,播放每一第二对象关联的第一音频数据,每一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关(102)。

Description

音频数据播放方法与装置
相关申请的交叉引用
本申请主张在2021年08月23日在中国提交的中国专利申请号202110971052.3的优先权,其全部内容通过引用包含于此。
技术领域
本申请属于通信技术领域,具体涉及一种音频数据播放方法与装置。
背景技术
目前,随着对残障人士的关注度的提高,越来越多的电子设备支持无障碍模式,以便残障用户能够便捷使用电子设备。通常来说,在无障碍模式下,电子设备可以对用户点击区域的内容进行识别,并将识别结果以音频的形式进行播放,以告知用户点击区域的内容。
然而,现有技术中,电子设备在显示图片时,往往是对图片局部的内容进行识别与音频播放,难以向用户传达图片的整体内容。
发明内容
本申请实施例的目的是提供一种音频数据播放方法与装置,能够解决现有技术对图片局部的内容进行识别与音频播放,难以向用户传达图片的整体内容的问题。
第一方面,本申请实施例提供了一种音频数据播放方法,该方法包括:
在显示图片的情况下,接收对图片中的第一对象的第一输入,其中,图片包括N个第二对象,第一对象为N个第二对象中的任一第二对象,N为大于1的整数;
响应于第一输入,播放每一第二对象关联的第一音频数据,其中,每 一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关。
第二方面,本申请实施例提供了一种音频数据播放装置,该装置包括:
第一接收模块,用于在显示图片的情况下,接收对图片中的第一对象的第一输入,其中,图片包括N个第二对象,第一对象为N个第二对象中的任一第二对象,N为大于1的整数;
第一播放模块,用于响应于第一输入,播放每一第二对象关联的第一音频数据,其中,每一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关。
第三方面,本申请实施例提供了一种电子设备,该电子设备包括处理器、存储器及存储在存储器上并可在处理器上运行的程序或指令,程序或指令被处理器执行时实现如第一方面的方法的步骤。
第四方面,本申请实施例提供了一种可读存储介质,可读存储介质上存储程序或指令,程序或指令被处理器执行时实现如第一方面的方法的步骤。
第五方面,本申请实施例提供了一种芯片,芯片包括处理器和通信接口,通信接口和处理器耦合,处理器用于运行程序或指令,实现如第一方面的方法。
第六方面,本申请实施例提供了一种计算机程序产品,包括有形地包含在计算机可读介质上的计算机程序,所述计算机程序包含用于执行如第一方面所述的方法的程序代码。
第七方面,本申请实施例提供了一种电子设备,被配置成用于执行如第一方面所述的方法的步骤。
本申请实施例提供的音频数据播放方法,在显示图片的情况下,接收对图片中的第一对象的第一输入,该图片包括多个第二对象,第一对象为这多个第二对象中的任一第二对象;响应于第一输入,播放每一第二对象关联的第一音频数据,每一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关。本申请实施例有助于向用户传达图片 的整体内容,提高用户体验。
附图说明
图1是本申请实施例提供的音频数据播放方法的流程示意图;
图2是本申请实施例中图片的一个示例图;
图3是在一个具体应用例中,音频数据播放方法的流程示意图;
图4是本申请实施例提供的音频数据播放装置的结构示意图;
图5是本申请实施例提供的一种电子设备的结构示意图;
图6是本申请实施例提供的一种电子设备的硬件结构示意图。
具体实施方式
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员获得的所有其他实施例,都属于本申请保护的范围。
本申请的说明书和权利要求书中的术语“第一”、“第二”等是用于区别类似的对象,而不用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便本申请的实施例能够以除了在这里图示或描述的那些以外的顺序实施,且“第一”、“第二”等所区分的对象通常为一类,并不限定对象的个数,例如第一对象可以是一个,也可以是多个。此外,说明书以及权利要求中“和/或”表示所连接对象的至少其中之一,字符“/”,一般表示前后关联对象是一种“或”的关系。
下面结合附图,通过具体的实施例及其应用场景对本申请实施例提供的音频数据播放方法与装置进行详细地说明。
如图1所示,本申请实施例提供的音频数据播放方法,包括:
步骤101,在显示图片的情况下,接收对图片中的第一对象的第一输入,其中,图片包括N个第二对象,第一对象为N个第二对象中的任一第二对象,N为大于1的整数;
步骤102,响应于第一输入,播放每一第二对象关联的第一音频数 据,其中,每一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关。
本申请实施例提供的音频数据播放方法,可以应用于电子设备。该电子设备可以是移动终端或者个人电脑等,此处不做具体限定。
电子设备可以显示图片,该图片中可以包括多个第二对象。
举例来说,上述的图片可以是在海边拍摄得到风景图片,相应地,图片可能包括大海、树以及鸟等第二对象。或者,图片可以是在车站拍摄得到图片,相应地,图片可能包括人物A、人物B以及车辆等第二对象。
容易理解的是,图片中的第二对象,可以通过对图片的识别得到。
在一些举例中,可以预先基于深度学习模型,从图片中识别出各个第二对象。而在另一些举例中,也可以预先通过人工识别的方式,从图片中识别出各第二对象。
为简化描述,以下将主要以图片中的各个第二对象为通过深度学习模型进行识别得到为例进行说明。
在一个举例中,可以预先将图片发送至服务器,服务器使用深度学习模型对图片进行识别,并可以将相关的识别结果发送至电子设备。这些识别结果中可以包括了上述的N个第二对象。
在另一个举例中,也可以是电子设备直接使用深度学习模型对图片进行识别,以得到N个第二对象。
在步骤101中,电子设备可以接收对第一对象的第一输入,也就是任一第二对象的第一输入。
第一输入可以对应了单击、多击或者长按等形式的输入,此处不做具体限定。
容易理解的是,深度学习模型识别第二对象时,可以同时获取到各个第二对象在图片中的位置,具体来说,可以获取到各个第二对象对应的图像区域在图片中的位置。
在接收到第一输入的情况下,电子设备可以获取第一输入在图片中的输入位置,基于该输入位置,以及各个第二对象在图片中的位置,电子设备从图片中确定出上述的第一对象。
当然,在一些可行的实施方式中,电子设备也可以根据第一输入在图片中的输入位置,实时对输入位置及其预设距离范围内的图像区域进行识别,进而确定出第一对象。
步骤102中,电子设备可以响应于第一输入,播放每一第二对象关联的第一音频数据。
如上文所示的,通过深度学习模型可以识别出各第二对象。在一些举例中,每一类第二对象可以关联有一标识,该标识在一定程度上可以体现对第二对象的分类与识别结果。
比如,对于人物这一类第二对象,可以关联有标识“人物”,该标识可以是通过文本的方式进行表达的。当然,在一些可行的实施方式中,上述的标识也可以是通过编号等方式进行表达的。
每一类第二对象可以关联有相应的第一音频数据,在一些可行的实施方式中,第二对象与第一音频数据的关联关系,可以体现在第二对象与第一音频数据的关联关系中。
结合一些应用场景的举例,电子设备可以根据各个第二对象的标识,在预设的音频数据库中查询各第二对象关联的第一音频数据,以便对各第二对象关联的第一音频数据进行播放。
当然,在另一些应用场景中,也可以是上述服务器在完成对第二对象的识别的情况下,从音频数据块中查询各第二对象关联的第一音频数据,并将各第一音频数据以及音频数据第二对象的关联关系发送至电子设备,以供电子设备进行播放。
为便于理解第一音频数据,以下主要以播放第一音频数据时发出的声音内容为例,来对第一音频数据进行代替说明。
在一些示例中,第一音频数据可以是对关联的第二对象简单的描述词。比如,对于人物这类第二对象,关联的第一音频数据可以是“人”;对于小狗这类第二对象,关联的第一音频数据可以是“小狗”;对于大海这类第二对象,关联的第一音频数据可以是“大海”。
而在另一示例中,第一音频数据也可以是关联的第二对象在实际环境下经常发出的声音。比如,对于人物这类第二对象,关联的第一音频数据 可以是“你好”;对于小狗这类第二对象,关联的第一音频数据可以是“汪汪汪”;对于大海这类第二对象,关联的第一音频数据可以是“哗哗哗”(海浪的声音)。
为了比较准确地向用户传达图片的整体内容,本实施例中,可以播放N个第二对象中每一第二对象关联的第一音频数据,同时,各个第二对象关联的第一音频数据的播放参数,可以与各第二对象与第一对象之间的距离相关。
比如,第一音频数据的播放参数可以包括播放音量,各个第二对象关联的第一音频数据的播放音量,可以与各第二对象与第一对象之间的距离负相关。换而言之,对于一个第二对象,其与第一对象之间的距离越近,则其关联的第一音频数据的播放音量则可以越高。
再比如,第一音频数据的播放参数可以包括播放速度,各个第二对象关联的第一音频数据的播放频率,可以与各第二对象与第一对象之间的距离负相关。换而言之,对于一个第二对象,其与第一对象之间的距离越近,则其关联的第一音频数据的播放速度则可以越快。
当然,以上是对播放参数的类型,以及播放参数与距离的相关关系的一些举例说明,在实际应用中,播放参数的类型与相关关系均可以根据需要进行设定。
为了简化说明,以下主要以播放参数包括播放音量为例进行说明。
对于第一对象,其与自身之间的距离为0。对于除第一对象以外的任一第二对象,其与第一对象之间的距离,可以是在图像坐标系中的距离,也可以是在大地坐标系等坐标系中的距离。
比如,在图像坐标系中,任一第二对象的坐标,可以是该第二对象对应的图像区域的中点在图像坐标系中的坐标;第一对象与第二对象之间的距离,可以通过第一对象的坐标与第二对象的坐标进行计算。
再比如,可以从图片中识别出直线线条,并将直线线条作为参考线,来确定第一对象与第二对象之间的距离。结合一些应用场景,直线线条通常可以是道路边界、水面与天空的交界线等等。以直线线条作为参考线,可以得到第一对象与第二对象在大地坐标系中大致的距离。
为便于更好地理解各个第二对象至第一对象的距离,与各个第二对象关联的第一音频数据的播放音量的关系,以下结合一个举例进行说明。
对于第一对象,其与自身之间的距离为0,因此,第一对象关联的第一音频数据可以按照较高的音量进行播放,例如,以电子设备最大音量的50%进行播放。N个第二对象中包括第一对象、第二对象A以及第二对象B,其中,第二对象A与第一对象的距离,小于第二对象B与第一对象的距离。则可以以40%的最大音量播放第二对象A关联的第一音频数据,以30%的最大音量播放第二对象B关联的第一音频数据。
可见,通过播放图片中各个第二对象关联的第一音频数据,可以便于用户了解图片中所包括的对象;而各第二对象关联的第一音频数据的播放音量,与各第二对象与第一对象之间的距离负相关,有助于用户能够根据播放音量确定各第二对象的相对位置。综合各第二对象关联的第一音频数据,以及从播放音量中体现的各第二对象的相对位置,用户可以获知图片中各第二对象的类型与分布,从而比较准确地理解图片整体表达的内容。
类似地,当各第二对象关联的第一音频数据的播放参数包括播放速度时,也可以通过播放速度来体现各第二对象的相对位置。
本申请实施例提供的音频数据播放方法,在显示图片的情况下,接收对图片中的第一对象的第一输入,该图片包括多个第二对象,第一对象为这多个第二对象中的任一第二对象;响应于第一输入,播放每一第二对象关联的第一音频数据,每一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关。本申请实施例有助于向用户传达图片的整体内容,提高用户体验。
与此同时,用户对图片进行第一输入,即可比较准确地理解图片的内容,因此,本申请实施例提供的音频数据播放方法,也能够有效提升操作便捷度。
在一个实施方式中,电子设备响应于第一输入,可以先播放第一对象关联的第一音频数据,在间隔预设时长后,再播放除第一对象以外的第二对象所关联的第一音频数据。
进一步地,除第一对象以外的第二对象所关联的第一音频数据的播放 顺序,也可以是与各第二对象至第一对象的距离相关的,第二对象距离第一对象越远,相关联的第一音频数据的播放顺序约靠后。
可选地,每一第二对象关联的第一音频数据的播放音色,与每一第二对象对应的图像区域的颜色相匹配。
如上文所示的,在完成对图片中的第二对象的识别的情况下,每一第二对象在图片中的图像区域也可以得到确定。而各个图像区域中像素的颜色通常是已知的,根据图像区域中各像素的颜色,可以确定图像区域的颜色。
通常情况下,像素的颜色可以通过RGB值或者灰度值进行表示,为简化描述,以下主要以像素的颜色通过RGB值进行表示为例进行说明。
结合一些举例,第二对象在图片中的图像区域中可能包括了多个像素,该图像区域的颜色,可以对应该图像区域中各像素的RGB值的众数或者平均数等。
播放音色可以与颜色存在预设的匹配关系,例如,播放音色可以与颜色的冷暖程度或者色系相匹配。相应地,在确定了各个第二对象对应的图像区域的颜色的情况下,可以根据该预设的匹配关系,确定各个第二对象关联的第一音频数据的播放音色。
举例来说,颜色的冷暖程度可以通过如下公式进行计算:
CW=r*0.299+g*0.578+b*0.114
其中,CW是用于衡量颜色冷暖程度的值,r是RGB值中红色通道的数值,g是RGB值中绿色通道的数值,b是RGB值中蓝色通道的数值。r、g以及b的取值范围通常均为0~255。
在一个示例中,当CW大于或等于192时,可以认为是冷色,当CW小于192时,可以认为是暖色。对于冷色,可以匹配有比较沉闷的音色;对于暖和,可以匹配有比较轻快的音色。
或者,在另一个示例中,可以将CW的取值范围(0~253)按照预设步长划分为多个区间,按照CW从小到大的顺序,针对各个区间按轻快到沉闷的趋势匹配音色。
结合一个应用场景,白天拍摄的包括大海的图片中,大海这一第二对 象的颜色为蓝色,相应的CW较低,匹配的音色比较轻快,因此,在播放大海对应的第一音频数据时,可以以较轻快的音色进行播放。
而晚上拍摄的包括大海的图片中,大海这一第二对象的颜色近似黑色或墨绿色,相应的CW较高,匹配的音色比较沉闷,在在播放大海对应的第一音频数据时,可以以较沉闷的音色进行播放。
当然,在实际应用中,上述用于衡量颜色冷暖程度的值的计算方式,也可以根据实际需要进行调整。
如上文所示的,音色与颜色的匹配关系,还可以体现在音色与颜色所属的色系的匹配关系,也就是说,针对不同的色系,可以匹配不同的音色。为了简化说明,此处不再对根据各第二对象对应的图像区域的颜色的色系,确定各第二对象的第一音频数据的播放音色的过程进行详细说明。
结合上文中的举例,各个第二对象关联的第一音频数据的播放音色,可以是服务器完成对图片中的各第二对象识别后,根据各第二对象对应的图像区域的颜色确定的。后续服务器可以将播放音色与第二对象之间的关联关系发送至电子设备,以便电子设备选择合适的播放音色对各第二对象关联的第一音频数据进行播放。
当然,也可以是电子设备在接收到第一输入的情况下,再对图片进行识别得到N个第二对象,根据各第二对象对应的图像区域的颜色,以及预设的播放音色与颜色之间的匹配关系,确定出各第二对象对应的播放音色,并以确定的播放音色,播放第二对象关联的第一音频数据。
本实施例中,对于播放音色的确定时机可以不作具体限定,保证在播放各第二对象关联的第一音频数据时,播放音色与各第二对象对应的图像区域的颜色相匹配即可。本实施例中播放音色与第二对象对应图像区域的颜色匹配,有助于更有效地向用户传达图片的内容,提升用户对图片内容的收听体验。
实际应用中,例如大海、道路等类型的第二对象,在图片中对应的图像区域可能被其他的第二对象分割为多个子图像区域。相应地。在一些实施方式中,可以根据这些子图像区域的颜色,将颜色相同或相近的子图像区域归并为同一个第二对象对应的图像区域,再根据归并得到的图像区域 的颜色的冷暖程度或色系,确定匹配的播放音色。
可选地,上述步骤102,响应于第一输入,播放每一第二对象关联的第一音频数据之后,音频数据播放方法还可以包括:
接收对第一对象的第二输入;
在第一对象的对象类型为预设对象类型的情况下,响应于第二输入,播放第一对象关联的第二音频数据;
其中,第二音频数据包括以下至少一项:
用于提示至少一个第二对象的状态的音频数据;
用于提示至少一个第三对象至第一对象的距离的音频数据,第三对象为N个第二对象中除第一对象之外的第二对象。
容易理解的是,第一对象可以是N个第二对象中的任一个第二对象,而第一对象的确定可以是与用户的输入相关的。比如,若用户的第一输入,是对大海这一第二对象的点击输入,则可以将点击的大海确定为第一对象;若用户的第二输入,是对人物这一第二对象的长按输入,则可以将第一对象更新为长按的人物。
当然,此时的第二输入的具体方式,可以并不限于长按输入,还可以是其他预设的手势输入等。比如,第二输入可以对应划问号或者划勾的手势输入。
本实施例中,在第一对象的对象类型为预设对象类型的情况下,电子设备可以响应于第二输入,播放第一对象关联的第二音频数据。
举例来说,预设对象类型可以是人物,当用户对对象类型为人物的第一对象进行第二输入后,电子设备可以以人物的口吻,来播放第一对象关联的第二音频数据。
而在另一些举例中,预设对象类型还可以是鹦鹉、喇叭或者其他对象类型,可以根据实际需要进行设置。
结合以上举例可见,作为一些可行的实施方式,预设对象类型可以对应在实际环境中能够发出提示性语言的对象。当然,在其他一些实施方式中,预设对象类型也可以动物或者植物等对象类型,对应的对象可以通过拟人化的方式发出提示性的语言。
为了简化描述,以下将主要以预设对象类型为人物为例进行说明。
在一个示例中,第二音频数据可以包括用于提示至少一个第二对象的状态的音频数据。
结合一个应用场景,第一对象为人物时,可以以第一人称的口吻,对包括自身在内的各个第二对象进行状态的介绍。第二音频数据可以对应用于状态的介绍的音频数据。比如,第二音频数据可以是“我是人物A,正在读报”、“我的身后是蓝色的大海”等。
也就是说,第二对象的状态,可以指第二对象的行为状态、颜色状态等等,可以根据实际需要进行设置,此处可以不做具体限定。
在另一个应用场景下,第一对象为人物时,可以以第一人称的口吻,对至少一个第三对象与自身的距离进行介绍。第二音频数据可以对应用于提示至少一个第三对象与第一对象之间的距离。比如,第二音频数据可以是“在我右侧约两指宽的距离处是人物B”、“我的左侧是一棵树,它大约距离我一指宽”等。
当然,在又一应用场景下,第二音频数据也可以同时用于提示上述的状态与距离。比如,第二音频数据可以是“在我左侧约两指宽的距离处是人物B,他正在打电话”。
可见,本实施例中,接收对第一对象的第二输入,在第一对象的对象类型为预设对象类型的情况下,响应于第二输入,播放第一对象关联的第二音频数据,有助于向用户传达图片中各个第二对象的状态或者距离等信息,以便用户更好地理解图片所表达的内容。
可选地,响应于第二输入,播放第一对象关联的第二音频数据之后,方法还包括:
在接收到第三输入的情况下,响应于第三输入,将第一对象更新为距离第三输入的终止位置最近,且对象类型为预设对象类型的第二对象;
播放更新后的第一对象所关联的第二音频数据。
以下结合一个应用例,来对本实施例的实施过程进行介绍。
该应用例中,图片中可以包括人物A和人物B这两个第二对象,用户可以点击(对应第一输入)人物A对应的图像区域,此时,电子设备可以 播放人物A关联的第一音频数据,例如“人物”或者“你好”等。
当用户在人物A对应的图像区域进行划问号的输入时(对应第二输入),此时,电子设备可以播放人物A关联的第二音频数据,例如“在我右侧约两指宽的距离处是人物B”。
当用户从人物A所在的图像区域向右时(即滑动输入,对应第三输入),电子设备可以根据第三输入的终止位置,来重新确定第一对象。
比如,当第三输入的终止位置在人物B所在的图像区域时,可以将第一对象更新为人物B,此时,可以播放与人物B关联的第二音频数据,例如“我是人物B,我正在打电话,在我的右侧约一指宽的距离处是人物C”。
当然,在实际应用中,第三输入的具体方式,也可以是长按输入等,此处可以不做具体限定。比如,用户可以根据人物A关联的第二音频数据的提示,确定出人物B对应的图像区域,进而可以对人物B对应的图像区域进行长按输入。
而至于预设对象类型,在上一实施例中进行了说明,此处不再赘述。
基于以上应用例可见,本实施例中,根据第三输入的终止位置,更新第一对象,并播放更新后的第一对象所关联的第二音频数据,有助于引导用户更为详细地获取图片中各个第二对象的状态或位置等信息,便于用户对图片的整体内容进行理解,提升用户体验。
在一个示例中,用户在进行第三输入时,可以是手指从人物A所在的图像区域划至人物B所在的图像区域。在滑动过程中,当用户手指距离人物A的图像区域较近时,可以保持播放人物A关联的第二音频数据,并且,随着用户手指至人物A的图像区域的距离的增加,第二音频数据的播放音量可以减小。
当用户的手指距离人物B所在的图像区域较近时,可以将第一对象从人物A更新为人物B,进而播放人物B关联的第二音频数据,并且,随着用户手指至人物B的图像区域的距离的减小,第二音频数据的播放音量可以增大。
也就是说,该示例中,可以根据第三输入对应的实时输入位置,确定 第二音频数据的内容与播放参数(例如播放音量与播放速度),从而使得用户能够对第三输入的输入位置与各个第二对象之间的距离关系进行实时获取,便于更好地引导用户获取图片所表达的内容。
在一个示例中,电子设备可以统计在第三输入期间播放的全部第二音频数据所关联的P个第二对象,以及图片中对象类型为预设对象类型的第二对象的数量Q。其中,Q为正整数,P为小于或等于Q的正整数。
在第三输入结束后,电子设备可以输出Q-P的值,结合以上应用例,Q-P的值,可以认为是未进行状态或距离等信息的介绍的第二对象。至于Q-P的值,也可以是以音频播放的方式输出。如此,可以使得用户能够比较完整地了解图片中所展现的内容。
可选地,响应于第一输入,播放每一第二对象关联的第一音频数据之后,音频数据播放方法还可以包括以下至少一项:
在接收到对第一对象的第四输入的情况下,响应于第四输入,停止播放对象类型与第一对象的对象类型不同的第二对象所关联的第一音频数据;
在接收到对图片的第五输入的情况下,响应于第五输入,从图片中确定与第五输入匹配的目标图像区域,停止播放第四对象关联的第一音频数据,第四对象为对应的图像区域位于目标图像区域之外的第二对象;
在接收到对第一对象的第六输入的情况下,响应于第六输入,播放于第一对象关联的第三音频数据;
在接收到第七输入的情况下,响应于第七输入,根据第七输入的输入参数,调整第一音频数据的播放音量。
如上文所示的,第一对象可以是N个第二对象中的任一个第二对象,而第一对象的确定可以是与用户的输入相关的。在不同的输入阶段,第一对象可以是N个第二对象中的同一第二对象或者不同的第二对象。
以下结合一些应用例来对本实施例进行说明。在这些应用例中,图片可以包括人物A、人物B、大海以及草地。
在第一个应用例中,第四输入可以是对第一对象的双击或多击输入。例如,用户可以对图片中大海进行双击,电子设备响应于用户对大海这一 第一对象的双击输入,保持对大海关联的第一音频数据的播放。例如,可以保持“哗哗哗”的海浪声音的播放;或者,大海关联的第一音频数据,还可以包括海风“呼呼呼”的声音,以及海鸟的叫声,保持对这些第一音频数据的播放即可。
而相应地,电子设备可以停止播放人物A、人物B以及草地分别关联的第一音频数据的播放。
例如,在接收到第一输入的情况下,电子设备可以响应于第一输入,播放各个第二对象关联的第一音频数据。其中,人物A关联的第一音频数据,可以包括“你好”、“我是人物”以及“有什么可以帮助你的呢”,这些第一音频数据可以间隔预设时长轮流播放或随机播放。而在接收到对大海的第四输入后,电子设备不再播放人物A关联的第一音频数据。
当然,如果第四输入为对人物A或者人物B的双击输入,则响应于第四输入,可以保持对人物A关联的第一音频数据与人物B关联的第一音频数据的播放,而停止播放大海关联的第一音频数据与草地关联的第一音频数据。
在第二个应用例中,第五输入可以是捏合输入。比如,捏合输入可以具体是至少三根手指相互靠拢的手势输入。
电子设备检测到上述捏合输入时,可以根据至少三根手指的终止位置,从图片中确定出与第五输入匹配的目标图像区域。比如,目标图像区域可以是三根手指终止位置的触点的连线围合的图像区域。
而在判断各个第二对象对应的图像区域是否位于目标图像区域之内时,可以是判断各第二对象对应的图像区域的中点是否位于目标图像区域之内;或者,可以是判断各第二对象对应的图像区域是否整体或部分位于目标图像区域之内等等,可以根据实际需要进行设置。
为简化说明,可以认为当某一第二对象对应的图像区域全部位于目标图像区域之外时,判定该第二对象对应的图像区域位于目标图像区域之外,该第二对象即可以确定为上述第四对象。后续电子设备可以停止对第四对象关联的第一音频数据的播放。
如此,电子设备可以根据用户的输入情况,对用户更加关注的第二对 象进行关联音频数据的播放。
当然,第五输入也可以是其他类型的手势输入,例如,第五输入可以是沿封闭轨迹的输入,则目标图像区域可以是对应封闭轨迹围合的区域。
在第三个应用例中,第六输入可以是往返滑动输入。
比如,用户对图片中的草地进行往返滑动输入时,电子设备可以发出“沙沙沙”的声音,以模仿草地被拨动的声音。
同一个第二对象关联的第三音频数据与第一音频数据之间可以存在差异。比如,对于草地这一第二对象,关联的第一音频数据,可以是频率较慢的“沙沙沙”的声音;而关联的第三音频数据,则可以是频率较快的“沙沙沙”的声音。
再比如,对于人物A这一第二对象,关联的第一音频数据,可以是“你好”,而关联的第三音频数据,可以是“请问有什么事情么”。
换而言之,用户对任一第二对象的第六输入,可以认为是与该第二对象进行的动作交互。相应地,从电子设备的角度来说,当接收到对第一对象的第六输入时,可以响应于第六输入,播放预设的体现第一对象被交互时发出的声音,即播放上述的第一对象关联的第三音频数据,从而使得用户能够获得较好的交互体验。
在第四个应用例中,用户可以在电子设备上以较小的幅度画圈,从而调整个第一音频数据的播放音量。该较小的幅度画圈的输入,可以认为对应第七输入。而较小的幅度的判断,可以是根据用户画圈的区域的大小进行判断。比如,当画圈的区域小于预设区域面积时,可以认为是以较小的幅度画圈。
容易理解的是,画圈的输入,也就是上述的第七输入,可以存在相应的输入参数,比如画圈的方向与圈数。
在一个示例中,当画圈的方向是顺时针时,可以调高各第一音频数据的播放音量;当画圈的方向是逆时针时,可以调低各第一音频数据的播放音量。而上述播放音量的调高或调低的程度,可以由圈数进行确定。
在一个示例中,在对各第一音频数据的播放音量进行调整后,各第一音频数据的播放音量之间的相对大小关系可以保持不变,也就是说,依然 存在每一第二对象关联的第一音频数据的播放音量,与每一第二对象至第一对象的距离负相关的关系。
结合以上应用例可见,本实施例中,电子设备可以根据用户的不同的手势输入,实现不同的音频数据播放功能,极大提高用户的操作便捷性。
可选地,上述步骤102,响应于第一输入,播放每一第二对象关联的第一音频数据之前,音频数据播放方法还可以包括:
确定图片中的背景图像区域以及每一第二对象对应的图像区域;
从背景图像区域中确定出直线线条;
根据每一第二对象对应的图像区域与直线线条的位置关系,确定任两个第二对象之间的距离。
一般情况下,图片中会存在背景图像区域,比如天空或者大地所在的图像区域等。在一个示例中,可以将识别为天空或大地的第二对象所在的图像区域,直接确定为背景图像区域。
在实际应用中,例如天空或大地所在的图像区域可能被分割为多个子图像区域,这些子图像区域的颜色可能相同或相近,因此可以根据子图像区域的颜色,将这些子图像区域归入到背景图像区域。
结合一些应用场景,图片可以是通过相机拍摄得到的,相应地,图片中的内容可以是呈透视图的形式进行呈现的。也就是说,图片中的第二对象,可以是呈近大远小的形式呈现的。
比如,如图2所示,图片中包括大地D1、公路D2、树D3、天空D4、人物D5以及车辆D6这些第二对象,公路D2在远端汇集成一个点TP,点TP可以是直线线条L1、直线线条L2以及直线线条L3之间的交点。其中,直线线条L1可以是大地D1与天空D4之间的分割线;直线线条L2与直线线条L3为大地D1与公路D2之间的分割线。
若将大地D1与天空D4作为背景图像区域,则直线线条L1、直线线条L2以及直线线条L3均可以从背景图像区域中确定出来。
而关于背景图像区域中直线线条的确定方式,可以通过图像分割或特征提取等技术进行获取,具体可以通过现有技术实现,此处不做赘述。
在从背景图像区域确定出直线线条的情况下,可以根据直线线条与各 第二对象对应的图像区域之间的位置关系,来确定任两个第二对象之间的距离。
为了简化说明,可以将第二对象对应的图像区域的中点,作为第二对象在图片中的位置。而至于第二对象对应的图像区域的确定方式,已在上文实施例中进行了说明,此处不做赘述
同样结合图2,对于人物D5与车辆D6,两个第二对象均与直线线条L2的距离较近,可以以直线线条L2作为参考,确定人物D5与车辆D6之间的距离。人物D5与车辆D6之间的连接线,可以分解至平行于直线线条L2的子线段,以及垂直于直线线条L2的子线段,根据两条子线段的长度,可以大致确定人物D5与车辆D6之间在两个方向上的距离,进而可以得到人物D5与车辆D6之间的距离。
可见,本实施例中,通过从背景图像区域确定出直线线条,基于直线线条来确定任两个第二对象之间的距离,可以比较准确地获取任两个第二对象在大地坐标系中的距离。后续在根据距离播放各第二对象关联的第一音频数据时,可以比较准确地向用户传达各第二对象与第一对象之间的距离关系。
可选地,上述步骤101,播放每一第二对象关联的第一音频数据之前,音频数据播放方法还可以包括:
将图片发送至服务器,服务器用于识别图片,得到图片中N个第二对象中每一第二对象对应的图像区域的坐标;
响应于第一输入,播放每一第二对象关联的第一音频数据,包括:
响应于第一输入,将第一输入的输入参数发送至服务器,服务器用于根据输入参数以及每一第二对象对应的图像区域的坐标,生成音频播放规则,音频播放规则包括每一第二对象的关联的第一音频数据及其播放参数;
接收服务器发送的音频播放规则;
根据音频播放规则播放每一第二对象关联的第一音频数据。
本实施例中,对图片的识别以及各个音频播放规则的确定,可以是在服务器中进行的,如此,可以降低对电子设备的硬件配置的要求,降低电 子设备计算资源的消耗。
结合一个应用场景,电子设备在显示图片时,可以将图片发送至服务器,服务器则可以对图片进行识别,得到图片中的各个第二对象,以及各个第二对象对应的图像区域在图片中的坐标。
各个第二对象可以通过文字或者其他形式的标识进行代表。相应地,服务器可以将各第二对象的标识以及在图片中的坐标进行关联存储。为简化说明,可以认为服务器将各第二对象的标识以及在图片中的坐标存储在第一映射表中。
电子设备在接收到第一输入时,可以将第一输入的输入参数发送至服务器。举例来说,第一输入的输入参数可以包括用户点击的图像区域相对于图片的位置。
服务器根据第一输入的输入参数,以及上述的第一映射表,可以确定用户点击的图像区域对应的第二对象,即确定上述的第一对象。
结合一个举例,服务器在建立音频播放规则时,可以主要进行如下处理过程:
一是根据上述第一映射表中各第二对象对应的图像区域的坐标,确定各第二对象与第一对象之间的距离,以进一步确定各个第二对象对应的音频播放音量。如上文所示的,这里的距离与音频播放音量可以是负相关的,即距离越大,音频播放音量越低,反之亦然。
二是根据各个第二对象的标识,以及对象音频数据对应关系(该对应关系可以认为是存储在第二映射表中的),从预设的音频数据库中查询各第二对象关联的第一音频数据。
如此,服务器可以将第二对象—第一音频数据—音频播放音量的对应关系,作为音频播放规则发送至电子设备。电子设备则可以根据该音频播放规则,播放每一第二对象关联的第一音频数据。
当然,在一些实施方式中,服务器还可以进一步根据各个第二对象对应的图像区域的颜色,确定各个第二对象关联的第一音频数据的播放音色,并将播放音色加入到上述的音频播放规则中。
或者,服务器还可以将根据第二对象与第一对象之间的距离,确定各 个第二对象对应的音频播放速度,并将播放速度加入到上述的音频播放规则中。
如图3所示,以下结合一个具体应用例,对本申请实施例提供的音频数据播放方法进行说明。
该具体应用例中,音频数据播放方法可以应用于电子设备中,该电子设备可以与服务器进行数据交互。音频数据播放方法包括:
步骤301,服务器解析图片中内容,提取图片中的第二对象;
容易理解的是,服务器解析的图片,可以是由电子设备发送至服务器的。而服务器可以是采用深度学习模型对图片进行解析。
步骤302,电子设备接收用户的第一输入,将第一输入的输入参数发送至服务器;
比如,用户可以点击图片中的某一图像区域,电子设备可以将点击的图像区域相对图片整体的位置信息,作为输入参数发送至服务器。
步骤303,服务器获取输入参数,确定第一对象,并按照相对第一对象从远到近的顺序,将各第二对象存储至预设数组中;
步骤304,服务器根据预设数组中各第二对象的顺序,按从小到大规则为各第二对象分配音频播放音量;
当然,在一些实施方式中,服务器也可以为第二对象分配音频播放速度或者其他类型的播放参数。
步骤305,计算各第二对象的图像区域的颜色的冷暖程度CW(以下可以简称第二对象的CW);
CW的一种可行的计算方式如下:
CW=r*0.299+g*0.578+b*0.114
其中,CW是用于衡量颜色冷暖程度的值,r是RGB值中红色通道的数值,g是RGB值中绿色通道的数值,b是RGB值中蓝色通道的数值。r、g以及b的取值范围通常均为0~255。
步骤306,判断CW是否大于或等于192,若是,执行步骤307,若否,执行步骤308;
步骤307,根据255与第二对象的CW的差值,确定第二对象对应的 音频播放音色的沉闷程度,执行步骤309;
步骤308,根据第二对象的CW与0的差值,确定第二对象对应的音频播放音色的清脆程度,执行步骤309;
步骤309,按照对各第二对象确定音频播放音量与音频播放音色,播放各第二对象关联的音频数据。
结合以上具体应用例可见,本申请实施例提供的音频数据播放方法,通过确定图片中各个第二对象音频播放音量与音频播放音色,可以比较准确地向用户传达图片的整体内容,满足残障用户对图片内容的理解需求。
需要说明的是,本申请实施例提供的音频数据播放方法,执行主体可以为音频数据播放装置,或者该音频数据播放装置中的用于执行音频数据播放方法的控制模块。本申请实施例中以音频数据播放装置执行音频数据播放方法为例,说明本申请实施例提供的音频数据播放装置。
如图4所示,本申请实施例提供的音频数据播放装置400,包括:
第一接收模块401,用于在显示图片的情况下,接收对图片中的第一对象的第一输入,其中,图片包括N个第二对象,第一对象为N个第二对象中的任一第二对象,N为大于1的整数;
第一播放模块402,用于响应于第一输入,播放每一第二对象关联的第一音频数据,其中,每一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关。
可选地,每一第二对象关联的第一音频数据的播放音色,与每一第二对象对应的图像区域的颜色相匹配。
可选地,音频数据播放装置400还可以包括:
第二接收模块,用于接收对第一对象的第二输入;
第二播放模块,用于在第一对象的对象类型为预设对象类型的情况下,响应于第二输入,播放第一对象关联的第二音频数据;
其中,第二音频数据包括以下至少一项:
用于提示至少一个第二对象的状态的音频数据;
用于提示至少一个第三对象至第一对象的距离的音频数据,第三对象为N个第二对象中除第一对象之外的第二对象。
可选地,音频数据播放装置400还可以包括:
更新模块,用于在接收到第三输入的情况下,响应于第三输入,将第一对象更新为距离第三输入的终止位置最近,且对象类型为预设对象类型的第二对象;
第三播放模块,用于播放更新后的第一对象所关联的第二音频数据。
可选地,音频数据播放装置400还可以包括以下至少一项:
第一停止播放模块,用于在接收到对第一对象的第四输入的情况下,响应于第四输入,停止播放对象类型与第一对象的对象类型不同的第二对象所关联的第一音频数据;
第二停止播放模块,用于在接收到对图片的第五输入的情况下,响应于第五输入,从图片中确定与第五输入匹配的目标图像区域,停止播放第四对象关联的第一音频数据,第四对象为对应的图像区域位于目标图像区域之外的第二对象;
第四播放模块,用于在接收到对第一对象的第六输入的情况下,响应于第六输入,播放于第一对象关联的第三音频数据;
调整模块,用于在接收到第七输入的情况下,响应于第七输入,根据第七输入的输入参数,调整第一音频数据的播放音量。
可选地,音频数据播放装置400还可以包括:
第一确定模块,用于确定图片中的背景图像区域以及每一第二对象对应的图像区域;
第二确定模块,用于从背景图像区域中确定出直线线条;
第四确定模块,用于根据每一第二对象对应的图像区域与直线线条的位置关系,确定任两个第二对象之间的距离。
可选地,音频数据播放装置400还可以包括:
发送模块,用于将图片发送至服务器,服务器用于识别图片,得到图片中N个第二对象中每一第二对象对应的图像区域的坐标;
相应地,第一播放模块401,可以包括:
发送单元,用于响应于第一输入,将第一输入的输入参数发送至服务器,服务器用于根据输入参数以及每一第二对象对应的图像区域的坐标, 生成音频播放规则,音频播放规则包括每一第二对象的关联的第一音频数据及其播放参数;
接收单元,用于接收服务器发送的音频播放规则;
播放单元,用于根据音频播放规则播放每一第二对象关联的第一音频数据。
本申请实施例提供的音频数据播放装置,在显示图片的情况下,接收对图片中的第一对象的第一输入,响应于第一输入,播放每一第二对象关联的第一音频数据,每一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关,如此,通过对各第二对象关联的第一音频数据的播放参数的处理,可以比较准确地向用户传达图片的内容。各第二对象关联的第一音频数据的播放音色,与各第二对象对应的图像区域的颜色相匹配,进一步方便了用户对图片内容的理解。此外,音频数据播放装置还可以响应用户相关的输入,来调整音频播放的焦点,从而满足用户对比较关注的图片内容的获取需求,提升用户使用体验。
本申请实施例中的音频数据播放装置可以是装置,也可以是终端中的部件、集成电路、或芯片。该装置可以是移动电子设备,也可以为非移动电子设备。示例性的,移动电子设备可以为手机、平板电脑、笔记本电脑、掌上电脑、车载电子设备、可穿戴设备、超级移动个人计算机(ultra-mobile personal computer,UMPC)、上网本或者个人数字助理(personal digital assistant,PDA)等,非移动电子设备可以为服务器、网络附属存储器(Network Attached Storage,NAS)、个人计算机(personal computer,PC)、电视机(television,TV)、柜员机或者自助机等,本申请实施例不作具体限定。
本申请实施例中的音频数据播放装置可以为具有操作系统的装置。该操作系统可以为安卓(Android)操作系统,可以为iOS操作系统,还可以为其他可能的操作系统,本申请实施例不作具体限定。
本申请实施例提供的音频数据播放装置能够实现图1至图3的方法实施例实现的各个过程,为避免重复,这里不再赘述。
可选地,如图5所示,本申请实施例还提供一种电子设备500,包括 处理器501,存储器502,存储在存储器502上并可在处理器501上运行的程序或指令,该程序或指令被处理器501执行时实现上述音频数据播放方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
需要说明的是,本申请实施例中的电子设备包括上述的移动电子设备和非移动电子设备。
图6为实现本申请实施例的一种电子设备的硬件结构示意图。
该电子设备600包括但不限于:射频单元601、网络模块602、音频输出单元603、输入单元604、传感器605、显示单元606、用户输入单元607、接口单元608、存储器609、以及处理器610等部件。
本领域技术人员可以理解,电子设备600还可以包括给各个部件供电的电源(比如电池),电源可以通过电源管理系统与处理器610逻辑相连,从而通过电源管理系统实现管理充电、放电、以及功耗管理等功能。图6中示出的电子设备结构并不构成对电子设备的限定,电子设备可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置,在此不再赘述。
其中,用户输入单元607,用于在显示图片的情况下,接收对图片中的第一对象的第一输入,其中,图片包括N个第二对象,第一对象为N个第二对象中的任一第二对象,N为大于1的整数;
音频输出单元603,用于播放每一第二对象关联的第一音频数据,其中,每一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关。
本申请实施例提供的电子设备,在显示图片的情况下,接收对图片中的第一对象的第一输入,该图片包括多个第二对象,第一对象为这多个第二对象中的任一第二对象;响应于第一输入,播放每一第二对象关联的第一音频数据,每一第二对象关联的第一音频数据的播放参数与每一第二对象至第一对象的距离相关。本申请实施例有助于向用户传达图片的整体内容,提高用户体验。
可选地,每一第二对象关联的第一音频数据的播放音色,与每一第二 对象对应的图像区域的颜色相匹配。
可选地,用户输入单元607,还可用于接收对第一对象的第二输入;
音频输出单元603,还可用于在第一对象的对象类型为预设对象类型的情况下,响应于第二输入,播放第一对象关联的第二音频数据;
其中,第二音频数据包括以下至少一项:
用于提示至少一个第二对象的状态的音频数据;
用于提示至少一个第三对象至第一对象的距离的音频数据,第三对象为N个第二对象中除第一对象之外的第二对象。
可选地,处理器610,可用于在接收到第三输入的情况下,响应于第三输入,将第一对象更新为距离第三输入的终止位置最近,且对象类型为预设对象类型的第二对象;
音频输出单元603,还可用于播放更新后的第一对象所关联的第二音频数据。
可选地,处理器610,可用于在接收到对第一对象的第四输入的情况下,响应于第四输入,停止播放对象类型与第一对象的对象类型不同的第二对象所关联的第一音频数据;
可选地,处理器610,可用于在接收到对图片的第五输入的情况下,响应于第五输入,从图片中确定与第五输入匹配的目标图像区域,停止播放第四对象关联的第一音频数据,第四对象为对应的图像区域位于目标图像区域之外的第二对象;
可选地,音频输出单元603,还可用于在接收到对第一对象的第六输入的情况下,响应于第六输入,播放于第一对象关联的第三音频数据;
可选地,处理器610,可用于在接收到第七输入的情况下,响应于第七输入,根据第七输入的输入参数,调整第一音频数据的播放音量。
可选地,处理器610,可用于确定图片中的背景图像区域以及每一第二对象对应的图像区域;从背景图像区域中确定出直线线条;根据每一第二对象对应的图像区域与直线线条的位置关系,确定任两个第二对象之间的距离。
可选地,射频单元601,可用于将图片发送至服务器,响应于第一输 入,将第一输入的输入参数发送至服务器,以及接收服务器发送的音频播放规则;
其中,服务器用于识别图片,得到图片中N个第二对象中每一第二对象对应的图像区域的坐标,服务器还用于根据输入参数以及每一第二对象对应的图像区域的坐标,生成音频播放规则,音频播放规则包括每一第二对象的关联的第一音频数据及其播放参数;
相应地,音频输出单元603,还可用于根据音频播放规则播放每一第二对象关联的第一音频数据。
应理解的是,本申请实施例中,输入单元604可以包括图形处理器(Graphics Processing Unit,GPU)6041和麦克风6042,图形处理器6041对在视频捕获模式或图像捕获模式中由图像捕获装置(如摄像头)获得的静态图片或视频的图像数据进行处理。显示单元606可包括显示面板6061,可以采用液晶显示器、有机发光二极管等形式来配置显示面板6061。用户输入单元607包括触控面板6071以及其他输入设备6072。触控面板6071,也称为触摸屏。触控面板6071可包括触摸检测装置和触摸控制器两个部分。其他输入设备6072可以包括但不限于物理键盘、功能键(比如音量控制按键、开关按键等)、轨迹球、鼠标、操作杆,在此不再赘述。存储器609可用于存储软件程序以及各种数据,包括但不限于应用程序和操作系统。处理器610可集成应用处理器和调制解调处理器,其中,应用处理器主要处理操作系统、用户界面和应用程序等,调制解调处理器主要处理无线通信。可以理解的是,上述调制解调处理器也可以不集成到处理器610中。
本申请实施例还提供一种可读存储介质,可读存储介质上存储有程序或指令,该程序或指令被处理器执行时实现上述音频数据播放方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
其中,处理器为上述实施例中的电子设备中的处理器。可读存储介质,包括计算机可读存储介质,如计算机只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、磁碟或者光盘等。
本申请实施例另提供了一种芯片,芯片包括处理器和通信接口,通信接口和处理器耦合,处理器用于运行程序或指令,实现上述音频数据播放方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
应理解,本申请实施例提到的芯片还可以称为系统级芯片、系统芯片、芯片系统或片上系统芯片等。
需要说明的是,在本文中,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者装置不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者装置所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者装置中还存在另外的相同要素。此外,需要指出的是,本申请实施方式中的方法和装置的范围不限按示出或讨论的顺序来执行功能,还可包括根据所涉及的功能按基本同时的方式或按相反的顺序来执行功能,例如,可以按不同于所描述的次序来执行所描述的方法,并且还可以添加、省去、或组合各种步骤。另外,参照某些示例所描述的特征可在其他示例中被组合。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以计算机软件产品的形式体现出来,该计算机软件产品存储在一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端(可以是手机,计算机,服务器,或者网络设备等)执行本申请各个实施例的方法。
上面结合附图对本申请的实施例进行了描述,但是本申请并不局限于上述的具体实施方式,上述的具体实施方式仅仅是示意性的,而不是限制性的,本领域的普通技术人员在本申请的启示下,在不脱离本申请宗旨和权利要求所保护的范围情况下,还可做出很多形式,均属于本申请的保护 之内。

Claims (13)

  1. 一种音频数据播放方法,包括:
    在显示图片的情况下,接收对图片中的第一对象的第一输入,所述图片包括N个第二对象,所述第一对象为所述N个第二对象中的任一第二对象,N为大于1的整数;
    响应于所述第一输入,播放每一所述第二对象关联的第一音频数据,每一所述第二对象关联的第一音频数据的播放参数与每一所述第二对象至所述第一对象的距离相关。
  2. 根据权利要求1所述的方法,其中,每一所述第二对象关联的第一音频数据的播放音色,与每一所述第二对象对应的图像区域的颜色相匹配。
  3. 根据权利要求1所述的方法,其中,所述响应于所述第一输入,播放每一所述第二对象关联的第一音频数据之后,所述方法还包括:
    接收对所述第一对象的第二输入;
    在所述第一对象的对象类型为预设对象类型的情况下,响应于所述第二输入,播放所述第一对象关联的第二音频数据,
    其中,所述第二音频数据包括以下至少一项:
    用于提示至少一个所述第二对象的状态的音频数据;
    用于提示至少一个第三对象至所述第一对象的距离的音频数据,所述第三对象为所述N个第二对象中除所述第一对象之外的第二对象。
  4. 根据权利要求3所述的方法,其中,所述响应于所述第二输入,播放所述第一对象关联的第二音频数据之后,所述方法还包括:
    在接收到第三输入的情况下,响应于所述第三输入,将所述第一对象更新为距离所述第三输入的终止位置最近,且对象类型为所述预设对象类型的第二对象;
    播放更新后的第一对象所关联的第二音频数据。
  5. 根据权利要求1所述的方法,其中,所述响应于所述第一输入,播放每一所述第二对象关联的第一音频数据之后,所述方法还包括以下至少 一项:
    在接收到对第一对象的第四输入的情况下,响应于所述第四输入,停止播放对象类型与所述第一对象的对象类型不同的第二对象所关联的第一音频数据;
    在接收到对所述图片的第五输入的情况下,响应于所述第五输入,从所述图片中确定与所述第五输入匹配的目标图像区域,停止播放第四对象关联的第一音频数据,所述第四对象为对应的图像区域位于所述目标图像区域之外的第二对象;
    在接收到对所述第一对象的第六输入的情况下,响应于所述第六输入,播放于所述第一对象关联的第三音频数据;
    在接收到第七输入的情况下,响应于所述第七输入,根据第七输入的输入参数,调整所述第一音频数据的播放音量。
  6. 根据权利要求1所述的方法,其中,所述响应于所述第一输入,播放每一所述第二对象关联的第一音频数据之前,所述方法还包括:
    确定所述图片中的背景图像区域以及每一所述第二对象对应的图像区域;
    从所述背景图像区域中确定出直线线条;
    根据每一所述第二对象对应的图像区域与所述直线线条的位置关系,确定任两个所述第二对象之间的距离。
  7. 根据权利要求1所述的方法,其中,所述播放每一所述第二对象关联的第一音频数据之前,所述方法还包括:
    将所述图片发送至服务器,所述服务器用于识别所述图片,得到所述图片中N个第二对象中每一所述第二对象对应的图像区域的坐标;
    所述响应于所述第一输入,播放每一所述第二对象关联的第一音频数据,包括:
    响应于所述第一输入,将所述第一输入的输入参数发送至所述服务器,所述服务器用于根据所述输入参数以及每一所述第二对象对应的图像区域的坐标,生成音频播放规则,所述音频播放规则包括每一所述第二对象的关联的第一音频数据及其播放参数;
    接收所述服务器发送的音频播放规则;
    根据所述音频播放规则播放每一所述第二对象关联的第一音频数据。
  8. 一种音频数据播放装置,包括:
    第一接收模块,用于在显示图片的情况下,接收对图片中的第一对象的第一输入,所述图片包括N个第二对象,所述第一对象为所述N个第二对象中的任一第二对象,N为大于1的整数;
    第一播放模块,用于响应于所述第一输入,播放每一所述第二对象关联的第一音频数据,每一所述第二对象关联的第一音频数据的播放参数与每一所述第二对象至所述第一对象的距离相关。
  9. 一种电子设备,包括处理器,存储器及存储在所述存储器上并可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如权利要求1-7任一项所述的音频数据播放方法的步骤。
  10. 一种可读存储介质,所述可读存储介质上存储程序或指令,所述程序或指令被处理器执行时实现如权利要求1-7任一项所述的音频数据播放方法的步骤。
  11. 一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如权利要求1-7中任一项所述的显示方法的步骤。
  12. 一种计算机程序产品,包括有形地包含在计算机可读介质上的计算机程序,所述计算机程序包含用于执行权利要求1-7中任一项所述的方法的程序代码。
  13. 一种电子设备,被配置成用于执行如权利要求1-7中任一项所述的方法的步骤。
PCT/CN2022/113074 2021-08-23 2022-08-17 音频数据播放方法与装置 Ceased WO2023025005A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202110971052.3A CN113672193B (zh) 2021-08-23 2021-08-23 音频数据播放方法与装置
CN202110971052.3 2021-08-23

Publications (1)

Publication Number Publication Date
WO2023025005A1 true WO2023025005A1 (zh) 2023-03-02

Family

ID=78545352

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/113074 Ceased WO2023025005A1 (zh) 2021-08-23 2022-08-17 音频数据播放方法与装置

Country Status (2)

Country Link
CN (1) CN113672193B (zh)
WO (1) WO2023025005A1 (zh)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113672193B (zh) * 2021-08-23 2024-05-14 维沃移动通信有限公司 音频数据播放方法与装置

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104157171A (zh) * 2014-08-13 2014-11-19 三星电子(中国)研发中心 一种点读系统及其方法
CN107885823A (zh) * 2017-11-07 2018-04-06 广东欧珀移动通信有限公司 音频信息的播放方法、装置、存储介质及电子设备
CN107895006A (zh) * 2017-11-07 2018-04-10 广东欧珀移动通信有限公司 音频播放方法、装置、存储介质及电子设备
CN109271628A (zh) * 2018-09-03 2019-01-25 东北大学 一种图像描述生成方法
CN110519636A (zh) * 2019-09-04 2019-11-29 腾讯科技(深圳)有限公司 语音信息播放方法、装置、计算机设备及存储介质
WO2020063614A1 (zh) * 2018-09-26 2020-04-02 上海肇观电子科技有限公司 一种智能眼镜跟踪方法、装置及智能眼镜、存储介质
US20200258422A1 (en) * 2019-02-12 2020-08-13 Can-U-C Ltd. Stereophonic apparatus for blind and visually-impaired people
CN113672193A (zh) * 2021-08-23 2021-11-19 维沃移动通信有限公司 音频数据播放方法与装置

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2006331240A (ja) * 2005-05-30 2006-12-07 Hitachi Ltd 情報処理装置および仮想現実動画像表示システム
JP2011250100A (ja) * 2010-05-26 2011-12-08 Sony Corp 画像処理装置および方法、並びにプログラム
KR20170039379A (ko) * 2015-10-01 2017-04-11 삼성전자주식회사 전자 장치 및 이의 제어 방법
US10249044B2 (en) * 2016-12-30 2019-04-02 Facebook, Inc. Image segmentation with touch interaction
CN108269460B (zh) * 2018-01-04 2020-05-08 高大山 一种电子屏幕的阅读方法、系统及终端设备
CN108509863A (zh) * 2018-03-09 2018-09-07 北京小米移动软件有限公司 信息提示方法、装置和电子设备

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104157171A (zh) * 2014-08-13 2014-11-19 三星电子(中国)研发中心 一种点读系统及其方法
CN107885823A (zh) * 2017-11-07 2018-04-06 广东欧珀移动通信有限公司 音频信息的播放方法、装置、存储介质及电子设备
CN107895006A (zh) * 2017-11-07 2018-04-10 广东欧珀移动通信有限公司 音频播放方法、装置、存储介质及电子设备
CN109271628A (zh) * 2018-09-03 2019-01-25 东北大学 一种图像描述生成方法
WO2020063614A1 (zh) * 2018-09-26 2020-04-02 上海肇观电子科技有限公司 一种智能眼镜跟踪方法、装置及智能眼镜、存储介质
US20200258422A1 (en) * 2019-02-12 2020-08-13 Can-U-C Ltd. Stereophonic apparatus for blind and visually-impaired people
CN110519636A (zh) * 2019-09-04 2019-11-29 腾讯科技(深圳)有限公司 语音信息播放方法、装置、计算机设备及存储介质
CN113672193A (zh) * 2021-08-23 2021-11-19 维沃移动通信有限公司 音频数据播放方法与装置

Also Published As

Publication number Publication date
CN113672193A (zh) 2021-11-19
CN113672193B (zh) 2024-05-14

Similar Documents

Publication Publication Date Title
WO2021135601A1 (zh) 辅助拍照方法、装置、终端设备及存储介质
JP5784141B2 (ja) 重畳筆記による手書き入力方法
US9354716B1 (en) Virtualization of tangible interface objects
US20220058436A1 (en) Method and apparatus for generating training sample of semantic segmentation model, storage medium, and electronic device
CN105303149B (zh) 人物图像的展示方法和装置
CN111757175A (zh) 视频处理方法及装置
WO2020063009A1 (zh) 图像处理方法、装置、存储介质及电子设备
WO2021213067A1 (zh) 物品显示方法、装置、设备及存储介质
CN110378287A (zh) 文档方向识别方法、装置及存储介质
CN108039995A (zh) 消息发送控制方法、终端及计算机可读存储介质
WO2020244074A1 (zh) 表情交互方法、装置、计算机设备及可读存储介质
CN111047511A (zh) 一种图像处理方法及电子设备
WO2018214115A1 (zh) 一种评价脸妆的方法及装置
US20250039537A1 (en) Screenshot processing method, electronic device, and computer readable medium
WO2023025060A1 (zh) 界面显示的适配处理方法、装置和电子设备
CN114095754A (zh) 视频处理方法、装置及电子设备
WO2023025005A1 (zh) 音频数据播放方法与装置
CN111080747B (zh) 一种人脸图像处理方法及电子设备
CN108805095A (zh) 图片处理方法、装置、移动终端及计算机可读存储介质
WO2024022149A1 (zh) 数据增强方法、装置及电子设备
CN113157966B (zh) 显示方法、装置及电子设备
CN114943872A (zh) 目标检测模型的训练方法、装置、目标检测方法、装置、介质及设备
CN112149599B (zh) 表情追踪方法、装置、存储介质和电子设备
CN117292384B (zh) 文字识别方法、相关装置及存储介质
CN111857499A (zh) 信息提示方法和装置

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22860351

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 22860351

Country of ref document: EP

Kind code of ref document: A1

122 Ep: pct application non-entry in european phase

Ref document number: 22860351

Country of ref document: EP

Kind code of ref document: A1

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 21.08.2024)

122 Ep: pct application non-entry in european phase

Ref document number: 22860351

Country of ref document: EP

Kind code of ref document: A1