WO2016107226A1 - 图像处理方法及装置 - Google Patents

图像处理方法及装置 Download PDF

Info

Publication number
WO2016107226A1
WO2016107226A1 PCT/CN2015/090279 CN2015090279W WO2016107226A1 WO 2016107226 A1 WO2016107226 A1 WO 2016107226A1 CN 2015090279 W CN2015090279 W CN 2015090279W WO 2016107226 A1 WO2016107226 A1 WO 2016107226A1
Authority
WO
WIPO (PCT)
Prior art keywords
frame
matching
source video
target video
current
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2015/090279
Other languages
English (en)
French (fr)
Inventor
孙茂杰
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Shenzhen TCL Digital Technology Co Ltd
Original Assignee
Shenzhen TCL Digital Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Shenzhen TCL Digital Technology Co Ltd filed Critical Shenzhen TCL Digital Technology Co Ltd
Publication of WO2016107226A1 publication Critical patent/WO2016107226A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof

Definitions

  • the present invention relates to the field of image processing technologies, and in particular, to an image processing method and apparatus.
  • the invention provides an image processing method and device, and the main purpose thereof is to solve the technical problem of how to obtain the continuity of the current action of the user in real time.
  • an image processing method including:
  • a preset accuracy matching algorithm is used to obtain an accuracy matching pair of the current action and an accuracy matching pair of the previous action, and obtain a source video of the current action accuracy matching pair.
  • the frame number and the target video frame sequence number, and the accuracy of the previous action match the source video frame sequence number and the target video frame sequence number;
  • a preset accuracy matching algorithm is used to obtain an accuracy matching pair of the current action, and the source video frame sequence number and the target video of the current action matching matching pair are obtained.
  • the steps of the frame number include:
  • the key points include: a head, a hand, an abdomen and a foot of the human body; and the key point group sequence of the source video matches a corresponding key point in a key point group sequence of the target video,
  • the steps to get the correct matching pair of the current action include:
  • the four angles are angles and heads formed by the head and the left hand
  • the source video frame sequence number and the target video frame sequence number of the matching pair according to the accuracy of the current action, and the source video frame sequence number and the target video frame sequence number of the matching match of the previous action are preset.
  • the coherence algorithm, the steps of calculating the coherent results of the current action include:
  • the coherence score of the current action is calculated; the specific calculation formula is as follows:
  • n v i v - i vLst
  • n c i c - i cLst
  • the current coherence score formula: s c 100 ⁇ (l cmin +(1-l cmin ) ⁇ l c ), where l c is the degree of coherence; s c is the current coherence score; i v is the current best match pair Medium source video frame number; i c is the current best matching pair target video frame number; i vLst is the previous best matching pair medium source video frame number; i cLst is the previous best matching pair target video frame number; n v is the number of frame intervals between the current source video key frame and the previous key frame; n c is the frame interval of the current target video key frame and the previous key frame; l cmin is the lower limit of the consistency score threshold ratio.
  • the method further comprises:
  • a weighted calculation is performed to obtain a comprehensive score of the current action.
  • the method further comprises:
  • the accuracy matching score, the coherence score and the comprehensive score calculated each time are summed and averaged, and the average score of the whole process is obtained.
  • An embodiment of the present invention further provides an image processing apparatus, including:
  • An image acquisition module configured to acquire a source video image and a target video image
  • the matching calculation module is configured to obtain a matching match between the accuracy matching pair of the current action and the accuracy of the previous action based on the source video image and the target video image, and obtain the accuracy of the current action.
  • the source video frame sequence number of the sexual matching pair and the target video frame sequence number, and the accuracy of the previous action match the source video frame sequence number and the target video frame sequence number;
  • the consistency calculation module is configured to match the source video frame sequence number and the target video frame sequence number of the pair according to the accuracy of the current action, and the source video frame sequence number and the target video frame sequence number of the matching accuracy of the previous action, and adopt a pre-predetermined
  • the coherence algorithm is set to calculate the coherence result of the current action.
  • the matching calculation module is further configured to obtain a sequence of source video frames of the current action from the source video image, and obtain a sequence of target video frames of the current action from the target video image; Extracting key point information, obtaining a key point group sequence of the source video and a key point group sequence of the target video; matching the key point group sequence of the source video with a corresponding key point in the key point group sequence of the target video to obtain a current
  • the accuracy of the action matches the pair, and the accuracy of the current action is matched to the source video frame number and the target video frame number.
  • the key points include: a head, a hand, an abdomen, and a foot of the human body;
  • the matching calculation module is further configured to acquire a key point group sequence of the source video and each of a plurality of frames of a key point group sequence of the target video based on four angles of the key point, the four angles The angle formed by the head and the left hand, the angle formed by the head and the right hand, the angle formed by the abdomen and the left foot, and the angle formed by the abdomen and the right foot; respectively calculating the key point group sequence of the source video and the target video a difference between corresponding angles in each frame of the key point group sequence, and obtaining a total difference of four corners corresponding to each frame; obtaining a minimum total difference value of total errors of the plurality of frames; according to the minimum total difference The value obtains the accuracy matching pair of the current action, and the accuracy matches the score.
  • the consistency calculation module is further configured to obtain a difference between a source video frame sequence number of the current action matching matching pair and a source video frame sequence number of a previous action matching pair, to obtain a source. a video key frame interval; obtaining a difference between the target video frame number of the current action matching pair and the target video frame number of the matching of the accuracy of the previous action, to obtain a target video key frame interval;
  • the source video key frame interval, the target video key frame interval, and the preset coherence formula calculate the coherence score of the current action; the specific calculation formula is as follows:
  • n v i v - i vLst
  • n c i c - i cLst
  • the current coherence score formula: s c 100 ⁇ (l cmin +(1-l cmin ) ⁇ l c ), where l c is the degree of coherence; s c is the current coherence score; i v is the current best match pair Medium source video frame number; i c is the current best matching pair target video frame number; i vLst is the previous best matching pair medium source video frame number; i cLst is the previous best matching pair target video frame number; n v is the number of frame intervals between the current source video key frame and the previous key frame; n c is the frame interval of the current target video key frame and the previous key frame; l cmin is the lower limit of the consistency score threshold ratio.
  • the device further comprises:
  • An integrated calculation module configured to perform a weighted calculation according to the accuracy matching score and a coherence score to obtain a comprehensive score of the current action; and when the source video image is played, the accuracy matching score is calculated each time The coherence score and the composite score are summed and averaged to obtain the average score of the whole process.
  • the invention provides an image processing method and device, which acquires a source video image and a target video image; and based on the source video image and the target video image, adopts a preset accuracy matching algorithm to respectively obtain an accuracy matching of the current motion Matching the accuracy of the previous action, obtaining the source video frame sequence number and the target video frame sequence number of the current action matching matching pair, and the source video frame sequence number and the target video frame sequence number of the matching accuracy of the previous action; Matching the source video frame number and the target video frame number of the pair according to the accuracy of the current action, and the source of the matching match of the accuracy of the previous action
  • the video frame number and the target video frame number are calculated by using a preset coherence algorithm to obtain the coherence result of the current action.
  • the level of action imitation provides an evaluation reference.
  • FIG. 1 is a schematic flow chart of a first embodiment of an image processing method according to the present invention.
  • FIG. 2 is a schematic flow chart of a second embodiment of an image processing method according to the present invention.
  • FIG. 3 is a schematic diagram of functional modules of a first embodiment of an image processing apparatus according to the present invention.
  • FIG. 4 is a schematic diagram of functional modules of a second embodiment of the image processing apparatus of the present invention.
  • a first embodiment of the present invention provides an image processing method, including:
  • Step S101 acquiring a source video image and a target video image
  • the source video image refers to a reference video image played by the user
  • the target video image is an action video image learned by the user in the reference source video image acquired by the camera or other camera module.
  • the video images collected by the source video and the camera can be displayed in real time, and the user's imitating action is calculated continuously, and corresponding scores are performed.
  • the score of the user learning video is composed of two parts, the motion matching accuracy and the motion consistency, and the motion coherence algorithm depends on the algorithm of the motion matching accuracy.
  • Step S102 Perform a preset accuracy matching algorithm based on the source video image and the target video image to obtain an accuracy matching pair of the current action and an accuracy matching pair of the previous action, and obtain an accuracy matching pair of the current action.
  • the source video frame sequence number and the target video frame sequence number, and the accuracy of the previous action match the source video frame sequence number and the target video frame sequence number;
  • the quantitative analysis of coherence depends not only on the sequence of video frames learned by the camera end users, but also on the basis of the sequence of source video frames.
  • the source video if an action is completed quickly, then this action in the camera should also be completed quickly.
  • Coherence is an important indicator of the level of dancers in learning. By calculating the dancer's coherence, the dancer's comprehensive score can be given.
  • a sequence of source video frames of the current action is acquired from the source video image, and a target video frame sequence of the current action is obtained from the target video image.
  • the key point information is extracted from the acquired frame sequence, and the key point group sequence of the source video and the key point group sequence of the target video are obtained.
  • the key points may include: a head, a hand, an abdomen, a foot, and the like of the human body.
  • the four angles are angles and heads formed by the head and the left hand
  • the difference between the corresponding angles in each frame of the key point group sequence of the target video and the key point group sequence of the target video is separately calculated, and the total difference of the four corners corresponding to each frame is obtained.
  • obtaining four angles of the first frame of the current action of the camera end and the source video end respectively calculating a difference between each angle corresponding to the camera end and the source video end, and summing the difference values of the four angles to obtain the The total difference of the first frame.
  • the total difference value of each of the multiple frames of the current action is obtained, and the minimum total difference is obtained among the multiple total differences, and the first action is obtained according to the minimum total difference.
  • the optimal accuracy matching pair assuming that the frame number of the current action corresponding to the minimum total difference is 90, then the An optimal accuracy matches the frame of the frame number 90 of the camera end and the frame number 90 of the source video end.
  • frame preprocessing is required for each frame in the obtained camera end and source video end images, and the frame preprocessing process includes background frame acquisition, image gray binarization and target frame extraction, and key Point identification, where:
  • a background image is taken from the camera side and the source video side, respectively.
  • the background of the camera can be used to notify the user to click on the button to take an environmental background image.
  • the background of the source video side will perform background extraction on the first frame in the video stream. Model the above two backgrounds for real-time updates in the later matching process.
  • Both the current frame and the background frame are grayed out, that is, the RGB color is converted into a gray value
  • the grayed out result of the current frame and the background frame is subtracted and subjected to noise elimination processing, and the difference is less than the set threshold and is considered as the background area, otherwise the target area is the human body area;
  • the corresponding matrix is established, the point value of the background area is 0, and the target area is 1.
  • Key points include the head, the left and right hands, the abdomen (central), and the left and right feet.
  • the matrix obtained above is further processed to obtain a minimum rectangle which can include the target area, which is called a target square.
  • the target square is divided into upper and lower partial regions, and the left and right feet are identified in the lower region, and other points are identified in the upper region.
  • Left and right feet In the lower area, the matrix is scanned from bottom to top row, and the first (target area) line appears on the left and right sides, that is, the lines of the left and right feet are respectively obtained, and the key points of the two feet are obtained.
  • Left and right hands In the upper area, from left to right, from right to left, one column and one column of scanning matrix, the first appearing 1 (target area) column, which is the column of the left and right hands, respectively, and then get the two key points.
  • Abdomen (center) The critical line of the upper and lower areas is taken as the line of the key point of the abdomen (central). The line is scanned, and the midpoint of the longest segment of the continuous 1 is used as the key point of the abdomen.
  • Header Scans the rectangular area consisting of several columns on the left and right sides of the column where the key points of the abdomen are located. The highest point in this area is the key point of the head.
  • the matching criterion for accuracy is the angle, that is, the angle between the head and the left and right hands, the angle between the abdomen (center) and the left and right feet, and the four angles can be calculated after each frame is recognized.
  • the accuracy matching pair of the previous action can be obtained, and the source video frame number and the target video frame number of the matching pair of the accuracy of the previous action are obtained.
  • Step S103 matching the source video frame sequence number and the target video frame sequence number of the pair according to the accuracy of the current action, and the source video frame sequence number and the target video frame sequence number of the previous action matching matching pair, and adopting preset consistency.
  • the algorithm calculates the coherence result of the current action.
  • the difference between the source video frame sequence number of the current action matching matching pair and the original video frame sequence number matching the accuracy of the previous action is obtained, and the source video key frame interval is obtained;
  • the continuity score of the current action is calculated.
  • n v i v - i vLst
  • n c i c - i cLst
  • the source video image and the target video image are obtained by using the foregoing solution.
  • a preset accuracy matching algorithm is used to obtain the current motion respectively.
  • the accuracy matching pair matches the accuracy of the previous action, and obtains the source video frame sequence number and the target video frame sequence number of the current action accuracy matching pair, and the source video frame sequence number and target of the accuracy of the previous action matching pair.
  • the sexual algorithm calculates the coherence result of the current action, and thus, the overall coordination effect of the user's dancing action can be obtained in real time through the coherence of the action.
  • the second embodiment of the present invention provides an image processing method. Based on the foregoing embodiment, the method further includes:
  • Step S104 performing weighting calculation according to the accuracy matching score and the consistency score to obtain a comprehensive score of the current action.
  • Step S105 when the source video image is played, the accuracy matching score, the coherence score and the comprehensive score calculated each time are summed and averaged, and the average score of the whole process is obtained.
  • the current video frame sequence accuracy and the coherency matching are all calculated, and the current sequence comprehensive score is obtained according to the weighted sum.
  • the above two algorithms are repeated until the completion of the play, and the accuracy score, the coherence score and the comprehensive score are respectively summed and averaged, and the average score of the whole process is obtained.
  • the coherence algorithm is called after each action matching algorithm is called.
  • the frame sequence is continuously acquired from the source video and the camera (the key points are extracted immediately after acquiring one frame), that is, the key point group sequences Q pv and Q pc are obtained , and the two sequences are synchronized, when the number is simultaneously The n match++ (settable) is reached, and the action matching calculation is performed.
  • the result is that the best matching pair and the matching score are obtained, and the matching pair numbers are i v and i c , respectively.
  • n v i v - i vLst
  • n c i c - i cLst
  • Each accuracy score, coherence score, and composite score are summed and averaged to obtain an average score for the entire process.
  • the source video image and the target video image are obtained by using the foregoing solution.
  • a preset accuracy matching algorithm is used to obtain an accuracy matching pair of the current action and a previous action.
  • the accuracy matching pair obtains the source video frame sequence number and the target video frame sequence number of the current action matching matching pair, and the source video frame sequence number and the target video frame sequence number of the previous action accuracy matching pair; according to the current action
  • the source video frame number and the target video frame number of the matching pair are matched, and the source video frame number and the target video frame number of the matching of the previous action are matched, and the coherence of the current action is calculated by using a preset coherence algorithm.
  • the overall coordination effect of the user's dancing action can be obtained in real time through the coherence of the action; in addition, the current comprehensive score, the average score, and the total evaluation score can be calculated, thereby providing an evaluation reference for the user's action effect.
  • the first embodiment of the present invention provides an image processing apparatus, including: an image acquisition module 201, a matching calculation module 202, and a coherency calculation module 203, wherein:
  • An image obtaining module 201 configured to acquire a source video image and a target video image
  • the matching calculation module 202 is configured to obtain a matching match between the accuracy matching pair of the current action and the accuracy of the previous action based on the source video image and the target video image, and obtain the current action
  • the consistency calculation module 203 is configured to match the source video frame sequence number and the target video frame sequence number of the pair according to the accuracy of the current action, and the source video frame sequence number and the target video frame sequence number of the matching match of the previous action,
  • the preset coherence algorithm calculates the coherent result of the current action.
  • the source video image refers to a reference video image played by the user
  • the target video image is an action video image learned by the user in the reference source video image acquired by the camera or other camera module.
  • the video images collected by the source video and the camera can be displayed in real time, and the user's imitating action is calculated continuously, and corresponding scores are performed.
  • the score of the user learning video is composed of two parts, the motion matching accuracy and the motion consistency, and the motion coherence algorithm depends on the algorithm of the motion matching accuracy.
  • the quantitative analysis of coherence depends not only on the sequence of video frames learned by the camera end users, but also on the basis of the sequence of source video frames.
  • the source video if an action is completed quickly, then this action in the camera should also be completed quickly.
  • Coherence is an important indicator of the level of dancers in learning. By calculating the dancer's coherence, the dancer's comprehensive score can be given.
  • a sequence of source video frames of the current action is acquired from the source video image, and a target video frame sequence of the current action is obtained from the target video image.
  • the key point information is extracted from the acquired frame sequence, and the key point group sequence of the source video and the key point group sequence of the target video are obtained.
  • the key points may include: a head, a hand, an abdomen, a foot, and the like of the human body.
  • the four angles are angles and heads formed by the head and the left hand
  • obtaining four angles of the first frame of the current action of the camera end and the source video end respectively calculating a difference between each angle corresponding to the camera end and the source video end, and summing the difference values of the four angles to obtain the The total difference of the first frame.
  • the total difference value of each of the multiple frames of the current action is obtained, and the minimum total difference is obtained among the multiple total differences, and the first action is obtained according to the minimum total difference.
  • the optimal accuracy matching pair if the frame number of the current action corresponding to the minimum total difference is 90, the first optimal accuracy matching pair is the frame number of the camera end 90 and the frame number of the source video end 90 Frame.
  • frame preprocessing is required for each frame in the obtained camera end and source video end images, and the frame preprocessing process includes background frame acquisition, image gray binarization and target frame extraction, and key Point identification, where:
  • a background image is taken from the camera side and the source video side, respectively.
  • the background of the camera can be used to notify the user to click on the button to take an environmental background image.
  • the background of the source video side will perform background extraction on the first frame in the video stream. Model the above two backgrounds for real-time updates in the later matching process.
  • Both the current frame and the background frame are grayed out, that is, the RGB color is converted into a gray value
  • the grayed out result of the current frame and the background frame is subtracted and subjected to noise elimination processing, and the difference is less than the set threshold and is considered as the background area, otherwise the target area is the human body area;
  • the corresponding matrix is established, the point value of the background area is 0, and the target area is 1.
  • Key points include the head, the left and right hands, the abdomen (central), and the left and right feet.
  • the matrix obtained above is further processed to obtain a minimum rectangle capable of containing the target area, Target square.
  • the target square is divided into upper and lower partial regions, and the left and right feet are identified in the lower region, and other points are identified in the upper region.
  • Left and right feet In the lower area, the matrix is scanned from bottom to top row, and the first (target area) line appears on the left and right sides, that is, the lines of the left and right feet are respectively obtained, and the key points of the two feet are obtained.
  • Left and right hands In the upper area, from left to right, from right to left, one column and one column of scanning matrix, the first appearing 1 (target area) column, which is the column of the left and right hands, respectively, and then get the two key points.
  • Abdomen (center) The critical line of the upper and lower areas is taken as the line of the key point of the abdomen (central). The line is scanned, and the midpoint of the longest segment of the continuous 1 is used as the key point of the abdomen.
  • Header Scans the rectangular area consisting of several columns on the left and right sides of the column where the key points of the abdomen are located. The highest point in this area is the key point of the head.
  • the matching criterion for accuracy is the angle, that is, the angle between the head and the left and right hands, the angle between the abdomen (center) and the left and right feet, and the four angles can be calculated after each frame is recognized.
  • the accuracy matching pair of the previous action can be obtained, and the source video frame number and the target video frame number of the matching pair of the accuracy of the previous action are obtained.
  • the source video frame sequence number and the target video frame sequence number, and the accuracy of the previous action match the source video frame sequence number and the target video frame sequence number, and the coherence result of the current action is calculated by using a preset coherence algorithm.
  • the difference between the source video frame sequence number of the current action matching matching pair and the original video frame sequence number matching the accuracy of the previous action is obtained, and the source video key frame interval is obtained;
  • the continuity score of the current action is calculated.
  • n v i v - i vLst
  • n c i c - i cLst
  • the source video image and the target video image are obtained by using the foregoing solution.
  • a preset accuracy matching algorithm is used to obtain an accuracy matching pair and a previous action of the current action respectively.
  • the accuracy matching pair obtains the source video frame sequence number and the target video frame sequence number of the current action matching match pair, and the source video frame sequence number and the target video frame sequence number of the previous action accuracy matching pair; according to the current action
  • the accuracy matches the source video frame number and the target video frame number
  • the accuracy of the previous action matches the source video frame number and the target video frame number.
  • the preset coherence algorithm is used to calculate the coherence of the current action. As a result, the overall coordination effect of the user's dancing action can be obtained in real time through the continuity of the action.
  • the second embodiment of the present invention provides an image processing apparatus. Based on the foregoing embodiment, the method further includes:
  • the comprehensive calculation module 204 is configured to perform weighting calculation according to the accuracy matching score and the consistency score to obtain a comprehensive score of the current action; and when the source video image is played, the accuracy of each calculation is matched.
  • the value, coherence score, and composite score are summed and averaged, respectively, resulting in an average score for the entire process.
  • the current video frame sequence accuracy and the coherency matching are all calculated, and the current sequence comprehensive score is obtained according to the weighted sum.
  • the above two algorithms are repeated until the completion of the play, and the accuracy score, the coherence score and the comprehensive score are respectively summed and averaged, and the average score of the whole process is obtained.
  • the coherence algorithm is called after each action matching algorithm is called.
  • the frame sequence is continuously acquired from the source video and the camera (the key points are extracted immediately after acquiring one frame), that is, the key point group sequences Q pv and Q pc are obtained , and the two sequences are synchronized, when the number is simultaneously The n match++ (settable) is reached, and the action matching calculation is performed.
  • the result is that the best matching pair and the matching score are obtained, and the matching pair numbers are i v and i c , respectively.
  • n v i v - i vLst
  • n c i c - i cLst
  • Each accuracy score, coherence score, and composite score are summed and averaged to obtain an average score for the entire process.
  • the source video image and the target video image are obtained by using the foregoing solution.
  • a preset accuracy matching algorithm is used to obtain an accuracy matching pair and a previous action of the current action respectively.
  • the accuracy matching pair obtains the source video frame sequence number and the target video frame sequence number of the current action matching match pair, and the source video frame sequence number and the target video frame sequence number of the previous action accuracy matching pair; according to the current action Accuracy matches the source video of the pair
  • the frame number and the target video frame number, and the accuracy of the previous action match the source video frame number and the target video frame number, and use the preset coherence algorithm to calculate the coherence result of the current action, thereby
  • the consistency of the action obtains the overall coordination effect of the user's dancing action in real time; in addition, the current comprehensive score, the average score and the total evaluation score can be calculated, and then the evaluation of the user's action effect is provided.
  • the application field of the solution of the embodiment of the present invention can be in the field of learning and action behavior identification of action categories such as game scoring, dance, and the like involved in consumer electronics such as game machines, PCs, and TVs.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Image Analysis (AREA)

Abstract

公开了一种图像处理方法及装置,其方法包括:获取源视频图像和目标视频图像;基于源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;根据当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果,由此可以通过动作的连贯性实时得到用户模仿动作的整体协调效果,并能够对用户的动作模仿水平提供评价参考。

Description

图像处理方法及装置 技术领域
本发明涉及图像处理技术领域,尤其涉及一种图像处理方法及装置。
背景技术
目前,市面上有很多跳舞软件可以让用户参照跳舞节目的动作进行学习,例如,用户可以根据电视画面的跳舞箭头以及跳舞毯上的方向箭头,配合音乐进行跳舞运动。但是,现有的跳舞软件,只能指示用户当前动作是否正确,以及用户跳舞运动的最终得分,无法实时的向用户显示当前动作和上一个动作是否连贯,用户的最终得分中也不包括用户连贯的分数。
发明内容
本发明提供一种图像处理方法及装置,主要目的在于解决如何实时获得用户当前动作的连贯性的技术问题。
为实现上述目的,本发明提供一种图像处理方法,包括:
获取源视频图像和目标视频图像;
基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;
根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果。
优选地,所述基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,获取当前动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号的步骤包括:
从所述源视频图像中获取当前动作的源视频帧序列,从所述目标视频图像中获取当前动作的目标视频帧序列;
从获取的帧序列中提取关键点信息,得到源视频的关键点组序列和目标 视频的关键点组序列;
对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,得到当前动作的准确性匹配对,并获取当前动作的准确性匹配对的源视频帧序号和目标视频帧序号。
优选地,所述关键点包括:人体的头部、手部、腹部、脚部;所述对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,得到当前动作的准确性匹配对的步骤包括:
获取所述源视频的关键点组序列以及目标视频的关键点组序列的多个帧中每一帧基于所述关键点的四个角度,所述四个角度为头与左手形成的角度、头与右手形成的角度、腹部与左脚形成的角度、腹部与右脚形成的角度;
分别计算所述源视频的关键点组序列与目标视频的关键点组序列的每一帧中对应角之差,并得到每一帧对应的四个角的总差值;
获取多个帧的总差值中的最小总差值;
根据所述最小总差值获取所述当前动作的准确性匹配对,及准确性匹配分值。
优选地,所述根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果的步骤包括:
获取所述当前动作的准确性匹配对的源视频帧序号与前一动作的准确性匹配对的源视频帧序号之间的差值,得到源视频关键帧间隔;
获取所述当前动作的准确性匹配对的目标视频帧序号与前一动作的准确性匹配对的目标视频帧序号之间的差值,得到目标视频关键帧间隔;
根据所述源视频关键帧间隔、目标视频关键帧间隔,以及预设的连贯性公式,计算得到当前动作的连贯性得分;具体计算公式如下:
关键帧间隔:nv=iv-ivLst,nc=ic-icLst
则当前连贯度公式:
Figure PCTCN2015090279-appb-000001
当前连贯性得分公式:sc=100·(lcmin+(1-lcmin)·lc),其中,lc为连贯度;sc为当前连贯性得分;iv为当前最佳匹配对中源视频帧序号;ic为当前最佳匹配对中目标视频帧序号;ivLst为前一最佳匹配对中源视频帧序号;icLst为前一最佳 匹配对中目标视频帧序号;nv为当前源视频关键帧与前一关键帧的帧间隔数;nc为当前目标视频关键帧与前一关键帧的帧间隔数;lcmin为连贯性得分下限阈值比例。
优选地,该方法还包括:
根据所述准确性匹配分值以及连贯性得分,进行加权计算得到当前动作的综合得分。
优选地,该方法还包括:
当所述源视频图像播放完毕,将每次计算得到的准确性匹配分值、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
本发明实施例还提出一种图像处理装置,包括:
图像获取模块,用于获取源视频图像和目标视频图像;
匹配计算模块,用于基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;
连贯性计算模块,用于根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果。
优选地,所述匹配计算模块,还用于从所述源视频图像中获取当前动作的源视频帧序列,从所述目标视频图像中获取当前动作的目标视频帧序列;从获取的帧序列中提取关键点信息,得到源视频的关键点组序列和目标视频的关键点组序列;对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,得到当前动作的准确性匹配对,并获取当前动作的准确性匹配对的源视频帧序号和目标视频帧序号。
优选地,所述关键点包括:人体的头部、手部、腹部、脚部;
所述匹配计算模块,还用于获取所述源视频的关键点组序列以及目标视频的关键点组序列的多个帧中每一帧基于所述关键点的四个角度,所述四个角度为头与左手形成的角度、头与右手形成的角度、腹部与左脚形成的角度、腹部与右脚形成的角度;分别计算所述源视频的关键点组序列与目标视频的 关键点组序列的每一帧中对应角之差,并得到每一帧对应的四个角的总差值;获取多个帧的总差值中的最小总差值;根据所述最小总差值获取所述当前动作的准确性匹配对,及准确性匹配分值。
优选地,所述连贯性计算模块,还用于获取所述当前动作的准确性匹配对的源视频帧序号与前一动作的准确性匹配对的源视频帧序号之间的差值,得到源视频关键帧间隔;获取所述当前动作的准确性匹配对的目标视频帧序号与前一动作的准确性匹配对的目标视频帧序号之间的差值,得到目标视频关键帧间隔;根据所述源视频关键帧间隔、目标视频关键帧间隔,以及预设的连贯性公式,计算得到当前动作的连贯性得分;具体计算公式如下:
关键帧间隔:nv=iv-ivLst,nc=ic-icLst
则当前连贯度公式:
Figure PCTCN2015090279-appb-000002
当前连贯性得分公式:sc=100·(lcmin+(1-lcmin)·lc),其中,lc为连贯度;sc为当前连贯性得分;iv为当前最佳匹配对中源视频帧序号;ic为当前最佳匹配对中目标视频帧序号;ivLst为前一最佳匹配对中源视频帧序号;icLst为前一最佳匹配对中目标视频帧序号;nv为当前源视频关键帧与前一关键帧的帧间隔数;nc为当前目标视频关键帧与前一关键帧的帧间隔数;lcmin为连贯性得分下限阈值比例。
优选地,该装置还包括:
综合计算模块,用于根据所述准确性匹配分值以及连贯性得分,进行加权计算得到当前动作的综合得分;以及当所述源视频图像播放完毕,将每次计算得到的准确性匹配分值、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
本发明提供的一种图像处理方法及装置,通过获取源视频图像和目标视频图像;基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源 视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果,由此,可以通过动作的连贯性实时得到用户模仿动作的整体协调效果,并能够对用户的动作模仿水平提供评价参考。
附图说明
图1是本发明图像处理方法第一实施例的流程示意图;
图2是本发明图像处理方法第二实施例的流程示意图;
图3是本发明图像处理装置第一实施例的功能模块示意图;
图4是本发明图像处理装置第二实施例的功能模块示意图。
本发明目的的实现、功能特点及优点将结合实施例,参照附图做进一步说明。
具体实施方式
应当理解,此处所描述的具体实施例仅仅用以解释本发明,并不用于限定本发明。
如图1所示,本发明第一实施例提出一种图像处理方法,包括:
步骤S101,获取源视频图像和目标视频图像;
本实施例中,源视频图像是指用户播放的参考视频图像,目标视频图像是通过摄像头或者其他摄像模块获取的用户参照源视频图像中的动作而学习的动作视频图像。
本实施例方案,可以实时显示源视频和摄像头采集的视频图像,对用户的模仿动作进行连贯性计算,并进行相应的评分。
其中,用户学习视频的评分由两部分构成,动作匹配准确性和动作连贯性,动作连贯性的算法依赖于动作匹配准确性的算法。
步骤S102,基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;
其中,连贯性的定量分析,不仅依赖摄像头端用户学习的视频帧序列,还要以源视频帧序列为基准做评判。源视频中,若一个动作很快完成,那么摄像头中这个动作也应快速完成。
动作连贯性是衡量学习的舞者水平的一个重要指标,通过计算舞者的动作连贯性,可以给出舞者的跳舞综合分数。
具体地,在得到源视频图像和目标视频图像后,从所述源视频图像中获取当前动作的源视频帧序列,从所述目标视频图像中获取当前动作的目标视频帧序列。
之后,从获取的帧序列中提取关键点信息,得到源视频的关键点组序列和目标视频的关键点组序列。
对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,得到当前动作的准确性匹配对,并获取当前动作的准确性匹配对的源视频帧序号和目标视频帧序号。
其中,所述关键点可以包括:人体的头部、手部、腹部、脚部等。
在对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,可以采用如下方案:
获取所述源视频的关键点组序列以及目标视频的关键点组序列的多个帧中每一帧基于所述关键点的四个角度,所述四个角度为头与左手形成的角度、头与右手形成的角度、腹部与左脚形成的角度、腹部与右脚形成的角度;
分别计算所述源视频的关键点组序列与目标视频的关键点组序列的每一帧中对应角之差,并得到每一帧对应的四个角的总差值。
获取上述多个帧的总差值中的最小总差值;根据所述最小总差值获取所述当前动作的准确性匹配对,及准确性匹配分值,进而得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号。
具体的,获取摄像头端以及源视频端的当前动作的第一帧的4个角度,分别计算摄像头端以及源视频端对应的每个角度之差,将4个角度的差值进行求和得到所述第一帧总的差值。
以此原理,获取当前动作的多个帧中的每个帧的总差值,并在多个总差值中获取最小总差值,根据所述最小总差值获取所述当前动作的第一最优准确性匹配对,假设所述最小总差值对应的当前动作的帧序号为90,则所述第 一最优准确性匹配对为摄像头端的帧序号90的帧与源视频端的帧序号90的帧。
上述提取关键点的过程中,需要对获取的摄像头端和源视频端图像中的每一帧做帧预处理,帧预处理过程包括背景帧获取、图像灰度二值化和目标帧提取、关键点识别,其中:
背景帧获取:
初始时,分别从摄像头端和源视频端获取一张背景图片。摄像头端的背景可通过与用户交互,通知用户点击按钮拍摄一张环境背景图片。源视频端的背景将对视频流中的第一帧进行背景提取。对上述两个背景进行建模,用于稍后匹配过程中的实时更新。
图像灰度二值化和目标提取:
对于摄像头端和源视频端的处理相同,对每一帧做如下处理:
将当前帧与背景帧都进行灰度化处理,也就是将RGB色彩转化成灰度值;
将当前帧与背景帧的灰度化后的结果,进行相减并进行消除噪声处理,差值小于设定阈值的认为是背景区域,否则为目标区域,即人体区域;
建立相应矩阵,背景区域的点值为0,目标区域的为1。
完成上述过程,将处理得到的矩阵作为输入供关键点识别及后续过程。
关键点识别:
关键点包括头、左右手、腹部(中央)和左右脚共6个点。
将上述得到的矩阵进一步处理,获得能包含目标区域的最小长方形,称目标方形。将目标方形分为上下两部分区域,在下区域识别出左右脚,其他点在上区域识别。
左右脚:在下区域,从下到上一行一行扫描矩阵,左右两边最先出现1(目标区域)的行,即分别左右脚所在行,进而得到两脚关键点。
左右手:在上区域,分别从左至右、从右至左,一列一列扫描矩阵,最先出现的1(目标区域)的列,即分别为左右手所在列,进而得到两手关键点。
腹部(中央):将上下区域的临界行,作为腹部(中央)关键点所在行,扫描该行,将连续1的最长段的中点作为腹部关键点。
头:扫描腹部关键点所在列左右指定几列所组成的长方形区域,此区域内最高点为头部关键点。
准确性的匹配标准是角度,即头分别与左右手的角度、腹部(中央)分别与左右脚的角度,每一帧识别出关键点后即可计算这4个角度。
基于上述匹配原理,可以得到前一动作的准确性匹配对,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号。
步骤S103,根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果。
具体地,获取所述当前动作的准确性匹配对的源视频帧序号与前一动作的准确性匹配对的源视频帧序号之间的差值,得到源视频关键帧间隔;
获取所述当前动作的准确性匹配对的目标视频帧序号与前一动作的准确性匹配对的目标视频帧序号之间的差值,得到目标视频关键帧间隔;
根据所述源视频关键帧间隔、目标视频关键帧间隔,以及预设的连贯性公式,计算得到当前动作的连贯性得分。
具体的计算公式可以如下:
采用以下公式计算关键帧间隔:
关键帧间隔:nv=iv-ivLst,nc=ic-icLst;  (1)
则当前连贯度公式:
Figure PCTCN2015090279-appb-000003
当前连贯性得分公式:sc=100·(lcmin+(1-lcmin)·lc);  (3)
其中,lc为连贯度,可以取值[0,1];sc当前连贯性得分,可以取值[0,100];iv为当前最佳匹配对中源视频帧序号;ic为当前最佳匹配对中摄像头帧序号;ivLst为前一最佳匹配对中源视频帧序号;icLst为前一最佳匹配对中摄像头帧序号(目标视频帧序号);nv为当前源视频关键帧与前一关键帧的帧间隔数;nc为当前摄像头关键帧(目标视频关键帧)与前一关键帧的帧间隔数;lcmin连贯性得分下限阈值比例[0,1),kc连贯性得分比重,(0,1)
本实施例通过上述方案,获取源视频图像和目标视频图像;基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作 的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果,由此,可以通过动作的连贯性实时得到用户跳舞动作的整体协调效果。
如图2所示,本发明第二实施例提出一种图像处理方法,基于上述实施例,还包括:
步骤S104,根据所述准确性匹配分值以及连贯性得分,进行加权计算得到当前动作的综合得分。
步骤S105,当所述源视频图像播放完毕,将每次计算得到的准确性匹配分值、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
具体的,当前视频帧序列准确性和连贯性匹配都计算完毕,根据加权和得出当前序列综合得分。情况暂存序列,重复上述两个算法,直至播放完毕,将每次准确性得分、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
具体算法如下:
其中涉及的参数如下表1所示:
Figure PCTCN2015090279-appb-000004
Figure PCTCN2015090279-appb-000005
表1
每次动作匹配算法调用后,进而调用连贯性算法。
在动作准确性匹配算法中,从源视频和摄像头不断获取帧序列(每获取一帧立即提取关键点),即获得关键点组序列Qpv和Qpc,两个序列同步进行,当个数同时达到nmatch++(可设置),进行动作匹配性计算,结果是得到了最佳匹配对和匹配得分,匹配对序号分别为iv、ic
然后,采用以下公式计算关键帧间隔:
关键帧间隔:nv=iv-ivLst,nc=ic-icLst;  (1)
则当前连贯度公式:
Figure PCTCN2015090279-appb-000006
当前连贯性得分公式:sc=100·(lcmin+(1-lcmin)·lc);  (3)
最后,可以计算当前综合得分和评价分值:
当前综合得分:s=sm·(1-kc)+sc·kc;  (4)
平均连贯性得分:
Figure PCTCN2015090279-appb-000007
平均综合得分:sav=smAv·(1-kc)+scAv·kc;  (6)
至此,当前序列连贯性计算完毕,同时得出综合得分,清空Qpv和Qpc,重复上述过程,直至播放完毕。
将每次准确性得分、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
上述公式中涉及的各参数可以参照上述表1。
本实施例通过上述方案,获取源视频图像和目标视频图像;基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果,由此,可以通过动作的连贯性实时得到用户跳舞动作的整体协调效果;此外,还可以计算当前综合得分、平均分值和总的评价分值,进而对用户的动作效果提供评价参考。
对应地,提出本发明的图像处理装置。
如图3所示,本发明第一实施例提出一种图像处理装置,包括:图像获取模块201、匹配计算模块202以及连贯性计算模块203,其中:
图像获取模块201,用于获取源视频图像和目标视频图像;
匹配计算模块202,用于基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;
连贯性计算模块203,用于根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果。
本实施例中,源视频图像是指用户播放的参考视频图像,目标视频图像是通过摄像头或者其他摄像模块获取的用户参照源视频图像中的动作而学习的动作视频图像。
本实施例方案,可以实时显示源视频和摄像头采集的视频图像,对用户的模仿动作进行连贯性计算,并进行相应的评分。
其中,用户学习视频的评分由两部分构成,动作匹配准确性和动作连贯性,动作连贯性的算法依赖于动作匹配准确性的算法。
其中,连贯性的定量分析,不仅依赖摄像头端用户学习的视频帧序列,还要以源视频帧序列为基准做评判。源视频中,若一个动作很快完成,那么摄像头中这个动作也应快速完成。
动作连贯性是衡量学习的舞者水平的一个重要指标,通过计算舞者的动作连贯性,可以给出舞者的跳舞综合分数。
具体地,在得到源视频图像和目标视频图像后,从所述源视频图像中获取当前动作的源视频帧序列,从所述目标视频图像中获取当前动作的目标视频帧序列。
之后,从获取的帧序列中提取关键点信息,得到源视频的关键点组序列和目标视频的关键点组序列。
对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,得到当前动作的准确性匹配对,并获取当前动作的准确性匹配对的源视频帧序号和目标视频帧序号。
其中,所述关键点可以包括:人体的头部、手部、腹部、脚部等。
在对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,可以采用如下方案:
获取所述源视频的关键点组序列以及目标视频的关键点组序列的多个帧中每一帧基于所述关键点的四个角度,所述四个角度为头与左手形成的角度、头与右手形成的角度、腹部与左脚形成的角度、腹部与右脚形成的角度;
分别计算所述源视频的关键点组序列与目标视频的关键点组序列的每一 帧中对应角之差,并得到每一帧对应的四个角的总差值。
获取上述多个帧的总差值中的最小总差值;根据所述最小总差值获取所述当前动作的准确性匹配对,及准确性匹配分值,进而得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号。
具体的,获取摄像头端以及源视频端的当前动作的第一帧的4个角度,分别计算摄像头端以及源视频端对应的每个角度之差,将4个角度的差值进行求和得到所述第一帧总的差值。
以此原理,获取当前动作的多个帧中的每个帧的总差值,并在多个总差值中获取最小总差值,根据所述最小总差值获取所述当前动作的第一最优准确性匹配对,假设所述最小总差值对应的当前动作的帧序号为90,则所述第一最优准确性匹配对为摄像头端的帧序号90的帧与源视频端的帧序号90的帧。
上述提取关键点的过程中,需要对获取的摄像头端和源视频端图像中的每一帧做帧预处理,帧预处理过程包括背景帧获取、图像灰度二值化和目标帧提取、关键点识别,其中:
背景帧获取:
初始时,分别从摄像头端和源视频端获取一张背景图片。摄像头端的背景可通过与用户交互,通知用户点击按钮拍摄一张环境背景图片。源视频端的背景将对视频流中的第一帧进行背景提取。对上述两个背景进行建模,用于稍后匹配过程中的实时更新。
图像灰度二值化和目标提取:
对于摄像头端和源视频端的处理相同,对每一帧做如下处理:
将当前帧与背景帧都进行灰度化处理,也就是将RGB色彩转化成灰度值;
将当前帧与背景帧的灰度化后的结果,进行相减并进行消除噪声处理,差值小于设定阈值的认为是背景区域,否则为目标区域,即人体区域;
建立相应矩阵,背景区域的点值为0,目标区域的为1。
完成上述过程,将处理得到的矩阵作为输入供关键点识别及后续过程。
关键点识别:
关键点包括头、左右手、腹部(中央)和左右脚共6个点。
将上述得到的矩阵进一步处理,获得能包含目标区域的最小长方形,称 目标方形。将目标方形分为上下两部分区域,在下区域识别出左右脚,其他点在上区域识别。
左右脚:在下区域,从下到上一行一行扫描矩阵,左右两边最先出现1(目标区域)的行,即分别左右脚所在行,进而得到两脚关键点。
左右手:在上区域,分别从左至右、从右至左,一列一列扫描矩阵,最先出现的1(目标区域)的列,即分别为左右手所在列,进而得到两手关键点。
腹部(中央):将上下区域的临界行,作为腹部(中央)关键点所在行,扫描该行,将连续1的最长段的中点作为腹部关键点。
头:扫描腹部关键点所在列左右指定几列所组成的长方形区域,此区域内最高点为头部关键点。
准确性的匹配标准是角度,即头分别与左右手的角度、腹部(中央)分别与左右脚的角度,每一帧识别出关键点后即可计算这4个角度。
基于上述匹配原理,可以得到前一动作的准确性匹配对,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号。
在得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号后,根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果。
具体地,获取所述当前动作的准确性匹配对的源视频帧序号与前一动作的准确性匹配对的源视频帧序号之间的差值,得到源视频关键帧间隔;
获取所述当前动作的准确性匹配对的目标视频帧序号与前一动作的准确性匹配对的目标视频帧序号之间的差值,得到目标视频关键帧间隔;
根据所述源视频关键帧间隔、目标视频关键帧间隔,以及预设的连贯性公式,计算得到当前动作的连贯性得分。
具体的计算公式可以如下:
采用以下公式计算关键帧间隔:
关键帧间隔:nv=iv-ivLst,nc=ic-icLst;  (1)
则当前连贯度公式:
Figure PCTCN2015090279-appb-000008
当前连贯性得分公式:sc=100·(lcmin+(1-lcmin)·lc);  (3)
其中,lc为连贯度,可以取值[0,1];sc当前连贯性得分,可以取值[0,100];iv为当前最佳匹配对中源视频帧序号;ic为当前最佳匹配对中摄像头帧序号;ivLst为前一最佳匹配对中源视频帧序号;icLst为前一最佳匹配对中摄像头帧序号;nv为当前源视频关键帧与前一关键帧的帧间隔数;nc为当前摄像头关键帧与前一关键帧的帧间隔数;lcmin连贯性得分下限阈值比例[0,1),kc连贯性得分比重,(0,1)
本实施例通过上述方案,通过获取源视频图像和目标视频图像;基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果,由此,可以通过动作的连贯性实时得到用户跳舞动作的整体协调效果。
如图4所示,本发明第二实施例提出一种图像处理装置,基于上述实施例,还包括:
综合计算模块204,用于根据所述准确性匹配分值以及连贯性得分,进行加权计算得到当前动作的综合得分;以及当所述源视频图像播放完毕,将每次计算得到的准确性匹配分值、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
具体的,当前视频帧序列准确性和连贯性匹配都计算完毕,根据加权和得出当前序列综合得分。情况暂存序列,重复上述两个算法,直至播放完毕,将每次准确性得分、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
具体算法如下:
其中涉及的参数如上表1所示。
每次动作匹配算法调用后,进而调用连贯性算法。
在动作准确性匹配算法中,从源视频和摄像头不断获取帧序列(每获取一帧立即提取关键点),即获得关键点组序列Qpv和Qpc,两个序列同步进行,当个数同时达到nmatch++(可设置),进行动作匹配性计算,结果是得到了最佳匹配对和匹配得分,匹配对序号分别为iv、ic
然后,采用以下公式计算关键帧间隔:
关键帧间隔:nv=iv-ivLst,nc=ic-icLst;  (1)
则当前连贯度公式:
Figure PCTCN2015090279-appb-000009
当前连贯性得分公式:sc=100·(lcmin+(1-lcmin)·lc);  (3)
最后,可以计算当前综合得分和评价分值:
当前综合得分:s=sm·(1-kc)+sc·kc;  (4)
平均连贯性得分:
Figure PCTCN2015090279-appb-000010
平均综合得分:sav=smAv·(1-kc)+scAv·kc;  (6)
至此,当前序列连贯性计算完毕,同时得出综合得分,清空Qpv和Qpc,重复上述过程,直至播放完毕。
将每次准确性得分、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
上述公式中涉及的各参数可以参照上述表1。
本实施例通过上述方案,通过获取源视频图像和目标视频图像;基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;根据所述当前动作的准确性匹配对的源视频 帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果,由此,可以通过动作的连贯性实时得到用户跳舞动作的整体协调效果;此外,还可以计算当前综合得分、平均分值和总的评价分值,进而对用户的动作效果提供评价参考。
本发明实施例方案的应用领域可以在游戏机、PC、TV等消费电子涉及到的游戏打分、舞蹈等动作类的学习及动作行为鉴定领域。
以上仅为本发明的优选实施例,并非因此限制本发明的专利范围,凡是利用本发明说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本发明的专利保护范围内。

Claims (20)

  1. 一种图像处理方法,其特征在于,包括:
    获取源视频图像和目标视频图像;
    基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;
    根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果。
  2. 根据权利要求1所述的方法,其特征在于,所述基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,获取当前动作的准确性匹配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号的步骤包括:
    从所述源视频图像中获取当前动作的源视频帧序列,从所述目标视频图像中获取当前动作的目标视频帧序列;
    从获取的帧序列中提取关键点信息,得到源视频的关键点组序列和目标视频的关键点组序列;
    对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,得到当前动作的准确性匹配对,并获取当前动作的准确性匹配对的源视频帧序号和目标视频帧序号。
  3. 根据权利要求2所述的方法,其特征在于,所述从获取的帧序列中提取关键点信息的过程中包括:
    对获取的摄像头端和源视频端图像中的每一帧做帧预处理,帧预处理过程包括背景帧获取、图像灰度二值化和目标帧提取、关键点识别。
  4. 根据权利要求3所述的方法,其特征在于,所述背景帧获取包括:
    初始时,分别从摄像头端和源视频端获取一张背景图片,其中:摄像头端的背景通过与用户交互,通知用户点击按钮拍摄一张环境背景图片;源视频端的背景将对视频流中的第一帧进行背景提取。
  5. 根据权利要求3所述的方法,其特征在于,所述图像灰度二值化和目标帧提取的步骤包括:
    将当前帧与背景帧都进行灰度化处理;
    将当前帧与背景帧的灰度化后的结果,进行相减并进行消除噪声处理,差值小于设定阈值的认为是背景区域,否则为目标区域;
    建立相应矩阵,背景区域的点值为0,目标区域的为1;
    完成上述过程,将处理得到的矩阵作为输入供关键点识别。
  6. 根据权利要求2所述的方法,其特征在于,所述关键点包括:人体的头部、手部、腹部、脚部;所述对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,得到当前动作的准确性匹配对的步骤包括:
    获取所述源视频的关键点组序列以及目标视频的关键点组序列的多个帧中每一帧基于所述关键点的四个角度,所述四个角度为头与左手形成的角度、头与右手形成的角度、腹部与左脚形成的角度、腹部与右脚形成的角度;
    分别计算所述源视频的关键点组序列与目标视频的关键点组序列的每一帧中对应角之差,并得到每一帧对应的四个角的总差值;
    获取多个帧的总差值中的最小总差值;
    根据所述最小总差值获取所述当前动作的准确性匹配对,及准确性匹配分值。
  7. 根据权利要求6所述的方法,其特征在于,所述根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果的步骤包括:
    获取所述当前动作的准确性匹配对的源视频帧序号与前一动作的准确性匹配对的源视频帧序号之间的差值,得到源视频关键帧间隔;
    获取所述当前动作的准确性匹配对的目标视频帧序号与前一动作的准确性匹配对的目标视频帧序号之间的差值,得到目标视频关键帧间隔;
    根据所述源视频关键帧间隔、目标视频关键帧间隔,以及预设的连贯性公式,计算得到当前动作的连贯性得分。
  8. 根据权利要求7所述的方法,其特征在于,所述根据所述源视频关键帧间隔、目标视频关键帧间隔,以及预设的连贯性公式,计算得到当前动作的连贯性得分的具体计算公式如下:
    关键帧间隔:nv=iv-ivLst,nc=ic-icLst
    则当前连贯度公式:
    Figure PCTCN2015090279-appb-100001
    当前连贯性得分公式:sc=100·(lcmin+(1-lcmin)·lc),其中,lc为连贯度;sc为当前连贯性得分;iv为当前最佳匹配对中源视频帧序号;ic为当前最佳匹配对中目标视频帧序号;icLst为前一最佳匹配对中源视频帧序号;icLst为前一最佳匹配对中目标视频帧序号;nv为当前源视频关键帧与前一关键帧的帧间隔数;nc为当前目标视频关键帧与前一关键帧的帧间隔数;lcmin为连贯性得分下限阈值比例。
  9. 根据权利要求8所述的方法,其特征在于,还包括:
    根据所述准确性匹配分值以及连贯性得分,进行加权计算得到当前动作的综合得分。
  10. 根据权利要求9所述的方法,其特征在于,还包括:
    当所述源视频图像播放完毕,将每次计算得到的准确性匹配分值、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
  11. 一种图像处理装置,其特征在于,包括:
    图像获取模块,用于获取源视频图像和目标视频图像;
    匹配计算模块,用于基于所述源视频图像和目标视频图像,采用预设的准确性匹配算法,分别获取当前动作的准确性匹配对和前一动作的准确性匹 配对,得到当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号;
    连贯性计算模块,用于根据所述当前动作的准确性匹配对的源视频帧序号和目标视频帧序号,以及前一动作的准确性匹配对的源视频帧序号和目标视频帧序号,采用预设的连贯性算法,计算得到当前动作的连贯性结果。
  12. 根据权利要求11所述的装置,其特征在于,
    所述匹配计算模块,还用于从所述源视频图像中获取当前动作的源视频帧序列,从所述目标视频图像中获取当前动作的目标视频帧序列;从获取的帧序列中提取关键点信息,得到源视频的关键点组序列和目标视频的关键点组序列;对所述源视频的关键点组序列和目标视频的关键点组序列中对应的关键点进行匹配,得到当前动作的准确性匹配对,并获取当前动作的准确性匹配对的源视频帧序号和目标视频帧序号。
  13. 根据权利要求11所述的装置,其特征在于,
    所述匹配计算模块,还用于对获取的摄像头端和源视频端图像中的每一帧做帧预处理,帧预处理过程包括背景帧获取、图像灰度二值化和目标帧提取、关键点识别。
  14. 根据权利要求11所述的装置,其特征在于,所述背景帧获取包括:
    初始时,分别从摄像头端和源视频端获取一张背景图片,其中:摄像头端的背景通过与用户交互,通知用户点击按钮拍摄一张环境背景图片;源视频端的背景将对视频流中的第一帧进行背景提取。
  15. 根据权利要求11所述的装置,其特征在于,所述图像灰度二值化和目标帧提取包括:
    将当前帧与背景帧都进行灰度化处理;
    将当前帧与背景帧的灰度化后的结果,进行相减并进行消除噪声处理,差值小于设定阈值的认为是背景区域,否则为目标区域;
    建立相应矩阵,背景区域的点值为0,目标区域的为1;
    完成上述过程,将处理得到的矩阵作为输入供关键点识别。
  16. 根据权利要求12所述的装置,其特征在于,所述关键点包括:人体的头部、手部、腹部、脚部;
    所述匹配计算模块,还用于获取所述源视频的关键点组序列以及目标视频的关键点组序列的多个帧中每一帧基于所述关键点的四个角度,所述四个角度为头与左手形成的角度、头与右手形成的角度、腹部与左脚形成的角度、腹部与右脚形成的角度;分别计算所述源视频的关键点组序列与目标视频的关键点组序列的每一帧中对应角之差,并得到每一帧对应的四个角的总差值;获取多个帧的总差值中的最小总差值;根据所述最小总差值获取所述当前动作的准确性匹配对,及准确性匹配分值。
  17. 根据权利要求16所述的装置,其特征在于,
    所述连贯性计算模块,还用于获取所述当前动作的准确性匹配对的源视频帧序号与前一动作的准确性匹配对的源视频帧序号之间的差值,得到源视频关键帧间隔;获取所述当前动作的准确性匹配对的目标视频帧序号与前一动作的准确性匹配对的目标视频帧序号之间的差值,得到目标视频关键帧间隔;根据所述源视频关键帧间隔、目标视频关键帧间隔,以及预设的连贯性公式,计算得到当前动作的连贯性得分。
  18. 根据权利要求17所述的装置,其特征在于,所述连贯性计算模块根据所述源视频关键帧间隔、目标视频关键帧间隔,以及预设的连贯性公式,计算得到当前动作的连贯性得分的具体计算公式如下:
    关键帧间隔:nv=iv-ivLst,nc=ic-icLst
    则当前连贯度公式:
    Figure PCTCN2015090279-appb-100002
    当前连贯性得分公式:sc=100·(lcmin+(1-lcmin)·lc),其中,lc为连贯度;sc为当前连贯性得分;iv为当前最佳匹配对中源视频帧序号;ic为当前最佳匹配对中目标视频帧序号;ivLst为前一最佳匹配对中源视频帧序号;icLst为前一最佳匹配对中目标视频帧序号;nv为当前源视频关键帧与前一关键帧的帧间隔数; nc为当前目标视频关键帧与前一关键帧的帧间隔数;lcmin为连贯性得分下限阈值比例。
  19. 根据权利要求18所述的装置,其特征在于,还包括:
    综合计算模块,用于根据所述准确性匹配分值以及连贯性得分,进行加权计算得到当前动作的综合得分。
  20. 根据权利要求19所述的装置,其特征在于,
    所述综合计算模块,还用于当所述源视频图像播放完毕,将每次计算得到的准确性匹配分值、连贯性得分和综合得分分别加和求均值,得出整个过程的平均得分。
PCT/CN2015/090279 2014-12-29 2015-09-22 图像处理方法及装置 Ceased WO2016107226A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201410836597.3 2014-12-29
CN201410836597.3A CN105809653B (zh) 2014-12-29 2014-12-29 图像处理方法及装置

Publications (1)

Publication Number Publication Date
WO2016107226A1 true WO2016107226A1 (zh) 2016-07-07

Family

ID=56284142

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2015/090279 Ceased WO2016107226A1 (zh) 2014-12-29 2015-09-22 图像处理方法及装置

Country Status (2)

Country Link
CN (1) CN105809653B (zh)
WO (1) WO2016107226A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112950951A (zh) * 2021-01-29 2021-06-11 浙江大华技术股份有限公司 智能信息显示方法、电子装置和存储介质
CN113705536A (zh) * 2021-09-18 2021-11-26 深圳市领存技术有限公司 连续动作打分方法、装置及存储介质

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108040289A (zh) * 2017-12-12 2018-05-15 天脉聚源(北京)传媒科技有限公司 一种视频播放的方法及装置
CN114827730B (zh) * 2022-04-19 2024-05-31 咪咕文化科技有限公司 视频封面选取方法、装置、设备及存储介质

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN1731316A (zh) * 2005-08-19 2006-02-08 北京航空航天大学 虚拟猿戏的人机交互方法
CN101615302A (zh) * 2009-07-30 2009-12-30 浙江大学 音乐数据驱动的基于机器学习的舞蹈动作生成方法
WO2010004953A1 (ja) * 2008-07-08 2010-01-14 株式会社コナミデジタルエンタテインメント ゲーム装置、コンピュータプログラムおよび記録媒体
CN103327356A (zh) * 2013-06-28 2013-09-25 Tcl集团股份有限公司 一种视频匹配方法、装置

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN1731316A (zh) * 2005-08-19 2006-02-08 北京航空航天大学 虚拟猿戏的人机交互方法
WO2010004953A1 (ja) * 2008-07-08 2010-01-14 株式会社コナミデジタルエンタテインメント ゲーム装置、コンピュータプログラムおよび記録媒体
CN101615302A (zh) * 2009-07-30 2009-12-30 浙江大学 音乐数据驱动的基于机器学习的舞蹈动作生成方法
CN103327356A (zh) * 2013-06-28 2013-09-25 Tcl集团股份有限公司 一种视频匹配方法、装置

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112950951A (zh) * 2021-01-29 2021-06-11 浙江大华技术股份有限公司 智能信息显示方法、电子装置和存储介质
CN112950951B (zh) * 2021-01-29 2023-05-02 浙江大华技术股份有限公司 智能信息显示方法、电子装置和存储介质
CN113705536A (zh) * 2021-09-18 2021-11-26 深圳市领存技术有限公司 连续动作打分方法、装置及存储介质
CN113705536B (zh) * 2021-09-18 2024-05-24 深圳市领存技术有限公司 连续动作打分方法、装置及存储介质

Also Published As

Publication number Publication date
CN105809653A (zh) 2016-07-27
CN105809653B (zh) 2019-01-01

Similar Documents

Publication Publication Date Title
CN111437583B (zh) 一种基于Kinect的羽毛球基本动作辅助训练系统
JP6124308B2 (ja) 動作評価装置及びそのプログラム
CN104598867B (zh) 一种人体动作自动评估方法及舞蹈评分系统
CN110448870B (zh) 一种人体姿态训练方法
US10186041B2 (en) Apparatus and method for analyzing golf motion
US20100208038A1 (en) Method and system for gesture recognition
CN108305283A (zh) 基于深度相机和基本姿势的人体行为识别方法及装置
CN106022213A (zh) 一种基于三维骨骼信息的人体动作识别方法
CN111263953A (zh) 动作状态评估系统、动作状态评估装置、动作状态评估服务器、动作状态评估方法以及动作状态评估程序
Vallabhaneni et al. The analysis of the impact of yoga on healthcare and conventional strategies for human pose recognition
CN107154058B (zh) 一种引导使用者还原魔方的方法
CN110561399A (zh) 用于运动障碍病症分析的辅助拍摄设备、控制方法和装置
CN115761901B (zh) 一种骑马姿势检测评估方法
CN113947811B (zh) 一种基于生成对抗网络的太极拳动作校正方法及系统
WO2017161734A1 (zh) 通过电视和体感配件矫正人体动作及系统
CN108875586A (zh) 一种基于深度图像与骨骼数据多特征融合的功能性肢体康复训练检测方法
CN117911264A (zh) 一种基于图像融合和注意力机制的手部穴位检测方法
CN112309540A (zh) 运动评估方法、装置、系统及存储介质
CN105809653A (zh) 图像处理方法及装置
Tarek et al. Yoga trainer for beginners via machine learning
CN1731316A (zh) 虚拟猿戏的人机交互方法
CN114863237B (zh) 一种用于游泳姿态识别的方法和系统
Yuan Application of posture estimation optimization algorithm in the analysis of college air volleyball teaching movements
KR20020011851A (ko) 인공시각과 패턴인식을 이용한 체감형 게임 장치 및 방법.
JP2021026292A (ja) スポーツ行動認識装置、方法およびプログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15874911

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205 DATED 08/11/2017)

122 Ep: pct application non-entry in european phase

Ref document number: 15874911

Country of ref document: EP

Kind code of ref document: A1