WO2024251044A1 - 视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质 - Google Patents

视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质 Download PDF

Info

Publication number
WO2024251044A1
WO2024251044A1 PCT/CN2024/096691 CN2024096691W WO2024251044A1 WO 2024251044 A1 WO2024251044 A1 WO 2024251044A1 CN 2024096691 W CN2024096691 W CN 2024096691W WO 2024251044 A1 WO2024251044 A1 WO 2024251044A1
Authority
WO
WIPO (PCT)
Prior art keywords
sight
line
parameter
prediction model
eye
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2024/096691
Other languages
English (en)
French (fr)
Inventor
王亮亮
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Publication of WO2024251044A1 publication Critical patent/WO2024251044A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/774Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/80Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/80Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
    • G06V10/806Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/18Eye characteristics, e.g. of the iris
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/18Eye characteristics, e.g. of the iris
    • G06V40/193Preprocessing; Feature extraction
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02TCLIMATE CHANGE MITIGATION TECHNOLOGIES RELATED TO TRANSPORTATION
    • Y02T10/00Road transport of goods or passengers
    • Y02T10/10Internal combustion engine [ICE] based vehicles
    • Y02T10/40Engine management systems

Definitions

  • the present disclosure relates to the field of computer technology, and in particular to a method for constructing a sight line prediction model, a sight line prediction method, a device, an electronic device, and a storage medium.
  • Line of sight estimation plays an important role in many research fields, such as human-computer interaction, virtual reality, social interaction analysis, and medical treatment.
  • a method for constructing a sight line prediction model comprising:
  • the sight line prediction model Inputting the first eye sample image and the second eye sample image into the sight line prediction model to be constructed to obtain predicted sight line information, where the predicted sight line information includes a first predicted sight line parameter of the first eye sample image and a second predicted sight line parameter of the second eye sample image;
  • the model parameters of the line of sight prediction model to be constructed are updated based on the line of sight difference description information and the predicted line of sight information; and/or, when the loss of the line of sight prediction model to be constructed meets the convergence conditions, the line of sight prediction model to be constructed is determined to be a line of sight prediction model.
  • a sight line prediction method comprising:
  • the target eye image is input into a sight line prediction model to obtain sight line parameters corresponding to the target eye image.
  • the sight line prediction model is obtained by the sight line prediction model construction method described in the exemplary embodiment of the present disclosure.
  • a device for constructing a sight line prediction model comprising:
  • the training module is configured to input the first eye sample image and the second eye sample image into the auxiliary prediction model to obtain sight difference description information, where the sight difference description information is used to characterize the sight difference between the first eye sample image and the second eye sample image; input the first eye sample image and the second eye sample image into the sight prediction model to be constructed to obtain predicted sight information, where the predicted sight information includes a first predicted sight parameter of the first eye sample image and a first predicted sight parameter of the second eye sample image.
  • Predicting sight line parameters determining the loss of the sight line prediction model to be constructed based on the sight line difference description information and the predicted sight line information; when the loss of the sight line prediction model to be constructed does not meet the convergence condition, updating the model parameters of the sight line prediction model to be constructed based on the sight line difference description information and the predicted sight line information;
  • the determination module is configured to determine that the sight line prediction model to be constructed is a sight line prediction model when the loss of the sight line prediction model to be constructed meets a convergence condition.
  • a sight line prediction device comprising:
  • An acquisition module is configured to acquire a target eye image
  • the prediction module is configured to input the target eye image into a sight line prediction model to obtain sight line parameters corresponding to the target eye image.
  • the sight line prediction model is obtained by the sight line prediction model construction device described in the exemplary embodiment of the present disclosure.
  • an electronic device including:
  • Memory for storing programs and/or instructions
  • the processor When the program and/or instruction is executed by a processor, the processor is caused to perform the method according to the exemplary embodiment of the present disclosure.
  • a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions, when executed by a computer, cause the computer to perform the method according to the exemplary embodiments of the present disclosure.
  • a computer program product including instructions and/or a computer program, which, when executed by a processor, enables the method according to the exemplary embodiments of the present disclosure to be performed.
  • a computer program comprising a program code, which, when executed by a processor, enables execution of the method according to the exemplary embodiments of the present disclosure.
  • FIG1 shows a flow chart of a method for constructing a sight line prediction model according to an exemplary embodiment of the present disclosure
  • FIG2 shows a schematic diagram of a model architecture of an exemplary embodiment of the present disclosure
  • FIG3 shows a flow chart of a sight line prediction method according to an exemplary embodiment of the present disclosure
  • FIG4 shows a schematic block diagram of modules of a device for constructing a sight line prediction model according to an exemplary embodiment of the present disclosure
  • FIG5 shows a schematic block diagram of modules of a sight line prediction device according to an exemplary embodiment of the present disclosure
  • FIG6 shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure
  • FIG. 7 shows a block diagram of an exemplary electronic device that can be used to implement an embodiment of the present disclosure.
  • the exemplary embodiments of the present disclosure provide a method for constructing a line of sight prediction model, a line of sight prediction method, an apparatus, an electronic device, and a storage medium.
  • the line of sight difference description information output by the auxiliary prediction model is used in the construction of the line of sight prediction model, for example, it is used as the supervision information of the eye sample image pair, so as to achieve favorable technical effects, such as but not limited to at least one of the following: on the one hand, the sample annotation cost can be reduced, and on the other hand, the line of sight deviation caused by individual differences such as the angle between the eye visual axis and the optical axis can be eliminated, and the overfitting caused by "discrete sample data" during the acquisition of the eye sample image pair can be avoided.
  • the line of sight prediction model obtained according to the embodiments of the present disclosure can be directly used to predict the line of sight direction corresponding to the eye image to be tested in the inference stage, without first acquiring a calibration image for calibration.
  • the method of the exemplary embodiments of the present disclosure may be executed by an electronic device, or may be executed by a chip applied to the electronic device.
  • the electronic device of the exemplary embodiment of the present disclosure may be an electronic device with a display function, and the electronic device may be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, and other terminals.
  • UMPC ultra-mobile personal computer
  • FIG1 shows a flow chart of a method for constructing a sight line prediction model according to an exemplary embodiment of the present disclosure.
  • the method for constructing a sight line prediction model according to an exemplary embodiment of the present disclosure may include:
  • Step 101 Input the first eye sample image and the second eye sample image into the auxiliary prediction model to obtain sight difference description information, where the sight difference description information is used to characterize the sight difference between the first eye sample image and the second eye sample image.
  • the first eye sample image and the second eye sample image may be two eye images belonging to the same object.
  • a single eye image it may be a single eye image or a binocular image, where the binocular image may be a spliced image formed by left and right eye images from the same object collected at the same time, or may be a binocular image of the same object collected at the same time.
  • they may be collected at the same time or at different times, which is determined according to the actual scene and is not limited here.
  • the auxiliary prediction model can process the first eye sample image and the second eye sample image in a shared weight manner, that is, the auxiliary prediction model can use the same weight for processing the first eye sample image and the second eye sample image.
  • an exemplary embodiment of the present disclosure can use an auxiliary prediction model to output sight difference description information between the first eye sample image and the second eye sample image, and the sight difference description information can be used to characterize the sight difference between the first sight direction and the second sight direction.
  • the above-mentioned auxiliary prediction model can be an untrained neural network model, an incompletely trained neural network model, or a pre-trained neural network model.
  • the exemplary embodiment of the present disclosure can use two accurately labeled eye images and the corresponding two labeled sight lines as training samples, and determine the loss based on the predicted value of the sight difference of the two eye images output by the neural network model during the training process and the true value of the sight difference determined by the two labeled sight lines.
  • the currently trained neural network model is determined to be an auxiliary prediction model.
  • the auxiliary prediction model can be used to predict the sight difference description information corresponding to the first eye sample image and the second eye sample image.
  • Step 102 Input the first eye sample image and the second eye sample image into the gaze prediction model to be constructed to obtain predicted gaze information, where the predicted gaze information includes a first predicted gaze parameter of the first eye sample image and a second predicted gaze parameter of the second eye sample image.
  • the processing weights of the first eye sample image and the second eye sample image by the prediction model to be constructed are consistent, so that the prediction model to be constructed can process the first eye sample image and the second eye sample image in a shared weight manner, thereby obtaining the first predicted line of sight parameter and the second predicted line of sight parameter.
  • the exemplary embodiments of the present disclosure may input the first eye sample image and the second eye sample image into the line of sight prediction model to be constructed at the same time, or may input them into the line of sight prediction model to be constructed separately.
  • the line of sight prediction model to be constructed may output a first predicted line of sight parameter based on the first eye sample image, and output a second predicted line of sight parameter based on the second eye sample image.
  • the first predicted line of sight parameter may be used to characterize the predicted result of the eye line of sight direction in the first eye sample image
  • the second predicted line of sight parameter may be used to characterize the predicted result of the eye line of sight direction in the second eye sample image. fruit.
  • Step 103 Determine the loss of the sight line prediction model to be constructed based on the sight line difference description information and the predicted sight line information.
  • eye images with marked sight directions are often used to train sight estimation models, and then the eye images to be tested are directly input into the trained sight estimation model to determine the sight direction of the eye images to be tested; and when training the sight estimation model, the loss is usually determined based on the predicted value of the sight estimation model and the marked sight direction of the eye images.
  • the sight estimation model obtained by training with limited eye images with precise annotations in its inference stage, when the object individual to which the eye image to be tested obtained by the sight estimation model belongs is different from the object individual to which the sample image processed in the training stage belongs, the sight direction determined by the sight estimation model often has a sight estimation deviation caused by individual differences.
  • the training sample set of eye images with precise annotations usually has the problem of "data discreteness", if the loss is determined only based on the predicted value of the sight estimation model and the marked sight direction of the eye image, it is easy to cause the sight estimation model to overfit during the training process.
  • the method of the exemplary embodiment of the present disclosure uses the sight difference description information output by the auxiliary prediction model as the supervision information of the eye sample image to achieve favorable technical effects, such as but not limited to at least one of the following: on the one hand, the sample annotation cost is reduced, and on the other hand, the sight line prediction model can learn the sight line difference information between the first eye sample image and the second eye sample image during the training process. Due to the data composition characteristics of the sight line difference information itself, it can eliminate the prediction bias introduced by the aforementioned individual differences, so that the sight line prediction model constructed thereby can avoid the problem of inaccurate sight line direction prediction caused by individual differences when predicting the sight line. In addition, by supervising the training process of the sight line prediction model through the sight line difference description information output by the auxiliary prediction model, the model overfitting problem caused by the "data discreteness" of the eye sample image training data set can also be avoided.
  • the loss of the sight prediction model to be constructed in the exemplary embodiment of the present disclosure may include at least a first loss parameter, and the first loss parameter may be determined by the first predicted sight parameter, the second predicted sight parameter, and the supervision information of the eye sample image.
  • the sight difference description parameter of the exemplary embodiment of the present disclosure may include at least a reference sight difference parameter
  • the supervision information of the above-mentioned eye sample image may include the reference sight difference parameter
  • the above-mentioned first loss parameter may include a first sub-loss parameter
  • the first sub-loss parameter may be determined based on the reference sight difference parameter and the sight difference parameter
  • the sight difference parameter may be determined by the difference between the first predicted sight parameter and the second predicted sight parameter.
  • the reference sight line difference parameter may actually be the difference between the first sight line direction determined by the first eye sample image and the second sight line direction determined by the second eye sample image.
  • it may also be a related parameter used to represent the sight line direction difference.
  • the sight line direction may be defined by the sight line angle.
  • the exemplary embodiment of the present disclosure can define a first sight direction determined by a first eye sample image by a first sight angle, define a second sight direction determined by a second eye sample image by a second sight angle, and then determine the first sight angle and the second sight angle.
  • a sight angle difference is determined, and the sight angle difference is defined as a related parameter of the sight direction difference.
  • the exemplary embodiment of the present disclosure may determine the line of sight difference parameter of the first eye sample image and the second eye sample image by solving the difference between the first predicted line of sight parameter and the second predicted line of sight parameter.
  • the essence of the above-mentioned line of sight difference parameter may be the difference between the first predicted line of sight parameter of the first eye sample image and the second predicted line of sight parameter of the second eye sample image.
  • it may also be a related parameter used to represent the line of sight difference.
  • the first predicted line of sight parameter may be defined as the first predicted line of sight angle
  • the second predicted line of sight parameter may be defined as the second predicted line of sight angle by the line of sight angle.
  • the line of sight difference parameter may be the difference between the first predicted line of sight angle and the second predicted line of sight angle.
  • the method of the exemplary embodiment of the present invention can determine the first sub-loss parameter based on the reference line of sight difference parameter and the line of sight difference parameter, and use the first sub-loss parameter to determine the line of sight direction prediction loss of the line of sight prediction model to be constructed, so as to continuously reduce the line of sight direction prediction loss of the line of sight prediction model to be constructed during the training iteration process, thereby eliminating the prediction bias caused by individual differences and avoiding the model overfitting problem caused by the "data discreteness" of the eye sample image training data set.
  • the construction of a visual prediction model can be performed based on the determined loss of the line of sight prediction model to be constructed, such as the line of sight direction prediction loss of the line of sight prediction model to be constructed.
  • the loss of the line of sight prediction model to be constructed can be judged, such as judging whether the loss meets a specific condition, and based on the judgment result, processing related to the construction of the visual prediction model can be performed.
  • the model construction can be continued or terminated based on the judgment result.
  • the above judgment process can be performed in an appropriate manner, such as steps 104 to 106 described below.
  • Step 104 Determine whether the loss of the sight line prediction model to be constructed meets the convergence condition. If the convergence condition is not met, it means that the sight line prediction accuracy of the currently trained sight line prediction model to be constructed is relatively poor and cannot be accurately used to predict the sight line information of the eye image. Further training is required to meet the sight line prediction accuracy. At this time, step 105 can be executed; if the convergence condition is met, it means that the sight line prediction accuracy of the currently trained sight line prediction model to be constructed is relatively high and can be accurately used to predict the sight line information of the eye image. At this time, step 106 can be executed.
  • the above-mentioned convergence conditions may include various conditions.
  • the loss of the line of sight prediction model to be constructed may be less than or equal to a preset threshold.
  • it may also refer to relevant information/conditions on the stability of the loss of the line of sight prediction model to be constructed, such as the loss change/fluctuation being less than a specific range, etc.
  • the specific conditions are determined according to the actual application scenario and are not limited here.
  • Step 105 Update the model parameters of the line of sight prediction model to be constructed based on the line of sight difference description information and the predicted line of sight information.
  • the model parameters here may include weights and/or offset values.
  • the line of sight prediction model to be constructed after the parameters are updated can be reused to construct the line of sight prediction model, for example, it can be reapplied to the operations of the aforementioned steps 101-104, so as to construct the next round of line of sight prediction model.
  • Step 106 Determine that the sight line prediction model to be constructed is a sight line prediction model. At this point, the sight line prediction model has been completed. The construction process.
  • the construction of the sight prediction model can be performed iteratively.
  • the initial sight prediction model to be constructed and the initial auxiliary prediction model can be applied, and then the sample image is used as input to perform the aforementioned construction process, such as steps 101-104. If it is determined in step 104 that the loss of the sight prediction model to be constructed does not meet the convergence condition, the parameters of the sight prediction model to be constructed are updated, and then the sight prediction model to be constructed with the updated parameters is used to start the construction operation in the next iteration.
  • step 104 can be performed first, and then other steps can be performed. For example, first determine whether the loss of the sight prediction model to be constructed meets the convergence condition (it should be noted that in the first operation, various model parameters, loss functions, etc. can all be default values or initial settings). If the convergence condition is not met, then perform operations such as steps 101-103 and 105.
  • the convergence condition can also be related to the number of iterations. For example, if the number of iterations is greater than a specific number, it can be considered that the iteration condition is met.
  • the convergence condition may include at least one convergence condition, such as the convergence condition described herein, and the iteration may stop if any convergence condition is satisfied.
  • the specific conditions are determined according to the actual application scenario and are not limited here.
  • the disparity estimation model uses eye image pairs as input and outputs sight difference information during the training process and the reasoning process. Therefore, in the actual reasoning process, it is necessary to first guide the user to look at the known direction (hereinafter referred to as the calibration direction) marker, and collect the corresponding eye image as the subsequent calibration image.
  • the calibration direction the known direction
  • the electronic device needs to estimate the user's sight direction in real time, it is necessary to collect the user's eye image to be tested in real time, and estimate the sight difference between the eye image to be tested and the aforementioned calibration image through the disparity estimation model. Since the sight direction of the calibration image is known, the sight direction represented by the eye image to be tested can be determined based on the sight difference estimated by the disparity estimation model and the calibration direction.
  • the collected calibration image may not match the calibration direction due to some reasons, such as the user not looking at the marker in the calibration direction, or not looking at the marker at all. In this case, there is a large deviation between the sight difference estimated by the disparity estimation model and the sight direction represented by the eye image to be tested determined by the calibration direction.
  • the sight difference description information output by the auxiliary prediction model is used as the supervision information of the eye sample image, which can not only reduce the sample annotation cost, but also eliminate the sight deviation caused by individual differences such as the angle between the eye visual axis and the optical axis, and avoid overfitting caused by "discrete sample data" during the eye sample image acquisition process.
  • the sight prediction model constructed by the method of the exemplary embodiment of the present disclosure can directly output the corresponding sight direction based on the input image to be tested, without the need to collect calibration images in the inference phase, thereby improving the user experience.
  • the sight difference description information of the exemplary embodiment of the present disclosure may further include the first eye image feature
  • the predicted sight line information may further include the second eye image feature.
  • the loss of the test model may include not only the first loss parameter but also the second loss parameter.
  • the first loss parameter can be determined by the first predicted sight line parameter, the second predicted sight line parameter and the supervision information of the eye sample image, which is actually the prediction loss.
  • the specific content can be found in the previous text and will not be repeated here.
  • the second loss parameter is determined by the first eye image feature and the second eye image feature, which is actually the feature loss.
  • the exemplary embodiment of the present disclosure may define eye sample images with marked sight directions as labeled eye sample images, and define eye sample images without marked sight directions as unlabeled eye sample images. From the perspective of whether the eye sample images have marked sight directions, the eye sample images of the exemplary embodiment of the present disclosure may include unlabeled eye sample images, may include labeled eye sample images, or may include both unlabeled eye sample images and labeled eye sample images. At this time, when determining the first loss parameter, the supervision information of the eye sample images needs to be determined according to the actual scenario.
  • the supervision information of the eye sample image can include a reference sight difference parameter
  • the first loss parameter can include a first sub-loss parameter, which can be determined by the reference sight difference parameter and the sight difference parameter.
  • the relevant content of the first sub-loss parameter can be specifically referred to in the previous text, and will not be repeated here.
  • the exemplary embodiment of the present disclosure does not need to accurately label the unlabeled eye sample image, nor does it need to abandon the unlabeled eye sample image.
  • the reference line of sight difference parameter output by the auxiliary prediction model can be used as the supervision information of the eye sample image, and then the prediction loss of the line of sight prediction model to be constructed is determined (i.e., the first sub-loss parameter is determined). Therefore, in the training stage of the line of sight prediction model, the unlabeled eye sample images that have not been accurately labeled can be fully learned, thereby reducing the sample labeling cost and improving the utilization efficiency of the unlabeled eye sample images.
  • the supervision information of the eye sample image may also include the annotated sight lines of all eye sample images.
  • the annotated sight lines of all eye sample images include the first annotated sight line direction of the first eye sample image and the second annotated sight line direction of the second eye sample image.
  • the first annotated sight line direction can be used to annotate the eye sight line direction in the first eye sample image
  • the second annotated sight line direction can be used to annotate the eye sight line direction in the second eye sample image.
  • the first loss parameter of the exemplary embodiment of the present disclosure may include a second sub-loss parameter.
  • the second sub-loss parameter may be determined by the first predicted loss parameter and the second predicted loss parameter, wherein the first predicted loss parameter may be determined by the first annotated sight line direction and the first predicted sight line parameter, and the second predicted loss parameter is determined by the second annotated sight line direction and the second predicted sight line parameter.
  • an exemplary embodiment of the present disclosure may determine a first predicted loss parameter based on a first annotated sight line direction and a first predicted sight line parameter, determine a second predicted loss parameter based on a second predicted sight line parameter and a second annotated sight line direction, and then determine a second sub-loss parameter based on a weighted result of the first predicted loss parameter and the second predicted loss parameter.
  • the exemplary embodiment of the present disclosure can directly convert The annotated gaze direction of the eye sample image is used as the supervision information of the eye sample image.
  • the prediction loss of the gaze prediction model to be constructed is determined using the accurately annotated gaze direction of the eye sample image (i.e., determining the second sub-loss parameter).
  • the supervision information of the labeled eye sample image can be fully utilized, thereby improving the utilization rate of the supervision information of the labeled eye sample image.
  • the loss of the sight line prediction model to be constructed in the exemplary embodiment of the present disclosure can be determined by the weighted result of the first loss parameter and the second loss parameter.
  • feature scale both the eye image features and the predicted sight line information are represented in the form of matrices, and the scale difference between the two is relatively large.
  • the weight of the second loss parameter can be set to be greater than the weight of the first loss parameter, and the second loss parameter can be amplified by the weight of the second loss parameter, so that the loss of the sight line prediction model to be constructed can fully consider the loss of eye image features, thereby improving the sight line prediction accuracy of the constructed sight line prediction model.
  • the exemplary embodiment of the present disclosure can comprehensively determine the loss of the line of sight prediction model to be constructed from two aspects: prediction loss and feature loss, so that the model parameters of the obtained line of sight prediction model are optimal.
  • the line of sight parameters predicted by the line of sight prediction model in the inference stage are closer to the real line of sight parameters of the eye image, thereby improving the line of sight prediction accuracy of the line of sight prediction model.
  • the exemplary embodiments of the present disclosure may add some branches to the line of sight prediction model to be constructed, through which the image features extracted from the first eye sample image and the second eye sample image are respectively fused, and the predicted line of sight difference parameters of the first eye sample image and the second eye sample image are output based on the fused features.
  • the predicted line of sight information of the exemplary embodiments of the present disclosure may also include predicted line of sight difference parameters.
  • the predicted line of sight difference parameters are different from the line of sight difference parameters of the first eye sample image and the second eye sample image determined based on the first predicted line of sight parameters and the second predicted line of sight parameters in the previous text, and they can be directly output through the branches added in the line of sight prediction model to be constructed.
  • the loss of the sight line prediction model to be constructed may also include a third loss parameter, the essence of which is the prediction loss, and the third loss parameter may be determined by a reference sight line difference parameter and a prediction sight line difference parameter.
  • the loss of the line of sight prediction model to be constructed may include a first loss parameter, a second loss parameter and a third loss parameter
  • the loss of the line of sight prediction model to be constructed may be determined by a weighted result of the first loss parameter, the second loss parameter and the third loss parameter, and the weight of the first loss parameter and the weight of the third loss parameter are both smaller than the weight of the second loss parameter.
  • the third loss parameter and the first loss parameter belong to the prediction loss, and the scale of the annotated sight line direction of the corresponding eye sample image is also large. Therefore, when determining the loss of the sight line prediction model to be constructed, the first loss is set.
  • the weight of the loss parameter and the weight of the third loss parameter are both less than the weight of the second loss parameter.
  • the weight of the second loss parameter can be used to amplify the second loss parameter, so that the loss of the sight line prediction model to be constructed can fully consider the loss of eye image features, thereby improving the sight line prediction accuracy of the constructed sight line prediction model.
  • the values of X1 and X2 can be set to 1
  • the value of X3 can be set to 100.
  • the exemplary embodiment of the present disclosure can determine whether the loss of the line of sight prediction model to be constructed meets the convergence condition during the training stage. If so, the currently trained line of sight prediction model to be constructed is determined to be the line of sight prediction model; otherwise, the model parameters of the line of sight prediction model to be constructed are updated based on the line of sight difference description information and the predicted line of sight information.
  • the loss of the sight prediction model to be constructed is determined based on the weighted result of the first sub-loss parameter, the second loss parameter, and the third loss parameter, and then it is determined whether the loss of the sight prediction model to be constructed meets the convergence condition of the sight prediction model to be constructed.
  • the model parameters of the sight prediction model to be constructed are updated based on the reference sight difference parameter and the sight difference parameter, the first eye image feature and the second eye image feature, and the reference sight difference parameter and the predicted sight difference parameter.
  • the convergence condition of the sight prediction model to be constructed may include that the loss of the sight prediction model to be constructed is less than or equal to the first preset threshold, or it may refer to that the loss of the sight prediction model to be constructed is stable, which is determined according to the actual application scenario and is not limited here.
  • the loss of the line of sight prediction model to be constructed is determined based on the weighted result of the second sub-loss parameter, the second loss parameter and the third loss parameter, and then it is determined whether the loss of the line of sight prediction model to be constructed meets the convergence condition of the line of sight prediction model to be constructed.
  • the model parameters of the line of sight prediction model to be constructed are updated based on the first annotated line of sight direction and the first predicted line of sight parameter, the second annotated line of sight direction and the second predicted line of sight parameter, the first eye image feature and the second eye image feature, and the reference line of sight difference parameter and the predicted line of sight difference parameter.
  • the convergence condition of the line of sight prediction model to be constructed may include that the loss of the line of sight prediction model to be constructed is less than or equal to the second preset threshold, or it may refer to that the loss of the line of sight prediction model to be constructed is stable, which is determined according to the actual application scenario and is not limited here.
  • the annotated sight line difference parameter determined based on the first annotated sight line direction and the second annotated sight line direction is also the “real label” of the eye sample image.
  • the parameters and the labeled sight difference parameters determine the loss of the currently trained auxiliary prediction model. If the loss of the currently trained auxiliary prediction model does not meet the convergence condition of the auxiliary prediction model, it means that the convergence degree of the currently trained auxiliary prediction model is not high. At this time, the model parameters of the auxiliary prediction model can be updated based on the reference sight difference parameters and the labeled sight difference parameters; otherwise, there is no need to update the model parameters of the auxiliary prediction model.
  • the method of the exemplary embodiment of the present disclosure may update the model parameters of the sight line prediction model to be constructed separately, or may simultaneously update the model parameters of the sight line prediction model to be constructed and the model parameters of the auxiliary prediction model. It can be understood that the update of the model parameters of the sight line prediction model to be constructed exists in each iterative training, while the update of the model parameters of the auxiliary prediction model may exist in the iterative training in which the first eye sample image has a first annotated sight line direction and the second eye sample image has a second annotated sight line direction.
  • the model parameters of the auxiliary prediction model will be updated based on the reference sight line difference parameters and the annotated sight line difference parameters.
  • the exemplary embodiment of the present disclosure can determine the loss of the auxiliary prediction model based on the reference sight line difference parameter and the annotated sight line difference parameter determined by the first annotated sight line direction and the second annotated sight line direction, and then determine whether the loss of the auxiliary prediction model currently being trained meets the convergence condition of the auxiliary prediction model.
  • the loss of the auxiliary prediction model currently being trained does not meet the convergence condition of the auxiliary prediction model, it means that the convergence degree of the auxiliary prediction model currently being trained is not high, and there is a certain gap between the reference sight line difference parameter output by the auxiliary prediction model currently being trained and the annotated sight line difference parameter.
  • the model parameters of the auxiliary prediction model can be updated based on the reference sight line difference parameter and the annotated sight line difference parameter, so that the reference sight line difference parameter output by the auxiliary prediction model in the next iterative training process is closer to the annotated sight line difference parameter, and the sight line difference prediction accuracy of the auxiliary prediction model is higher, thereby making the prediction sight line difference parameter output by the added branch in the sight line prediction model to be constructed in the next iterative training process more accurate.
  • the convergence condition of the auxiliary prediction model may include that the loss of the auxiliary prediction model is less than or equal to a second preset threshold, or may refer to that the loss of the auxiliary prediction model is stable, which is determined according to the actual application scenario and is not limited here.
  • the method of the exemplary embodiment of the present disclosure may further include: determining a marked sight difference parameter based on the first marked sight direction and the second marked sight direction; and calibrating the reference sight difference parameter based on the marked sight difference parameter.
  • the determination of the marked sight difference parameter can refer to the above text and will not be repeated here.
  • the loss of the auxiliary prediction model does not meet the loss convergence condition, it means that the convergence degree of the currently trained auxiliary prediction model is not high, and there is a certain gap between the reference sight difference parameter output by the currently trained auxiliary prediction model and the labeled sight difference parameter.
  • the exemplary embodiment of the present disclosure can update the model parameters of the auxiliary prediction model based on the reference sight difference parameter and the labeled sight difference parameter, and can also calibrate the reference sight difference parameter based on the labeled sight difference parameter, so that the calibrated The reference sight difference parameter is closer to the annotated sight difference parameter. Therefore, the third loss parameter determined based on the calibrated reference sight difference parameter and the predicted sight difference parameter output by the additional branch in the sight prediction model to be constructed is smaller.
  • the method of the exemplary embodiment of the present disclosure can also update the model parameters of the sight prediction model to be constructed based on the calibrated reference sight difference parameter and the predicted sight difference parameter output by the additional branch in the sight prediction model to be constructed during the current iteration, and can also improve the sight prediction accuracy of the sight prediction model to be constructed in the next iteration.
  • the line of sight prediction model to be constructed since the updating of the model parameters of the line of sight prediction model to be constructed occurs in each training iteration process, in the iterative training process of updating the model parameters of the auxiliary prediction model, when the model parameters of the line of sight prediction model to be constructed are updated, the line of sight prediction model to be constructed will iteratively update the model parameters based on the reference line of sight difference parameters output by the auxiliary line of sight model and the predicted line of sight difference parameters output by the branch added in the line of sight prediction model to be constructed. Therefore, the line of sight prediction model to be constructed can learn the reference line of sight difference parameters output by the auxiliary prediction model.
  • the line of sight difference prediction accuracy of the auxiliary prediction model will gradually improve, and the improvement process of the line of sight difference prediction accuracy of the auxiliary line of sight model will also be learned by the line of sight prediction model to be constructed, so that the line of sight prediction model to be constructed can learn the generalization ability of the auxiliary line of sight model.
  • FIG2 shows a schematic diagram of a model architecture of an exemplary embodiment of the present disclosure.
  • the model architecture of the exemplary embodiment of the present disclosure may include an auxiliary sight line model 210 and a sight line prediction model 220
  • the auxiliary prediction model 210 may include a first backbone network 211, a first feature splicing network 212, and a first differential prediction network 213
  • the sight line prediction model 220 to be constructed may include a second backbone network 221 and a direction prediction network 222.
  • the first backbone network 211 contains a first basic module, which can be used to output a first eye image feature; the first feature stitching network 212 can be used to determine a first line of sight stitching feature based on the first eye image feature; the first difference prediction network 213 can be used to determine a reference line of sight difference parameter based on the first line of sight stitching feature.
  • the first backbone network 211 can obtain the first eye image feature based on the first eye sample image and the second eye sample image, and then the first feature stitching network 212 can obtain the first sight feature of the first eye sample image and the second sight feature of the second eye sample image based on the first eye image feature, and stitch the first sight feature of the first eye sample image and the second sight feature of the second eye sample image into the first sight stitching feature.
  • the first differential prediction network 213 fuses the first sight stitching feature to obtain the reference sight difference parameter. Based on this, since the first sight feature of the first eye sample image and the second sight feature of the second eye sample image have been spliced, individual differences such as the angle between the visual axis and the optical axis are eliminated. Therefore, when the reference sight difference parameter is used as the supervision information of the eye sample image, the sight deviation caused by individual differences can be eliminated, and the overfitting problem caused by "data discreteness" can be avoided.
  • the second backbone network 221 includes a second basic module, which can be used to output a second eye image feature; the direction prediction network 222 can be used to determine a first predicted sight line parameter and a second predicted sight line parameter based on the second eye image feature.
  • the second backbone network 221 may obtain the second eye image feature based on the first eye sample image and the second eye sample image, and the direction prediction network 222 may obtain the third eye feature of the first eye sample image based on the first eye sample image.
  • Line of sight features obtaining the fourth line of sight features of the second eye sample image based on the second eye sample image.
  • the direction prediction network 222 can also obtain the first predicted line of sight parameters of the first eye sample image based on the third line of sight features of the first eye sample image, and obtain the second predicted line of sight parameters of the second eye sample image based on the fourth line of sight features of the second eye sample image.
  • the eye feature difference parameters of each layer can be determined by the first eye image features and the second eye image features with the same number of channels.
  • the first backbone network contains multiple layers of first basic modules and the second backbone network contains multiple layers of second basic modules
  • the first basic modules and the second basic modules have the same number of layers, and each layer of the first basic modules corresponds to the second basic modules.
  • the first backbone network 211 may include 4 layers of first basic modules connected in series, each layer of first basic modules is used to extract first eye image features, the first layer of first basic modules can extract first layer first eye image features, the first layer of first eye image features are used as input of the second layer of first basic modules, the second layer of first basic modules extracts second layer first eye image features based on the first layer first eye image features, and so on, until the fourth layer of first basic modules extracts fourth layer first eye image features, and the fourth layer first eye image features are determined as the first eye image features.
  • the second backbone network 221 may include four layers of second basic modules connected in series, each layer of second basic modules is used to extract second eye image features, the first layer of second basic modules can extract first layer of second eye image features, the first layer of second eye image features are used as input to the second layer of second basic modules, the second layer of second basic modules extract second layer of second eye image features based on the first layer of second eye image features, and so on, until the fourth layer of second basic modules extracts fourth layer of second eye image features, and the fourth layer of second eye image features are determined as second eye image features.
  • the second loss parameter of the exemplary embodiment of the present disclosure is determined by multiple layers of eye feature difference parameters, and the eye feature difference parameters of each layer are determined by the first eye image features output by the first basic module of the corresponding layer and the second eye image features output by the second basic module of the corresponding layer.
  • the exemplary embodiments of the present disclosure may first determine the eye feature difference parameter of each layer based on the first eye image feature output by the first basic module of the corresponding layer and the second eye image feature output by the second basic module of the corresponding layer; and then determine the second loss function based on the eye feature difference parameters of the multiple layers.
  • the exemplary embodiments of the present disclosure may determine the second loss function based on the weighted results of the eye feature difference parameters of the multiple layers, and the weight of the eye feature difference parameters of each layer may be The same or different ones can be determined according to the actual scenario and are not specifically limited here.
  • the method of the exemplary embodiment of the present disclosure can determine the second loss parameter of the second eye image feature relative to the first eye image feature layer by layer, so that the line of sight prediction model can deeply learn the ability of the auxiliary prediction model to extract eye image features during the training stage, thereby improving the prediction accuracy of the line of sight prediction model in the reasoning stage.
  • the sight line prediction model 220 to be constructed can also include a second feature stitching network 223 and a second differential prediction network 224.
  • the second feature stitching network 223 can be used to determine the second sight line stitching feature based on the second eye image feature
  • the second differential prediction network 224 can be used to determine the predicted sight line difference parameters based on the second sight line stitching feature.
  • the second feature stitching network 223 and the second differential prediction network 224 can be understood as branches added to the sight line prediction model 220 to be constructed mentioned above.
  • the second backbone network 221 can obtain the second eye image feature based on the first eye sample image and the first eye sample image, and then the second feature stitching network 223 can obtain the fifth sight feature of the first eye sample image and the sixth sight feature of the second eye sample image based on the second eye image feature, and stitch the fifth sight feature of the first eye sample image and the sixth sight feature of the second eye sample image into the second sight stitching feature.
  • the second difference prediction network 224 fuses the second sight stitching feature to obtain the predicted sight difference parameter.
  • the line of sight prediction model 220 to be constructed may include a second backbone network 221 and a direction prediction network 222, as well as a second feature splicing network 223 and a second differential prediction network 224.
  • the second feature splicing network 223 and the second differential prediction network 224 may be used as additional branches in the line of sight prediction model 220 to be constructed mentioned above.
  • the second feature splicing network 223 and the second differential prediction network 224 output predicted line of sight difference parameters in the training stage, they may learn the process in which the first feature splicing network 212 and the first differential prediction network 213 in the auxiliary prediction model 210 output reference line of sight difference parameters, thereby improving the line of sight difference prediction accuracy of the auxiliary line of sight model through multiple iterative trainings.
  • the line of sight prediction model to be constructed may also be learned, thereby enabling the line of sight prediction model to be constructed to learn the generalization ability of the auxiliary line of sight model.
  • the model architecture of the line of sight prediction model in the reasoning stage may include a second backbone network 221 and a direction prediction network 222. At this time, compared with the model architecture of the line of sight prediction model 220 to be constructed in the training stage, it is smaller and simpler, and more conducive to deployment on edge devices such as wearable devices.
  • the exemplary embodiments of the present disclosure also provide a line of sight prediction method, which can improve the line of sight prediction accuracy in the inference stage based on the line of sight prediction model trained by the exemplary embodiments of the present disclosure. It should be understood that the method of the exemplary embodiments of the present disclosure can be executed by an electronic device or by a chip applied to an electronic device. For details, please refer to the above text and will not be repeated here.
  • Step 301 Obtain a target eye image.
  • the specific content of the target eye image is described in the above description of the eye image, which will not be described again here.
  • Step 302 Input the target eye image into a sight line prediction model to obtain sight line parameters corresponding to the target eye image, wherein the sight line prediction model is trained by the method of the exemplary embodiment of the present disclosure.
  • the sight line prediction model can be used to predict the sight line parameters corresponding to the target eye image. If the target eye image is a labeled eye sample image, the sight line parameters corresponding to the target eye image predicted by the sight line prediction model can be used to calibrate the supervision information of the labeled eye sample image.
  • the method of the exemplary embodiment of the present disclosure can use the sight prediction model trained by the exemplary embodiment of the present disclosure to predict the sight parameters corresponding to the target eye image.
  • the sight prediction model uses the sight difference description information output by the auxiliary prediction model as the supervision information of the eye sample image in the training stage to achieve favorable technical effects, such as but not limited to at least one of the following: on the one hand, it can reduce the sample annotation cost, on the other hand, it can eliminate the sight deviation caused by individual differences such as the angle between the eye visual axis and the optical axis, and on the other hand, it can avoid overfitting caused by "discrete sample data" during the eye sample image acquisition process.
  • the obtained sight prediction model can be directly used to predict the sight direction corresponding to the eye image to be tested in the reasoning stage without first collecting a calibration image for calibration. Additionally or alternatively, when the sight prediction model is trained, the model architecture of the sight prediction model in the reasoning stage is smaller and simpler than the model architecture of the sight prediction model to be constructed in the training stage, and is more conducive to deployment on edge devices such as wearable devices.
  • One or more technical solutions provided in the exemplary embodiments of the present disclosure can input the first eye sample image and the second eye sample image into the auxiliary prediction model to obtain the sight difference description information, and input the first eye sample image and the second eye sample image into the sight prediction model to be constructed to obtain the predicted sight information, and the predicted sight information can include the first predicted sight parameter of the first eye sample image and the second predicted sight parameter of the second eye sample image.
  • the sight difference description information can be used to characterize the sight difference between the first eye sample image and the second eye sample image, when the exemplary embodiments of the present disclosure use the sight difference description information output by the auxiliary prediction model as the supervision information of the eye sample image, it can achieve favorable technical effects, such as but not limited to at least one of the following: on the one hand, it can reduce the sample annotation cost, on the other hand, it can also eliminate the sight deviation caused by individual differences such as the angle between the eye visual axis and the optical axis, and on the other hand, it can also avoid overfitting caused by "discrete sample data" during the eye sample image acquisition process.
  • the sight prediction model finally obtained can be directly used to predict the sight direction corresponding to the eye image to be tested in the inference stage without first collecting a calibration image for calibration.
  • the apparatus provided by the embodiment of the present disclosure can execute the various methods provided by any embodiment of the present disclosure.
  • the embodiment of the present disclosure can divide the apparatus into functional units according to the above method examples.
  • each functional module/unit can be divided corresponding to each function, or two or more functions can be integrated into one processing module.
  • the various modules/units included in the above-mentioned apparatus are only divided according to functional logic, but are not limited to the division described in the text. There may be other division methods in actual implementation as long as the corresponding functions can be achieved; in addition, the specific names of the modules/units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
  • each module/unit can be implemented in various appropriate ways, such as The present invention may be implemented in hardware, firmware, or any suitable combination.
  • the exemplary embodiment of the present disclosure provides a device for constructing a sight line prediction model, which can be an electronic device or a chip applied to an electronic device.
  • FIG4 shows a module schematic block diagram of the device for constructing a sight line prediction model of the exemplary embodiment of the present disclosure. As shown in FIG4, the device 400 includes:
  • the training module 401 is configured to input the first eye sample image and the second eye sample image into the auxiliary prediction model to obtain line of sight difference description information, where the line of sight difference description information is used to characterize the line of sight difference between the first eye sample image and the second eye sample image; input the first eye sample image and the second eye sample image into the line of sight prediction model to be constructed to obtain predicted line of sight information, where the predicted line of sight information includes a first predicted line of sight parameter of the first eye sample image and a second predicted line of sight parameter of the second eye sample image; determine the loss of the line of sight prediction model to be constructed based on the line of sight difference description information and the predicted line of sight information; in particular, when the loss of the line of sight prediction model to be constructed does not meet the convergence condition, the model parameters of the line of sight prediction model to be constructed can be updated based on the line of sight difference description information and the predicted line of sight information;
  • the determination module 402 is configured to determine that the sight line prediction model to be constructed is a sight line prediction model when the loss of the sight line prediction model to be constructed meets a convergence condition.
  • the sight difference description information includes a first eye image feature and a reference sight difference parameter
  • the predicted sight information also includes a second eye image feature
  • the loss of the sight prediction model to be constructed includes a first loss parameter and a second loss parameter
  • the first loss parameter is determined by the first predicted line of sight parameter, the second predicted line of sight parameter and the supervision information of the eye sample image
  • the supervision information of the eye sample image includes a reference line of sight difference parameter
  • the second loss parameter is determined by the first eye image feature and the second eye image feature.
  • the loss of the sight line prediction model to be constructed is determined by a weighted result of a first loss parameter and a second loss parameter, and the weight of the second loss parameter is greater than the weight of the first loss parameter.
  • the first loss parameter includes a first sub-loss parameter and a second sub-loss parameter
  • the first sub-loss parameter is determined by a reference sight difference parameter and a sight difference parameter, and the sight difference parameter is determined by a difference between a first predicted sight parameter and a second predicted sight parameter;
  • the supervision information of the eye sample image also includes a first annotated sight direction of the first eye sample image and a second annotated sight direction of the second eye sample image.
  • the second sub-loss parameter is determined by the first predicted loss parameter and the second predicted loss parameter.
  • the first predicted loss parameter is determined by the first annotated sight direction and the first predicted sight parameter.
  • the second predicted loss parameter is determined by the second annotated sight direction and the second predicted sight parameter.
  • the predicted line of sight information also includes a predicted line of sight difference parameter
  • the loss of the line of sight prediction model to be constructed also includes a third loss parameter
  • the third loss parameter is determined by the predicted line of sight difference parameter and the reference line of sight difference parameter.
  • the training module 401 is also configured to update the model parameters of the auxiliary prediction model based on the reference line of sight difference parameter and the labeled line of sight difference parameter, and the labeled line of sight difference parameter is determined by the first labeled line of sight direction and the second labeled line of sight direction.
  • the loss of the sight prediction model to be constructed is determined by the weighted result of the first loss parameter, the second loss parameter and the third loss parameter, and the weight of the first loss parameter and the weight of the third loss parameter are both smaller than the weight of the second loss parameter.
  • the determination module 402 is further configured to determine a labeled sight line difference parameter based on the first labeled sight line direction and the second labeled sight line direction; it should be noted that, alternatively, the determination of the labeled sight line difference parameter may be implemented by the training module 401 .
  • the apparatus 400 further includes a calibration module 403 configured to calibrate the reference sight difference parameter based on the marked sight difference parameter.
  • the calibration module can be implemented separately from other modules, or included in other modules, such as in the training module, or its function can be implemented by the training module.
  • the auxiliary prediction model includes a first backbone network, a first feature concatenation network, and a first differential prediction network
  • the sight line prediction model to be constructed includes a second backbone network and a direction prediction network
  • the first backbone network includes a first basic module configured to output a first eye image feature, a first feature stitching network used to determine a first sight stitching feature based on the first eye image feature, and a first difference prediction network used to determine a reference sight difference parameter based on the first sight stitching feature;
  • the second backbone network contains a second basic module configured to output a second eye image feature, and the direction prediction network is used to determine a first predicted sight line parameter and a second predicted sight line parameter based on the second eye image feature.
  • the first backbone network contains multiple layers of first basic modules
  • the second backbone network contains multiple layers of second basic modules.
  • Each layer of the first basic modules corresponds to the second basic modules.
  • the second loss parameter is determined by the multiple layers of eye feature difference parameters
  • each layer of the eye feature difference parameters is determined by the first eye image features output by the first basic module of the corresponding layer and the second eye image features output by the second basic module of the corresponding layer.
  • the line of sight prediction model to be constructed also includes a second feature stitching network and a second differential prediction network
  • the second feature stitching network is used to determine the second line of sight stitching feature based on the second eye image feature
  • the second differential prediction network is used to determine the predicted line of sight difference parameters based on the second line of sight stitching feature.
  • the exemplary embodiment of the present disclosure further provides a sight line prediction device, which can be an electronic device or a chip applied to an electronic device.
  • FIG5 shows a schematic block diagram of a sight line prediction device of an exemplary embodiment of the present disclosure.
  • the device 500 includes:
  • An acquisition module 501 is configured to acquire a target eye image
  • the prediction module 502 is configured to input the target eye image into a sight line prediction model to obtain sight line parameters corresponding to the target eye image, and the sight line prediction model is trained by the device described in the exemplary embodiment of the present disclosure.
  • the electronic device includes hardware structures and/or software modules corresponding to executing each function.
  • the present disclosure can be implemented in hardware. Or a combination of hardware and computer software. Whether a function is performed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
  • FIG6 shows a schematic block diagram of a chip of an exemplary embodiment of the present disclosure.
  • the chip 600 includes one or more (including two) processors 601 and a communication interface 602.
  • the communication interface 602 can support the server to perform the data sending and receiving steps in the above method, and the processor 601 can support the server to perform the data processing steps in the above method.
  • the chip 600 further includes a memory 603, which may include a read-only memory and a random access memory, and provides operation instructions and data to the processor.
  • a portion of the memory may also include a non-volatile random access memory (NVRAM).
  • NVRAM non-volatile random access memory
  • the processor 601 performs corresponding operations by calling operation instructions stored in the memory (the operation instructions may be stored in the operating system).
  • the processor 601 controls the processing operations of any one of the terminal devices, and the processor may also be referred to as a central processing unit (CPU).
  • the memory 603 may include a read-only memory and a random access memory, and provides instructions and data to the processor 601.
  • a portion of the memory 603 may also include NVRAM.
  • the memory, the communication interface, and the memory are coupled together through a bus system, wherein the bus system may include a power bus, a control bus, and a status signal bus in addition to a data bus.
  • various buses are labeled as bus system 604 in FIG6 .
  • the method disclosed in the above-mentioned embodiment of the present disclosure can be applied to a processor or implemented by a processor.
  • the processor may be an integrated circuit chip with signal processing capabilities.
  • each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software.
  • the above-mentioned processor can be a general-purpose processor, a digital signal processor (digital signal processing, DSP), an ASIC, a field-programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
  • DSP digital signal processing
  • ASIC application-programmable gate array
  • FPGA field-programmable gate array
  • the methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed.
  • the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
  • the steps of the method disclosed in conjunction with the embodiment of the present disclosure can be directly embodied as a hardware decoding processor to be executed, or a combination of hardware and software modules in the decoding processor can be executed.
  • the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc.
  • the storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
  • the exemplary embodiment of the present disclosure also provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication.
  • the memory stores a computer program and/or instructions that can be executed by the at least one processor, and the computer program and/or instructions, when executed by the at least one processor, cause the electronic device to perform a method according to an embodiment of the present disclosure.
  • the exemplary embodiments of the present disclosure also provide a non-transitory computer readable storage medium storing a computer program and/or instructions.
  • a storage medium wherein the computer program, when executed by a processor of a computer, causes the computer to perform a method according to an embodiment of the present disclosure.
  • the exemplary embodiments of the present disclosure further provide a computer program product, including a computer program and/or instructions, wherein when the computer program is executed by a processor of a computer, the computer is enabled to perform the method according to the embodiments of the present disclosure.
  • the exemplary embodiments of the present disclosure also provide a computer program, including a program code, which, when executed by a processor, causes the processor to execute the method according to the embodiment of the present disclosure.
  • FIG. 7 a block diagram of an electronic device 700 that can be used as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure.
  • the electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.
  • the electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices.
  • the components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the present disclosure described and/or required herein.
  • the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703.
  • ROM read-only memory
  • RAM random access memory
  • various programs and data required for the operation of the device 700 can also be stored.
  • the computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704.
  • An input/output (I/O) interface 705 is also connected to the bus 704.
  • a plurality of components in the electronic device 700 are connected to the I/O interface 705, including: an input unit 706, an output unit 707, a storage unit 708, and a communication unit 709.
  • the input unit 707 may be any type of device capable of inputting information to the electronic device 700, and the input unit 706 may receive input digital or character information, and generate key signal inputs related to user settings and/or function control of the electronic device.
  • the output unit 707 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video/audio output terminal, a vibrator, and/or a printer.
  • the storage unit 708 may include, but is not limited to, a disk, an optical disk.
  • the communication unit 709 allows the electronic device 700 to exchange information/data with other devices via a computer network such as the Internet and/or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and/or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and/or the like.
  • the computing unit 701 may be a variety of general and/or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc.
  • the computing unit 701 performs the various methods and processes described above.
  • the method of the exemplary embodiments of the present disclosure may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708.
  • part or all of the computer program may be loaded and/or installed on the electronic device 700 via the ROM 702 and/or the communication unit 709.
  • the computing unit 701 may be configured to execute the method in any other appropriate manner (eg, by means of firmware).
  • the program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions/operations specified in the flow chart and/or block diagram to be implemented.
  • the program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
  • a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment.
  • a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
  • a machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing.
  • a more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
  • RAM random access memory
  • ROM read-only memory
  • EPROM or flash memory erasable programmable read-only memory
  • CD-ROM portable compact disk read-only memory
  • CD-ROM compact disk read-only memory
  • magnetic storage device or any suitable combination of the foregoing.
  • machine-readable medium and “computer-readable medium” refer to any computer program product, apparatus, and/or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and/or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal.
  • machine-readable signal refers to any signal for providing machine instructions and/or data to a programmable processor.
  • the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer.
  • a display device e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor
  • a keyboard and pointing device e.g., a mouse or trackball
  • Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
  • the systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components.
  • the components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
  • a computer system may include clients and servers.
  • Clients and servers are generally remote from each other and usually interact through a communication network.
  • the relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
  • the computer program product includes one or more computer programs or instructions.
  • the computer may be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user device or other programmable device.
  • the computer program or instruction may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instruction may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired or wireless means.
  • the computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server, data center, etc. that integrates one or more available media.
  • the available medium may be a magnetic medium, for example, a floppy disk, a hard disk, a tape; it may also be an optical medium, for example, a digital video disc (DVD); it may also be a semiconductor medium, for example, a solid state drive (SSD).

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Health & Medical Sciences (AREA)
  • Multimedia (AREA)
  • General Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Evolutionary Computation (AREA)
  • Computing Systems (AREA)
  • Software Systems (AREA)
  • Medical Informatics (AREA)
  • Databases & Information Systems (AREA)
  • Artificial Intelligence (AREA)
  • Ophthalmology & Optometry (AREA)
  • Human Computer Interaction (AREA)
  • Eye Examination Apparatus (AREA)

Abstract

本公开提供一种视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质,所述方法包括:将第一眼部样本图像和第二眼部样本图像输入辅助预测模型,获得视线差描述信息;将第一眼部样本图像和第二眼部样本图像输入待构建视线预测模型,获得预测视线信息;若基于视线差描述信息和预测视线信息确定的待构建视线预测模型的损失未满足收敛条件,基于视线差描述信息和预测视线信息更新待构建视线预测模型的模型参数;否则,确定待构建视线预测模型为视线预测模型。

Description

视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质
相关申请的交叉引用
本申请是以申请号为202310668262.4、申请日为2023年6月7日的中国申请为基础,并主张其优先权,该中国申请的公开内容在此作为整体引入本申请中。
技术领域
本公开涉及计算机技术领域,尤其涉及一种视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质。
背景技术
随着计算机视觉、人工智能等领域的高速发展,视线估计技术的研究引起了广泛的关注。视线估计在很多研究领域都发挥着重要的作用,如:人机交互、虚拟现实、社会交互分析、医疗等。
发明内容
根据本公开的一方面,提供了一种视线预测模型的构建方法,包括:
将第一眼部样本图像和第二眼部样本图像输入辅助预测模型,获得视线差描述信息,视线差描述信息用于表征第一眼部样本图像和第二眼部样本图像的视线差;
将第一眼部样本图像和第二眼部样本图像输入待构建视线预测模型,获得预测视线信息,预测视线信息包括第一眼部样本图像的第一预测视线参数和第二眼部样本图像的第二预测视线参数;
基于视线差描述信息和预测视线信息确定待构建视线预测模型的损失;
在待构建视线预测模型的损失未满足收敛条件的情况下,基于视线差描述信息和预测视线信息更新待构建视线预测模型的模型参数;和/或,在待构建视线预测模型的损失满足收敛条件的情况下,确定待构建视线预测模型为视线预测模型。
根据本公开的另一方面,提供了一种视线预测方法,包括:
获取目标眼部图像;
将目标眼部图像输入至视线预测模型,获得目标眼部图像对应的视线参数,该视线预测模型由本公开示例性实施例所述的视线预测模型的构建方法得到。
根据本公开的另一方面,提供了一种视线预测模型的构建装置,包括:
训练模块,被配置为将第一眼部样本图像和第二眼部样本图像输入辅助预测模型,获得视线差描述信息,视线差描述信息用于表征第一眼部样本图像和第二眼部样本图像的视线差;将第一眼部样本图像和第二眼部样本图像输入待构建视线预测模型,获得预测视线信息,预测视线信息包括第一眼部样本图像的第一预测视线参数和第二眼部样本图像的第 二预测视线参数;基于视线差描述信息和预测视线信息确定待构建视线预测模型的损失;在待构建视线预测模型的损失未满足收敛条件的情况下,基于视线差描述信息和预测视线信息更新待构建视线预测模型的模型参数;
确定模块,被配置为在待构建视线预测模型的损失满足收敛条件的情况下,确定待构建视线预测模型为视线预测模型。
根据本公开的还另一方面,提供了一种视线预测装置,包括:
获取模块,被配置为获取目标眼部图像;
预测模块,被配置为将目标眼部图像输入至视线预测模型,获得目标眼部图像对应的视线参数,该视线预测模型由本公开示例性实施例所述的视线预测模型的构建装置得到。
根据本公开的还另一方面,提供了一种电子设备,包括:
处理器;以及,
存储程序和/或指令的存储器;
其中,该程序和/或指令在由处理器执行时使处理器执行根据本公开示例性实施例所述的方法。
根据本公开的另一方面,提供了一种存储有计算机指令的非瞬时计算机可读存储介质,其中,计算机指令在由计算机执行时使得计算机执行根据本公开示例性实施例所述的方法。
根据本公开的还另一方面,提供了一种计算机程序产品,包括指令和/或计算机程序,所述指令和/或计算机程序被处理器执行时使得执行根据本公开示例性实施例所述的方法。
根据本公开的还另一方面,提供了一种计算机程序,包括程序代码,该程序代码在由处理器执行时使得执行根据本公开示例性实施例所述的方法。
附图说明
在下面结合附图对于示例性实施例的描述中,本公开的更多细节、特征和优点被公开,在附图中:
图1示出了本公开示例性实施例的视线预测模型的构建方法的流程图;
图2示出了本公开示例性实施例的模型架构示意图;
图3示出了本公开示例性实施例的视线预测方法的流程图;
图4示出了本公开示例性实施例的视线预测模型的构建装置的模块示意性框图;
图5示出了本公开示例性实施例的视线预测装置的模块示意性框图;
图6示出了本公开示例性实施例的芯片的示意性框图;
图7示出了能够用于实现本公开的实施例的示例性电子设备的结构框图。
具体实施方式
下面将参照附图更详细地描述本公开的实施例。虽然附图中显示了本公开的某些实施 例,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例,相反提供这些实施例是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实施例仅用于示例性作用,并非用于限制本公开的保护范围。
应当理解,本公开的方法实施方式中记载的各个步骤可以按照不同的顺序执行,和/或并行执行。此外,方法实施方式可以包括附加的步骤和/或省略执行示出的步骤。本公开的范围在此方面不受限制。
本文使用的术语“包括”及其变形是开放性包括,即“包括但不限于”。术语“基于”是“至少部分地基于”。术语“一个实施例”表示“至少一个实施例”;术语“另一实施例”表示“至少一个另外的实施例”;术语“一些实施例”表示“至少一些实施例”。其他术语的相关定义将在下文描述中给出。需要注意,本公开中提及的“第一”、“第二”等概念仅用于对不同的装置、模块或单元进行区分,并非用于限定这些装置、模块或单元所执行的功能的顺序或者相互依存关系。
需要注意,本公开中提及的“一个”、“多个”的修饰是示意性而非限制性的,本领域技术人员应当理解,除非在上下文另有明确指出,否则应该理解为“一个或多个”。
本公开实施方式中的多个装置之间所交互的消息或者信息的名称仅用于说明性的目的,而并不是用于对这些消息或信息的范围进行限制。
本公开示例性实施例提供一种视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质,特别地,在视线预测模型构建中利用了辅助预测模型输出的视线差描述信息,例如将其作为眼部样本图像对的监督信息,从而实现有利的技术效果,例如包括但不限于以下中的至少一者,一方面可以降低样本标注成本,另一方面还可以消除由眼睛视轴和光轴夹角等个体差异性所带来的视线偏差,避免眼部样本图像对采集过程中因“样本数据离散”而带来的过拟合。附加地或者可替代地,根据本公开的实施例获得的视线预测模型在推理阶段可以直接用于预测待测眼部图像对应的视线方向,无需先采集校准图像进行校准。
本公开示例性实施例的方法可以由电子设备执行,也可以由应用于电子设备的芯片执行。
示例性的,本公开示例性实施例的电子设备可以是具有显示功能的电子设备执行,电子设备可以为手机、平板电脑、可穿戴设备、车载设备、笔记本电脑、超级移动个人计算机(ultra-mobile personal computer,UMPC)、上网本等终端。
图1示出了本公开示例性实施例的视线预测模型的构建方法的流程图。如图1所示,本公开示例性实施例的视线预测模型的构建方法可以包括:
步骤101:将第一眼部样本图像和第二眼部样本图像输入辅助预测模型,获得视线差描述信息,视线差描述信息用于表征第一眼部样本图像和第二眼部样本图像的视线差。上 述第一眼部样本图像和第二眼部样本图像可以是属于同一对象的两张眼部图像。针对单张眼部图像,其可以是单眼图像,也可以是双眼图像,此处双眼图像可以是通过同一时刻采集的来自同一对象的左右眼图像形成的拼接图像,也可以是同一时刻采集的同一对象的双眼图像。针对多张眼部图像,其可以在同一时刻采集,也可以在不同时刻采集,具体根据实际场景确定,此处不作限定。
当第一眼部样本图像和第二眼部样本图像输入辅助预测模型时,辅助预测模型在处理第一眼部样本图像和第二眼部样本图像的过程中,可以采用共享权重的方式进行处理,也就是辅助预测模型对于第一眼部样本图像的处理和第二眼部样本图像的处理权重可以一致。
假设第一眼部样本图像中的眼部视线方向定义为第一视线方向,第二眼部样本图像中的眼部视线方向定义为第二视线方向,本公开示例性实施例可以利用辅助预测模型输出第一眼部样本图像和第二眼部样本图像的视线差描述信息,该视线差描述信息可以用于表征第一视线方向和第二视线方向的视线差。
上述辅助预测模型可以是未经训练的神经网络模型,也可以是未完全训练的神经网络模型,还可以是预先训练完成的神经网络模型。示例性的,当上述辅助预测模型为预先训练完成的神经网络模型时,本公开示例性实施例可以采用精确标注的两张眼部图像和对应的两个标注视线方向为训练样本,基于训练过程中神经网络模型输出的两张眼部图像的视线差预测值和由两个标注视线方向确定的视线差真实值确定损失,当基于该损失确定当前训练的神经网络模型收敛时,确定当前训练的神经网络模型为辅助预测模型。当完成辅助预测模型的训练后,可以利用辅助预测模型预测第一眼部样本图像和第二眼部样本图像对应的视线差描述信息。
步骤102:将第一眼部样本图像和第二眼部样本图像输入待构建视线预测模型,获得预测视线信息,预测视线信息包括第一眼部样本图像的第一预测视线参数和第二眼部样本图像的第二预测视线参数。
当第一眼部样本图像和第二眼部样本图像输入待构建视线预测模型后,待构建预测模型对于第一眼部样本图像的处理和第二眼部样本图像的处理权重一致,使得待构建预测模型可以通过共享权重的方式处理第一眼部样本图像和第二眼部样本图像,从而获得第一预测视线参数和第二预测视线参数。
示例性的,本公开示例性实施例可以将第一眼部样本图像和第二眼部样本图像同时输入至待构建视线预测模型,也可以分别输入至待构建视线预测模型。待构建视线预测模型可以基于第一眼部样本图像输出第一预测视线参数,基于第二眼部样本图像输出第二预测视线参数。此处,第一预测视线参数可以用于表征第一眼部样本图像中的眼部视线方向的预测结果,第二预测视线参数可以用于表征第二眼部样本图像中的眼部视线方向的预测结 果。
步骤103:基于视线差描述信息和预测视线信息确定待构建视线预测模型的损失。
相关技术中,常常利用具有标注视线方向的眼部图像训练视线估计模型,然后将待测眼部图像直接输入训练完成的视线估计模型以确定该待测眼部图像的视线方向;并且,在训练该视线估计模型时通常基于该视线估计模型的预测值和眼部图像的标注视线方向确定损失。然而实践中,具有精确标注的眼部图像相对较少,并且由于不同个体的视轴和光轴夹角存在差异,通过有限的具有精确标注的眼部图像训练得到的视线估计模型,在其推理阶段,当视线估计模型获取的待测眼部图像所属的对象个体与其在训练阶段处理的样本图像所属的对象个体不同时,由视线估计模型确定的视线方向往往存在个体差异导致的视线估计偏差。此外,由于具有精确标注的眼部图像训练样本集通常存在“数据离散”的问题,如仅基于视线估计模型的预测值和眼部图像的标注视线方向确定损失,容易造成视线估计模型在训练过程中出现过拟合现象。
而本公开示例性实施例的方法以辅助预测模型输出的视线差描述信息作为眼部样本图像的监督信息,实现有利的技术效果,例如包括但不限于以下中的至少一者,一方面降低了样本标注成本,另一方面视线预测模型在训练过程中可以学习到第一眼部样本图像和第二眼部样本图像之间的视线差信息,由于视线差信息本身的数据构成特点,其可以消除因前述个体差异引入的预测偏差,从而使得由此构建得到的视线预测模型在预测视线时可以避免个体差异性所带来的视线方向预测不准确的问题。此外,通过辅助预测模型输出的视线差描述信息监督视线预测模型的训练过程,还可以避免由于眼部样本图像训练数据集“数据离散”带来的模型过拟合问题。
示例性的,本公开示例性实施例的待构建视线预测模型的损失至少可以包括第一损失参数,第一损失参数可以由第一预测视线参数、第二预测视线参数和眼部样本图像的监督信息确定。当本公开示例性实施例的视线差描述参数至少可以包括参考视线差参数时,上述眼部样本图像的监督信息可以包括参考视线差参数,上述第一损失参数可以包括第一子损失参数,第一子损失参数可以基于参考视线差参数和视线差异参数确定,视线差异参数可以由第一预测视线参数和第二预测视线参数的差值确定。
上述参考视线差参数实质可以是第一眼部样本图像确定的第一视线方向和第二眼部样本图像确定的第二视线方向的差值,当然,也可以是用以表示视线方向差值的相关参数,例如:可以通过视线角度定义视线方向。
举例来说,假设眼部正视前方时,眼部视线为基准视线,然后以该基准视线为基础,定义其他眼部视线方向与基准视线的夹角为其他眼部视线角度。基于此,本公开示例性实施例可以通过第一视线角度定义第一眼部样本图像确定的第一视线方向,通过第二视线角度定义第二眼部样本图像确定的第二视线方向,然后通过第一视线角度和第二视线角度确 定视线角度差值,将该视线角度差值定义为视线方向差值的相关参数。
当眼部样本图像的监督信息可以包括参考视线差参数时,本公开示例性实施例可以通过求解第一预测视线参数和第二预测视线参数的差值的方式确定第一眼部样本图像和第二眼部样本图像的视线差异参数。上述视线差异参数的实质可以是第一眼部样本图像的第一预测视线参数和第二眼部样本图像的第二预测视线参数的差值,当然,也可以是用以表示视线差异的相关参数。例如:可以通过视线角度定义第一预测视线参数为第一预测视线角度、定义第二预测视线参数为第二预测视线角度,此时,视线差异参数可以为第一预测视线角度和第二预测视线角度的差值。
基于此,本公开示例性实施例的方法可以基于参考视线差参数和视线差异参数确定第一子损失参数,利用该第一子损失参数确定待构建视线预测模型的视线方向预测损失,从而在训练迭代过程中,不断缩小待构建视线预测模型的视线方向预测损失,可以消除因个体差异带来的预测偏差,避免由于眼部样本图像训练数据集“数据离散”带来的模型过拟合问题。
根据本公开的实施例,可以基于所确定的待构建视线预测模型的损失,例如待构建视线预测模型的视线方向预测损失,来执行视觉预测模型的构建。在一些实施例中,可以对待构建视线预测模型的损失进行判断,例如判断损失是否满足特定条件,并且基于判断结果来执行与视觉预测模型构建相关的处理,作为示例,可以基于判断结果来继续进行模型构建或者终止模型构建。上述判断过程可以采用适当的方式来执行,例如如下所述的步骤104到106。
步骤104:判断待构建视线预测模型的损失是否满足收敛条件。如果不满足收敛条件,说明当前训练的待构建视线预测模型的视线预测准确性比较差,不能准确用于预测眼部图像的视线信息,需要进一步训练才能满足视线预测准确度,此时,可以执行步骤105;如果满足收敛条件,说明当前训练的待构建视线预测模型的视线预测准确性比较高,可以准确用于预测眼部图像的视线信息,此时,可以执行步骤106。
在实际应用中,上述收敛条件可以包括各种条件,作为一个示例,可以为待构建视线预测模型的损失小于或等于预设阈值,作为另一示例,也可以是指待构建视线预测模型的损失稳定的相关信息/条件,例如损失变化/波动小于特定范围等等,具体根据实际应用场景确定,此处不作限定。
步骤105:基于视线差描述信息和预测视线信息更新待构建视线预测模型的模型参数。此处的模型参数可以包括权值和/或偏移值。根据本公开的一些实施例,参数更新后的待构建视线预测模型可重新用于视线预测模型的构建,例如可以重新应用于前述步骤101-104的操作,从而进行下一轮的视线预测模型的构建。
步骤106:确定待构建视线预测模型为视线预测模型。此时,已经完成视线预测模型 的构建过程。
作为一个示例,视线预测模型的构建可以迭代地执行,特别地,在构建过程的开始,可以应用初始的待构建视线预测模型和初始的辅助预测模型,然后利用样本图像作为输入来执行前述的构建过程,例如步骤101-104,在步骤104中判断待构建视线预测模型的损失不满足收敛条件的情况下则更新待构建视线预测模型的参数,然后将参数更新后的待构建视线预测模型用于开始下一次迭代中的构建操作,另一方面,如果在步骤104中判断待构建视线预测模型的损失满足收敛条件,则迭代终止,将所得到的待构建视线预测模型为视线预测模型。应指出,上述步骤101到104的顺序仅是示例性的,还可以为其他适当的顺序。在一个示例中,可以首先执行步骤104,然后再执行其他步骤。例如,首先判断待构建视线预测模型的损失是否满足收敛条件,(应指出,在首次操作中,各种模型参数、损失函数等等可都为默认值或者初始设置值),如果不满足收敛条件,则执行操作,如步骤101-103以及105,如果满足收敛条件,则迭代终止,执行步骤106。作为示例,在迭代操作情况下,收敛条件也可以与迭代次数有关,例如迭代次数大于特定次数则可认为满足迭代条件。
应指出,收敛条件可以包含至少一个收敛条件,诸如如文中所述的收敛条件,并且任一收敛条件被满足,则迭代可以停止。具体根据实际应用场景确定,此处不作限定。
相关技术中,视差估计模型在训练过程、推理过程中,均以眼部图像对为输入,输出视线差信息,因此,在实际推理过程中,需要先引导用户注视已知方向(下文称作校准方向)标志物,并采集相应的眼部图像作为后续的校准图像。当电子设备需要实时估计用户视线方向时,需要实时采集用户的待测眼部图像,通过视差估计模型估计该待测眼部图像与前述校准图像之间的视线差,由于校准图像的视线方向已知,可以基于视差估计模型估计的视线差和校准方向,确定待测眼部图像表征的视线方向。但是,在获取校准视线时,可能由于一些原因,例如用户并未按照校准方向注视标志物,或者根本未注视标志物,导致所采集的校准图像与校准方向不匹配。这种情况下,基于视差估计模型估计的视线差和校准方向确定的待测眼部图像表征的视线方向存在较大的偏差。
而通过本公开示例性实施例的方法在视线预测模型的训练阶段,以辅助预测模型输出的视线差描述信息作为眼部样本图像的监督信息,不仅可以降低样本标注成本,还可以消除由眼睛视轴和光轴夹角等个体差异性所带来的视线偏差,避免眼部样本图像采集过程中因“样本数据离散”而带来的过拟合。在此基础上,本公开示例性实施例的方法构建得到的视线预测模型可以直接基于输入的待测图像输出相应的视线方向,无需在推理阶段采集校准图像,提升了用户体验。
在实际应用中,本公开示例性实施例的视线差描述信息还可以包括第一眼部图像特征,预测视线信息还可以包括第二眼部图像特征。此时,本公开示例性实施例的待构建视线预 测模型的损失不仅可以包括第一损失参数,还可以包括第二损失参数。
第一损失参数可以由第一预测视线参数、第二预测视线参数和眼部样本图像的监督信息确定,其实质为预测损失。具体内容可以参见前文,此处不再赘述。第二损失参数由第一眼部图像特征和第二眼部图像特征确定,其实质为特征损失。
本公开示例性实施例可以将具有标注视线方向的眼部样本图像定义为有标签眼部样本图像,将不具有标注视线方向的眼部样本图像定义为无标签眼部样本图像。从眼部样本图像是否具有标注视线方向来看,本公开示例性实施例的眼部样本图像可以包括无标签眼部样本图像,也可以包括有标签眼部样本图像,还可以同时包括无标签眼部样本图像和有标签眼部样本图像。此时,在确定第一损失参数时,眼部样本图像的监督信息需要根据实际场景进行确定。
可以理解的是,不管眼部样本图像是无标签眼部样本图像,还是有标签眼部样本图像,眼部样本图像的监督信息均可以包括参考视线差参数,第一损失参数可以包括第一子损失参数,第一子损失参数可以由参考视线差参数和视线差异参数确定。此处,第一子损失参数的相关内容具体可以参考前文,此处不再赘述。
因此,当眼部样本图像包括无标签眼部样本图像时,本公开示例性实施例无需对该无标签眼部样本图像进行精确标注,也无需舍弃使用该无标签眼部样本图像,而是可以将辅助预测模型输出的参考视线差参数作为眼部样本图像的监督信息,进而确定待构建视线预测模型的预测损失(即确定第一子损失参数),从而在视线预测模型的训练阶段可以充分学习那些没有精确标注过的无标签眼部样本图像,降低了样本标注成本,提高了无标签眼部样本图像的利用效率。
在此基础上,当眼部样本图像包括有标签眼部样本图像时,眼部样本图像的监督信息还可以包括所有眼部样本图像的标注视线方向。例如:在本公开示例性实施例的方法中,所有眼部样本图像的标注视线方向包括第一眼部样本图像的第一标注视线方向和第二眼部样本图像的第二标注视线方向。此处,第一标注视线方向可以用于标注第一眼部样本图像中的眼部视线方向,第二标注视线方向可以用于标注第二眼部样本图像中的眼部视线方向。基于此,本公开示例性实施例的第一损失参数可以包括第二子损失参数。第二子损失参数可以由第一预测损失参数和第二预测损失参数确定,其中,第一预测损失参数可以由第一标注视线方向和第一预测视线参数确定,所述第二预测损失参数由第二标注视线方向和第二预测视线参数确定。
例如,本公开示例性实施例可以基于第一标注视线方向和第一预测视线参数确定第一预测损失参数,基于第二预测视线参数和第二标注视线方向确定第二预测损失参数,然后,基于第一预测损失参数和第二预测损失参数的加权结果确定第二子损失参数。
可见,当眼部样本图像包括有标签眼部样本图像时,本公开示例性实施例可以直接将 眼部样本图像的标注视线方向作为眼部样本图像的监督信息,利用精确标注的眼部样本图像的标注视线方向确定待构建视线预测模型的预测损失(即确定第二子损失参数),此时,可以充分利用有标签眼部样本图像的监督信息,提高了有标签眼部样本图像的监督信息的利用率。
在此基础上,本公开示例性实施例的待构建视线预测模型的损失可以由第一损失参数和第二损失参数的加权结果确定。从特征尺度来说,眼部图像特征和预测视线信息均以矩阵的形式表示,二者的尺度差异较大。
举例来说,对于眼部图像特征来说,其尺度相对比较小,而预测视线信息的尺度较大,因此,第二损失参数比较小,而第一损失参数比较大。此时,可以设置第二损失参数的权重大于第一损失参数的权重,利用第二损失参数的权重放大第二损失参数,以使得待构建视线预测模型的损失可以充分考虑眼部图像特征损失,提高构建后的视线预测模型的视线预测准确性。
因此,本公开示例性实施例可以从预测损失和特征损失两个方面综合确定待构建视线预测模型的损失,使得获得的视线预测模型的模型参数最优,该视线预测模型在推理阶段预测的视线参数更加接近眼部图像的真实视线参数,从而可以提高视线预测模型的视线预测精度。
在实际应用中,本公开示例性实施例可以在待构建视线预测模型中增设一些分支,通过这些分支将分别第一眼部样本图像和第二眼部样本图像提取的图像特征进行特征融合,并基于融合特征输出第一眼部样本图像和第二眼部样本图像的预测视线差参数。基于此,本公开示例性实施例的预测视线信息还可以包括预测视线差参数。此处,预测视线差参数不同于前文中基于第一预测视线参数和第二预测视线参数确定的第一眼部样本图像和第二眼部样本图像的视线差异参数,其可以通过待构建视线预测模型中增设的分支直接输出。
这种情况下,待构建视线预测模型的损失还可以包括第三损失参数,该第三损失参数的实质为预测损失,第三损失参数可以由参考视线差参数和预测视线差参数确定。
当待构建视线预测模型的损失可以包括第一损失参数、第二损失参数和第三损失参数时,待构建视线预测模型的损失可以由第一损失参数、第二损失参数和第三损失参数的加权结果确定,第一损失参数的权重和第三损失参数的权重均小于第二损失参数的权重。
举例来说,待构建视线预测模型的损失的计算公式为:L=X1*L1+X2*L2+X3*L3,其中,L表示待构建视线预测模型的损失,L1、L2和L3分别表示第一损失参数、第二损失参数和第三损失参数,X1、X2和X3分别表示第一损失参数的权重、第二损失参数的权重和第三损失参数的权重。
此处,第三损失参数与第一损失参数一起,同属于预测损失,其对应的眼部样本图像的标注视线方向的尺度也较大,因此,在确定待构建视线预测模型的损失时,设置第一损 失参数的权重和第三损失参数的权重均小于第二损失参数的权重。此时,可以利用第二损失参数的权重放大第二损失参数,以使得待构建视线预测模型的损失可以充分考虑眼部图像特征损失,提高构建后的视线预测模型的视线预测准确性。例如,可以将X1、X2是取值设为1,将X3的取值设为100。
在此基础上,本公开示例性实施例可以在训练阶段判断待构建视线预测模型的损失是否满足收敛条件,若满足,确定当前训练的待构建视线预测模型为视线预测模型,否则,基于视线差描述信息和预测视线信息更新待构建视线预测模型的模型参数。
在实际应用中,在每次迭代训练过程中,当眼部样本图像的监督信息包括参考视线差参数、第一标注视线方向和第二标注视线方向时,基于第一子损失参数、第二损失参数和第三损失参数的加权结果确定待构建视线预测模型的损失,然后判断该待构建视线预测模型的损失是否满足待构建视线预测模型的收敛条件。若待构建视线预测模型的损失不满足待构建视线预测模型的收敛条件,基于参考视线差参数和视线差异参数、第一眼部图像特征和第二眼部图像特征、以及参考视线差参数和预测视线差参数更新待构建视线预测模型的模型参数。此处,待构建视线预测模型的收敛条件可以包括待构建视线预测模型的损失小于或等于第一预设阈值,也可以是指待构建视线预测模型的损失稳定,具体根据实际应用场景确定,此处不作限定。
当眼部样本图像的监督信息包括参考视线差参数时,基于第二子损失参数、第二损失参数和第三损失参数的加权结果确定待构建视线预测模型的损失,然后判断该待构建视线预测模型的损失是否满足待构建视线预测模型的收敛条件。若待构建视线预测模型的损失不满足待构建视线预测模型的收敛条件,基于第一标注视线方向和第一预测视线参数、第二标注视线方向和第二预测视线参数、第一眼部图像特征和第二眼部图像特征、以及参考视线差参数和预测视线差参数更新待构建视线预测模型的模型参数。此处,待构建视线预测模型的收敛条件可以包括待构建视线预测模型的损失小于或等于第二预设阈值,也可以是指待构建视线预测模型的损失稳定,具体根据实际应用场景确定,此处不作限定。
在一种可选的方式中,当眼部样本图像的监督信息包括参考视线差参数、第一标注视线方向和第二标注视线方向时,辅助预测模型可以是未经训练或者预训练的网络模型。此时,可以在训练待构建视线预测模型的同时,训练辅助预测模型。基于此,本公开示例性实施例的方法还可以包括:基于参考视线差参数和标注视线差参数,更新辅助预测模型的模型参数,标注视线差参数可以由第一标注视线方向和第二标注视线方向确定。此处的模型参数可以包括权值和/或偏移值。
由于第一标注视线方向和第二标注视线方向分别为第一眼部样本图像和第二眼部样本图像的“真实标签”,因此,基于第一标注视线方向和第二标注视线方向确定的标注视线差参数也是眼部样本图像的“真实标签”。此时,可以基于辅助预测模型输出的参考视线差 参数和标注视线差参数确定当前训练的辅助预测模型的损失,若当前训练的辅助预测模型的损失不满足辅助预测模型的收敛条件,说明当前训练的辅助预测模型收敛程度不高,此时可以基于参考视线差参数和标注视线差参数更新辅助预测模型的模型参数;否则,无需更新辅助预测模型的模型参数。
考虑到在每次迭代训练过程中所使用的眼部样本图像可能存在标注视线方向,也可能不存在标注视线方向,因此,本公开示例性实施例的方法可以单独更新待构建视线预测模型的模型参数,也可以同时更新待构建视线预测模型的模型参数和辅助预测模型的模型参数。可以理解的是,待构建视线预测模型的模型参数的更新存在于每次迭代训练中,而辅助预测模型的模型参数的更新可能存在于第一眼部样本图像具有第一标注视线方向和第二眼部样本图像具有第二标注视线方向的迭代训练中。也就是说,当第一眼部样本图像具有第一标注视线方向和第二眼部样本图像具有第二标注视线方向、且当前训练的辅助预测模型的损失不满足辅助预测模型的收敛条件时,才会基于参考视线差参数和标注视线差参数更新辅助预测模型的模型参数。
示例性的,在每次迭代训练过程中,若第一眼部样本图像具有第一标注视线方向、且第二眼部样本图像具有第二标注视线方向,本公开示例性实施例可以基于参考视线差参数和由第一标注视线方向和第二标注视线方向确定的标注视线差参数,确定辅助预测模型的损失,然后判断当前训练的辅助预测模型的损失是否满足辅助预测模型的收敛条件。若当前训练的辅助预测模型的损失不满足辅助预测模型的收敛条件,说明当前训练的辅助预测模型收敛程度不高,当前训练的辅助预测模型输出的参考视线差参数与标注视线差参数之间存在一定的差距,此时可以基于参考视线差参数和标注视线差参数更新辅助预测模型的模型参数,以使在下一次迭代训练过程中辅助预测模型输出的参考视线差参数更加靠近标注视线差参数,辅助预测模型的视线差预测准确度更高,进而使得待构建视线预测模型中的增设分支在下一次迭代训练过程中输出的预测视线差参数的精度更高。此处,辅助预测模型的收敛条件可以包括辅助预测模型的损失小于或等于第二预设阈值,也可以是指辅助预测模型的损失稳定,具体根据实际应用场景确定,此处不作限定。
基于此,若辅助预测模型的损失不满足辅助预测模型的收敛条件,本公开示例性实施例的方法还可以包括:基于第一标注视线方向和第二标注视线方向确定标注视线差参数;基于标注视线差参数对参考视线差参数进行校准。此处,标注视线差参数的确定可以参考前文,此处不再赘述。
若辅助预测模型的损失不满足损失收敛条件,说明当前训练的辅助预测模型收敛程度不高,当前训练的辅助预测模型输出的参考视线差参数与标注视线差参数之间存在一定的差距,本公开示例性实施例在基于参考视线差参数和标注视线差参数更新辅助预测模型的模型参数的同时,还可以基于标注视线差参数对参考视线差参数进行校准,以使校准后的 参考视线差参数更加靠近标注视线差参数。因此,基于校准后的参考视线差参数和由待构建视线预测模型中增设分支输出的预测视线差参数确定的第三损失参数更小,此时,由待构建视线预测模型中增设分支输出的预测视线差参数的预测精度更高。同时,本公开示例性实施例的方法还可以在当前迭代过程中,基于校准后的参考视线差参数和待构建视线预测模型中增设分支输出的预测视线差参数更新待构建视线预测模型的模型参数,还可以提高待构建视线预测模型在下一次迭代过程中的视线预测精度。
可见,在本公开示例性实施例的方法中,由于待构建视线预测模型的模型参数的更新发生在每次训练迭代过程中,在发生更新辅助预测模型的模型参数的迭代训练过程中,待构建视线预测模型的模型参数在更新时,待构建视线预测模型都会基于辅助视线模型输出的参考视线差参数和待构建视线预测模型中增设分支输出的预测视线差参数进行模型参数的迭代更新,因此,待构建视线预测模型可以学习到辅助预测模型输出的参考视线差参数,经过多次迭代训练,辅助预测模型的视线差预测精度会逐渐提高,辅助视线模型的视线差预测精度的提升过程也会被待构建视线预测模型学习到,进而使得待构建视线预测模型可以学习到辅助视线模型的泛化能力。
在一种可选的方式中,图2示出了本公开示例性实施例的模型架构示意图。如图2所示,本公开示例性实施例的模型架构可以包括辅助视线模型210和视线预测模型220,辅助预测模型210可以包括第一主干网络211、第一特征拼接网络212和第一差分预测网络213,待构建视线预测模型220可以包括第二主干网络221和方向预测网络222。
第一主干网络211含有第一基本模块,可以用于输出第一眼部图像特征;第一特征拼接网络212可以用于基于第一眼部图像特征确定第一视线拼接特征;第一差分预测网络213可以用于基于第一视线拼接特征确定参考视线差参数。
示例性的,第一主干网络211可以基于第一眼部样本图像和第二眼部样本图像获取第一眼部图像特征,然后第一特征拼接网络212可以基于第一眼部图像特征获得第一眼部样本图像的第一视线特征和第二眼部样本图像的第二视线特征,并将第一眼部样本图像的第一视线特征和第二眼部样本图像的第二视线特征拼接为第一视线拼接特征。第一差分预测网络213对第一视线拼接特征进行融合后获取参考视线差参数。基于此,由于已经对第一眼部样本图像的第一视线特征和第二眼部样本图像的第二视线特征进行拼接,消除了因为视轴和光轴夹角等个体差异性,因此,在将参考视线差参数作为眼部样本图像的监督信息时,可以消除个体差异所带来的视线偏差,避免“数据离散”带来的过拟合问题。
第二主干网络221含有第二基本模块,可以用于输出第二眼部图像特征;方向预测网络222可以用于基于第二眼部图像特征确定第一预测视线参数和第二预测视线参数。
示例性的,第二主干网络221可以基于第一眼部样本图像和第二眼部样本图像获取第二眼部图像特征,方向预测网络222基于第一眼部样本图像获取第一眼部样本图像的第三 视线特征,基于第二眼部样本图像获取第二眼部样本图像的第四视线特征,此时,方向预测网络222还可以基于第一眼部样本图像的第三视线特征获得第一眼部样本图像的第一预测视线参数,基于第二眼部样本图像的第四视线特征获得第二眼部样本图像的第二预测视线参数。
考虑到每层第一基本模块输出的第一眼部图像特征的通道数量和对应层第二基本模块输出的第二眼部图像特征的通道数量可能不同,此时,需要将同一层第一基本模块输出的第一眼部图像特征和第二基本模块输出的第二眼部图像特征的通道数量调整为相同的通道数量,因此,每层眼部特征差异参数可以由通道数量相同情况下的第一眼部图像特征和第二眼部图像特征确定。
举例来说,第一眼部图像特征的特征图尺寸为H1*W1*C1,第二眼部图像特征的特征图尺寸为H2*W2*C2,经过通道数量调整后,第一眼部图像特征的特征图尺寸为H1*W1,第二眼部图像特征的特征图尺寸为H2*W2,然后基于第一眼部图像特征和第二眼部图像特征确定第二损失参数。
当第一主干网络含有多层第一基本模块,第二主干网络含有多层第二基本模块时,第一基本模块和第二基本模块的层数相同,每层第一基本模块与第二基本模块对应。
例如,第一基本模块第二基本模块的层数均为4层时,第一主干网络211可以包括4层串接的第一基本模块,每层第一基本模块用于提取到第一眼部图像特征,第一层第一基本模块可以提取到第一层第一眼部图像特征,该第一层第一眼部图像特征作为第二层第一基本模块的输入,第二层第一基本模块基于第一层第一眼部图像特征提取第二层第一眼部图像特征,以此类推,直至第四层第一基本模块提取到第四层第一眼部图像特征,将第四层第一眼部图像特征确定为第一眼部图像特征。第二主干网络221可以包括4层串接的第二基本模块,每层第二基本模块用于提取到第二眼部图像特征,第一层第二基本模块可以提取到第一层第二眼部图像特征,该第一层第二眼部图像特征作为第二层第二基本模块的输入,第二层第二基本模块基于第一层第二眼部图像特征提取第二层第二眼部图像特征,以此类推,直至第四层第二基本模块提取到第四层第二眼部图像特征,将第四层第二眼部图像特征确定为第二眼部图像特征。
基于此,本公开示例性实施例的第二损失参数由多层眼部特征差异参数确定,每层眼部特征差异参数由对应层第一基本模块输出的第一眼部图像特征和对应层第二基本模块输出的第二眼部图像特征确定。
示例性的,本公开示例性实施例可以先基于对应层第一基本模块输出的第一眼部图像特征和对应层第二基本模块输出的第二眼部图像特征,确定每层的眼部特征差异参数;然后基于多层的眼部特征差异参数确定第二损失函数。例如,本公开示例性实施例可以基于多层的眼部特征差异参数的加权结果确定第二损失函数,每层眼部特征差异参数的权重可 以相同,也可以不同,根据实际场景确定,此处不做具体限定。
可见,本公开示例性实施例的方法可以逐层确定第二眼部图像特征相对于第一眼部图像特征的第二损失参数,使得视线预测模型在训练阶段可以深入学习辅助预测模型提取眼部图像特征的能力,进而提高视线预测模型在推理阶段的预测精度。
在一种可选的方式中,当预测视线信息还可以包括预测视线差参数时,待构建视线预测模型220还可以包括第二特征拼接网络223和第二差分预测网络224,第二特征拼接网络223可以用于基于第二眼部图像特征确定第二视线拼接特征,第二差分预测网络224可以用于基于第二视线拼接特征确定预测视线差参数。此处,第二特征拼接网络223和第二差分预测网络224可以理解为前文中提到的待构建视线预测模型220中增设的分支。
示例性的,第二主干网络221可以基于第一眼部样本图像和第一眼部样本图像获取第二眼部图像特征,然后第二特征拼接网络223可以基于第二眼部图像特征获得第一眼部样本图像的第五视线特征和第二眼部样本图像的第六视线特征,并将第一眼部样本图像的第五视线特征和第二眼部样本图像的第六视线特征拼接为第二视线拼接特征。第二差分预测网络224对第二视线拼接特征进行融合后获取预测视线差参数。
需要说明的是,在视线预测模型的训练阶段,待构建视线预测模型220可以包括第二主干网络221和方向预测网络222,以及第二特征拼接网络223和第二差分预测网络224,在每次迭代过程中,第二特征拼接网络223和第二差分预测网络224可以作为前文中提到的待构建视线预测模型220中的增设分支,第二特征拼接网络223和第二差分预测网络224在训练阶段输出预测视线差参数时,可以学习到辅助预测模型210中第一特征拼接网络212和第一差分预测网络213输出参考视线差参数的过程,从而通过多次迭代训练使辅助视线模型的视线差预测精度的提升过程也会被待构建视线预测模型学习到,进而使得待构建视线预测模型可以学习到辅助视线模型的泛化能力。而当视线预测模型训练完成后,视线预测模型在推理阶段的模型架构可以包括第二主干网络221和方向预测网络222,此时,其相对于训练阶段的待构建视线预测模型220的模型架构更小,更简单,更利于在可穿戴设备等边缘设备的部署。
本公开示例性实施例还提供一种视线预测方法,可以基于本公开示例性实施例训练的视线预测模型在推理阶段提高视线预测准确度。应理解,本公开示例性实施例的方法可以由电子设备执行,也可以由应用于电子设备的芯片执行。具体内容参见前文,此处不再赘述。
图3示出了本公开示例性实施例的视线预测方法的流程图。如图3所示,本公开示例性实施例的视线预测方法可以包括:
步骤301:获取目标眼部图像。目标眼部图像的具体内容参见前文眼部图像,此处不再赘述。
步骤302:将目标眼部图像输入至视线预测模型,获得目标眼部图像对应的视线参数,该视线预测模型由本公开示例性实施例的方法训练。
若目标眼部图像为无标签眼部样本图像,可以利用该视线预测模型预测目标眼部图像对应的视线参数。若目标眼部图像为有标签眼部样本图像,可以利用该视线预测模型预测的可以目标眼部图像对应的视线参数,校准该有标签眼部样本图像的监督信息。
可见,本公开示例性实施例的方法可以利用本公开示例性实施例训练的视线预测模型预测目标眼部图像对应的视线参数,该视线预测模型在训练阶段将辅助预测模型输出的视线差描述信息作为眼部样本图像的监督信息,实现有利的技术效果,例如包括但不限于以下中的至少一者,一方面可以降低样本标注成本,另一方面可以消除由眼睛视轴和光轴夹角等个体差异性所带来的视线偏差,还另一方面可以避免眼部样本图像采集过程中因“样本数据离散”而带来的过拟合。附加地或者可替代的,获得的视线预测模型在推理阶段可以直接用于预测待测眼部图像对应的视线方向,无需先采集校准图像进行校准。附加地或者可替代地,当视线预测模型训练完成后,视线预测模型在推理阶段的模型架构相对于训练阶段的待构建视线预测模型的模型架构更小,更简单,更利于在可穿戴设备等边缘设备的部署。
本公开示例性实施例中提供的一个或多个技术方案,可以将第一眼部样本图像和第二眼部样本图像输入辅助预测模型获得视线差描述信息,将第一眼部样本图像和第二眼部样本图像输入待构建视线预测模型获得预测视线信息,预测视线信息可以包括第一眼部样本图像的第一预测视线参数和第二眼部样本图像的第二预测视线参数。由于视线差描述信息可以用于表征第一眼部样本图像和第二眼部样本图像的视线差,因此,当本公开示例性实施例将辅助预测模型输出的视线差描述信息作为眼部样本图像的监督信息时,可以实现有利的技术效果,例如包括但不限于以下中的至少一者,一方面可以降低样本标注成本,另一方面也可以消除由眼睛视轴和光轴夹角等个体差异性所带来的视线偏差,还另一方面还可以避免眼部样本图像采集过程中因“样本数据离散”而带来的过拟合。附加地或者可替代地,最终获得的视线预测模型在推理阶段可以直接用于预测待测眼部图像对应的视线方向,无需先采集校准图像进行校准。
以下将描述根据本公开的实施例的装置。本公开实施例所提供的装置可执行本公开任意实施例所提供的各种方法,本公开实施例可以根据上述方法示例对装置进行功能单元的划分,例如,可以对应各个功能划分各个功能模块/单元,也可以将两个或两个以上的功能集成在一个处理模块中。值得注意的是,上述装置所包括的各个模块/单元只是按照功能逻辑进行划分的,但并不局限于文中所述的划分,实际实现时可以有另外的划分方式,只要能够实现相应的功能即可;另外,各模块/单元的具体名称也只是为了便于相互区分,并不用于限制本公开实施例的保护范围。此外,各模块/单元可采用各种适当方式来实现,例如 硬件、固件、或任何适当组合来实现的。
在采用对应各个功能划分各个功能模块的情况下,本公开示例性实施例提供一种视线预测模型的构建装置,该视线预测模型的构建装置可以为电子设备或应用于电子设备的芯片。图4示出了本公开示例性实施例的视线预测模型的构建装置的模块示意性框图。如图4所示,所述装置400包括:
训练模块401,被配置为将第一眼部样本图像和第二眼部样本图像输入辅助预测模型,获得视线差描述信息,视线差描述信息用于表征第一眼部样本图像和第二眼部样本图像的视线差;将第一眼部样本图像和第二眼部样本图像输入待构建视线预测模型,获得预测视线信息,预测视线信息包括第一眼部样本图像的第一预测视线参数和第二眼部样本图像的第二预测视线参数;基于视线差描述信息和预测视线信息确定待构建视线预测模型的损失;特别地,在所述待构建视线预测模型的损失未满足收敛条件的情况下,可以基于所述视线差描述信息和所述预测视线信息更新所述待构建视线预测模型的模型参数;
确定模块402,被配置用于在待构建视线预测模型的损失满足收敛条件的情况下,确定待构建视线预测模型为视线预测模型。
作为一种可能的实现方式,视线差描述信息包括第一眼部图像特征和参考视线差参数,预测视线信息还包括第二眼部图像特征,待构建视线预测模型的损失包括第一损失参数和第二损失参数;
其中,第一损失参数由第一预测视线参数、第二预测视线参数和眼部样本图像的监督信息确定,眼部样本图像的监督信息包括参考视线差参数,第二损失参数由第一眼部图像特征和第二眼部图像特征确定。
作为一种可能的实现方式,待构建视线预测模型的损失由第一损失参数和第二损失参数的加权结果确定,第二损失参数的权重大于第一损失参数的权重。
作为一种可能的实现方式,第一损失参数包括第一子损失参数和第二子损失参数;
第一子损失参数由参考视线差参数和视线差异参数确定,视线差异参数由第一预测视线参数和第二预测视线参数的差值确定;
眼部样本图像的监督信息还包括第一眼部样本图像的第一标注视线方向和第二眼部样本图像的第二标注视线方向,第二子损失参数由第一预测损失参数和第二预测损失参数确定,第一预测损失参数由第一标注视线方向和第一预测视线参数确定,第二预测损失参数由第二标注视线方向和第二预测视线参数确定。
作为一种可能的实现方式,预测视线信息还包括预测视线差参数,待构建视线预测模型的损失还包括第三损失参数,第三损失参数由预测视线差参数和参考视线差参数确定,训练模块401还被配置为基于参考视线差参数和标注视线差参数,更新辅助预测模型的模型参数,标注视线差参数由第一标注视线方向和第二标注视线方向确定。
作为一种可能的实现方式,待构建视线预测模型的损失由第一损失参数、第二损失参数和第三损失参数的加权结果确定,第一损失参数的权重和第三损失参数的权重均小于第二损失参数的权重。
作为一种可能的实现方式,确定模块402还被配置为基于第一标注视线方向和第二标注视线方向确定标注视线差参数;应指出,作为替代地,标注视线差参数的确定可由训练模块401实现。
装置400还包括,校准模块403,被配置为基于标注视线差参数对参考视线差参数进行校准。应指出,校准模块可以与其他模块分离地实现,也可被包含在其他模块中,例如可以包含在训练模块中,或者其功能可由训练模块实现。
作为一种可能的实现方式,辅助预测模型包括第一主干网络、第一特征拼接网络和第一差分预测网络,待构建视线预测模型包括第二主干网络和方向预测网络;
第一主干网络含有第一基本模块,被配置为输出第一眼部图像特征,第一特征拼接网络用于基于第一眼部图像特征确定第一视线拼接特征,第一差分预测网络用于基于第一视线拼接特征确定参考视线差参数;
第二主干网络含有第二基本模块,被配置为输出第二眼部图像特征,方向预测网络用于基于第二眼部图像特征确定第一预测视线参数和第二预测视线参数。
作为一种可能的实现方式,第一主干网络含有多层第一基本模块,第二主干网络含有多层第二基本模块,每层第一基本模块与第二基本模块对应,第二损失参数由多层眼部特征差异参数确定,每层眼部特征差异参数由对应层第一基本模块输出的第一眼部图像特征和对应层第二基本模块输出的第二眼部图像特征确定。
作为一种可能的实现方式,当预测视线信息还包括预测视线差参数时,待构建视线预测模型还包括第二特征拼接网络和第二差分预测网络,第二特征拼接网络用于基于第二眼部图像特征确定第二视线拼接特征,第二差分预测网络用于基于第二视线拼接特征确定预测视线差参数。
本公开示例性实施例还提供一种视线预测装置,该视线预测装置可以为电子设备或应用于电子设备的芯片。图5示出了本公开示例性实施例的视线预测装置的模块示意性框图。如图5所示,所述装置500包括:
获取模块501,被配置为获取目标眼部图像;
预测模块502,被配置为将目标眼部图像输入至视线预测模型,获得目标眼部图像对应的视线参数,视线预测模型由本公开示例性实施例所述的装置训练。
上述主要对本公开实施例提供的方案进行了介绍。可以理解的是,为了实现上述功能,电子设备包含了执行各个功能相应的硬件结构和/或软件模块。本领域技术人员应该很容易意识到,结合本文中所公开的实施例描述的各示例的单元及算法步骤,本公开能够以硬件 或硬件和计算机软件的结合形式来实现。某个功能究竟以硬件还是计算机软件驱动硬件的方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本公开的范围。
图6示出了本公开示例性实施例的芯片的示意性框图。如图6所示,该芯片600包括一个或两个以上(包括两个)处理器601和通信接口602。通信接口602可以支持服务器执行上述方法中的数据收发步骤,处理器601可以支持服务器执行上述方法中的数据处理步骤。
可选的,如图6所示,该芯片600还包括存储器603,存储器603可以包括只读存储器和随机存取存储器,并向处理器提供操作指令和数据。存储器的一部分还可以包括非易失性随机存取存储器(non-volatile random access memory,NVRAM)。
在一些实施方式中,如图6所示,处理器601通过调用存储器存储的操作指令(该操作指令可存储在操作系统中),执行相应的操作。处理器601控制终端设备中任一个的处理操作,处理器还可以称为中央处理单元(central processing unit,CPU)。存储器603可以包括只读存储器和随机存取存储器,并向处理器601提供指令和数据。存储器603的一部分还可以包括NVRAM。例如应用中存储器、通信接口以及存储器通过总线系统耦合在一起,其中总线系统除包括数据总线之外,还可以包括电源总线、控制总线和状态信号总线等。但是为了清楚说明起见,在图6中将各种总线都标为总线系统604。
上述本公开实施例揭示的方法可以应用于处理器中,或者由处理器实现。处理器可能是一种集成电路芯片,具有信号的处理能力。在实现过程中,上述方法的各步骤可以通过处理器中的硬件的集成逻辑电路或者软件形式的指令完成。上述的处理器可以是通用处理器、数字信号处理器(digital signal processing,DSP)、ASIC、现成可编程门阵列(field-programmable gate array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件。可以实现或者执行本公开实施例中的公开的各方法、步骤及逻辑框图。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。结合本公开实施例所公开的方法的步骤可以直接体现为硬件译码处理器执行完成,或者用译码处理器中的硬件及软件模块组合执行完成。软件模块可以位于随机存储器,闪存、只读存储器,可编程只读存储器或者电可擦写可编程存储器、寄存器等本领域成熟的存储介质中。该存储介质位于存储器,处理器读取存储器中的信息,结合其硬件完成上述方法的步骤。
本公开示例性实施例还提供一种电子设备,包括:至少一个处理器;以及与至少一个处理器通信连接的存储器。所述存储器存储有能够被所述至少一个处理器执行的计算机程序和/或指令,所述计算机程序和/或指令在被所述至少一个处理器执行时使所述电子设备执行根据本公开实施例的方法。
本公开示例性实施例还提供一种存储有计算机程序和/或指令的非瞬时计算机可读存 储介质,其中,所述计算机程序在被计算机的处理器执行时使所述计算机执行根据本公开实施例的方法。
本公开示例性实施例还提供一种计算机程序产品,包括计算机程序和/或指令,其中,所述计算机程序在被计算机的处理器执行时使所述计算机执行根据本公开实施例的方法。
本公开示例性实施例还提供一种计算机程序,包括程序代码,该程序代码在由处理器执行时使得所述处理器执行根据本公开实施例的方法。
参考图7,现将描述可以作为本公开的服务器或客户端的电子设备700的结构框图,其是可以应用于本公开的各方面的硬件设备的示例。电子设备旨在表示各种形式的数字电子的计算机设备,诸如,膝上型计算机、台式计算机、工作台、个人数字助理、服务器、刀片式服务器、大型计算机、和其它适合的计算机。电子设备还可以表示各种形式的移动装置,诸如,个人数字处理、蜂窝电话、智能电话、可穿戴设备和其它类似的计算装置。本文所示的部件、它们的连接和关系、以及它们的功能仅仅作为示例,并且不意在限制本文中描述的和/或者要求的本公开的实现。
如图7所示,电子设备700包括计算单元701,其可以根据存储在只读存储器(ROM)702中的计算机程序或者从存储单元708加载到随机访问存储器(RAM)703中的计算机程序,来执行各种适当的动作和处理。在RAM 703中,还可存储设备700操作所需的各种程序和数据。计算单元701、ROM 702以及RAM 703通过总线704彼此相连。输入/输出(I/O)接口705也连接至总线704。
电子设备700中的多个部件连接至I/O接口705,包括:输入单元706、输出单元707、存储单元708以及通信单元709。输入单元707可以是能向电子设备700输入信息的任何类型的设备,输入单元706可以接收输入的数字或字符信息,以及产生与电子设备的用户设置和/或功能控制有关的键信号输入。输出单元707可以是能呈现信息的任何类型的设备,并且可以包括但不限于显示器、扬声器、视频/音频输出终端、振动器和/或打印机。存储单元708可以包括但不限于磁盘、光盘。通信单元709允许电子设备700通过诸如因特网的计算机网络和/或各种电信网络与其他设备交换信息/数据,并且可以包括但不限于调制解调器、网卡、红外通信设备、无线通信收发机和/或芯片组,例如蓝牙TM设备、WiFi设备、WiMax设备、蜂窝通信设备和/或类似物。
如图7所示,计算单元701可以是各种具有处理和计算能力的通用和/或专用处理组件。计算单元701的一些示例包括但不限于中央处理单元(CPU)、图形处理单元(GPU)、各种专用的人工智能(AI)计算芯片、各种运行机器学习模型算法的计算单元、数字信号处理器(DSP)、以及任何适当的处理器、控制器、微控制器等。计算单元701执行上文所描述的各个方法和处理。例如,在一些实施例中,本公开示例性实施例的方法可被实现为计算机软件程序,其被有形地包含于机器可读介质,例如存储单元708。在一些实施例 中,计算机程序的部分或者全部可以经由ROM 702和/或通信单元709而被载入和/或安装到电子设备700上。在一些实施例中,计算单元701可以通过其他任何适当的方式(例如,借助于固件)而被配置为执行方法。
用于实施本公开的方法的程序代码可以采用一个或多个编程语言的任何组合来编写。这些程序代码可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理器或控制器,使得程序代码当由处理器或控制器执行时使流程图和/或框图中所规定的功能/操作被实施。程序代码可以完全在机器上执行、部分地在机器上执行,作为独立软件包部分地在机器上执行且部分地在远程机器上执行或完全在远程机器或服务器上执行。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
如本公开使用的,术语“机器可读介质”和“计算机可读介质”指的是用于将机器指令和/或数据提供给可编程处理器的任何计算机程序产品、设备、和/或装置(例如,磁盘、光盘、存储器、可编程逻辑装置(PLD)),包括,接收作为机器可读信号的机器指令的机器可读介质。术语“机器可读信号”指的是用于将机器指令和/或数据提供给可编程处理器的任何信号。
为了提供与用户的交互,可以在计算机上实施此处描述的系统和技术,该计算机具有:用于向用户显示信息的显示装置(例如,CRT(阴极射线管)或者LCD(液晶显示器)监视器);以及键盘和指向装置(例如,鼠标或者轨迹球),用户可以通过该键盘和该指向装置来将输入提供给计算机。其它种类的装置还可以用于提供与用户的交互;例如,提供给用户的反馈可以是任何形式的传感反馈(例如,视觉反馈、听觉反馈、或者触觉反馈);并且可以用任何形式(包括声输入、语音输入或者、触觉输入)来接收来自用户的输入。
可以将此处描述的系统和技术实施在包括后台部件的计算系统(例如,作为数据服务器)、或者包括中间件部件的计算系统(例如,应用服务器)、或者包括前端部件的计算系统(例如,具有图形用户界面或者网络浏览器的用户计算机,用户可以通过该图形用户界面或者该网络浏览器来与此处描述的系统和技术的实施方式交互)、或者包括这种后台部件、中间件部件、或者前端部件的任何组合的计算系统中。可以通过任何形式或者介质的数字数据通信(例如,通信网络)来将系统的部件相互连接。通信网络的示例包括:局域网(LAN)、广域网(WAN)和互联网。
计算机系统可以包括客户端和服务器。客户端和服务器一般远离彼此并且通常通过通信网络进行交互。通过在相应的计算机上运行并且彼此具有客户端-服务器关系的计算机程序来产生客户端和服务器的关系。
在上述实施例中,可以全部或部分地通过软件、硬件、固件或者其任意组合来实现。当使用软件实现时,可以全部或部分地以计算机程序产品的形式实现。所述计算机程序产品包括一个或多个计算机程序或指令。在计算机上加载和执行所述计算机程序或指令时,全部或部分地执行本公开实施例所述的流程或功能。所述计算机可以是通用计算机、专用计算机、计算机网络、终端、用户设备或者其它可编程装置。所述计算机程序或指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一个计算机可读存储介质传输,例如,所述计算机程序或指令可以从一个网站站点、计算机、服务器或数据中心通过有线或无线方式向另一个网站站点、计算机、服务器或数据中心进行传输。所述计算机可读存储介质可以是计算机能够存取的任何可用介质或者是集成一个或多个可用介质的服务器、数据中心等数据存储设备。所述可用介质可以是磁性介质,例如,软盘、硬盘、磁带;也可以是光介质,例如,数字视频光盘(digital video disc,DVD);还可以是半导体介质,例如,固态硬盘(solid state drive,SSD)。
尽管结合具体特征及其实施例对本公开进行了描述,显而易见的,在不脱离本公开的精神和范围的情况下,可对其进行各种修改和组合。相应地,本说明书和附图仅仅是所附权利要求所界定的本公开的示例性说明,且视为已覆盖本公开范围内的任意和所有修改、变化、组合或等同物。显然,本领域的技术人员可以对本公开进行各种改动和变型而不脱离本公开的精神和范围。这样,倘若本公开的这些修改和变型属于本公开权利要求及其等同技术的范围之内,则本公开也意图包括这些改动和变型在内。

Claims (19)

  1. 一种视线预测模型的构建方法,所述方法包括:
    将第一眼部样本图像和第二眼部样本图像输入辅助预测模型,获得视线差描述信息,所述视线差描述信息用于表征所述第一眼部样本图像和所述第二眼部样本图像的视线差;
    将所述第一眼部样本图像和所述第二眼部样本图像输入待构建视线预测模型,获得预测视线信息,所述预测视线信息包括所述第一眼部样本图像的第一预测视线参数和所述第二眼部样本图像的第二预测视线参数;
    基于所述视线差描述信息和所述预测视线信息确定所述待构建视线预测模型的损失;
    在所述待构建视线预测模型的损失未满足收敛条件的情况下,基于所述视线差描述信息和所述预测视线信息更新所述待构建视线预测模型的模型参数;和/或,在所述待构建视线预测模型的损失满足收敛条件的情况下,确定所述待构建视线预测模型为视线预测模型。
  2. 根据权利要求1所述的方法,其中,所述视线差描述信息包括第一眼部图像特征和参考视线差参数,所述预测视线信息还包括第二眼部图像特征,所述待构建视线预测模型的损失包括第一损失参数和第二损失参数;
    其中,所述第一损失参数由所述第一预测视线参数、所述第二预测视线参数和眼部样本图像的监督信息确定,所述眼部样本图像的监督信息包括所述参考视线差参数,所述第二损失参数由所述第一眼部图像特征和所述第二眼部图像特征确定。
  3. 根据权利要求2所述的方法,其中,所述待构建视线预测模型的损失由所述第一损失参数和所述第二损失参数的加权结果确定,所述第二损失参数的权重大于所述第一损失参数的权重。
  4. 根据权利要求2或3所述的方法,其中,所述第一损失参数包括第一子损失参数和第二子损失参数;
    所述第一子损失参数由所述参考视线差参数和视线差异参数确定,所述视线差异参数由所述第一预测视线参数和所述第二预测视线参数的差值确定;
    所述眼部样本图像的监督信息还包括所述第一眼部样本图像的第一标注视线方向和所述第二眼部样本图像的第二标注视线方向,所述第二子损失参数由第一预测损失参数和第二预测损失参数确定,所述第一预测损失参数由所述第一标注视线方向和所述第一预测视线参数确定,所述第二预测损失参数由所述第二标注视线方向和所述第二预测视线参数确定。
  5. 根据权利要求1~4中任一项所述的方法,其中,所述方法还包括:
    基于所述参考视线差参数和标注视线差参数,更新所述辅助预测模型的模型参数,所述标注视线差参数由所述第一眼部样本图像的第一标注视线方向和所述第二眼部样本图像的第二标注视线方向确定。
  6. 根据权利要求2~5中任一项所述的方法,其中,所述预测视线信息还包括预测视线差参数,所述待构建视线预测模型的损失还包括第三损失参数,所述第三损失参数由所述预测视线差参数和所述参考视线差参数确定,并且
    其中,所述待构建视线预测模型的损失由所述第一损失参数、所述第二损失参数和所述第三损失参数的加权结果确定,所述第一损失参数的权重和所述第三损失参数的权重均小于所述第二损失参数的权重。
  7. 根据权利要求1~6中任一项所述的方法,其中,所述方法还包括:
    基于所述第一眼部样本图像的第一标注视线方向和所述第二眼部样本图像的第二标注视线方向确定标注视线差参数;
    基于所述标注视线差参数对所述参考视线差参数进行校准。
  8. 根据权利要求2~7任一项所述的方法,其中,所述辅助预测模型包括第一主干网络、第一特征拼接网络和第一差分预测网络,所述待构建视线预测模型包括第二主干网络和方向预测网络;
    所述第一主干网络含有第一基本模块,用于输出所述第一眼部图像特征,所述第一特征拼接网络用于基于所述第一眼部图像特征确定第一视线拼接特征,所述第一差分预测网络用于基于所述第一视线拼接特征确定所述参考视线差参数;
    所述第二主干网络含有第二基本模块,用于输出所述第二眼部图像特征,所述方向预测网络用于基于所述第二眼部图像特征确定所述第一预测视线参数和所述第二预测视线参数。
  9. 根据权利要求8所述的方法,其中,所述第一主干网络含有多层第一基本模块,所述第二主干网络含有多层第二基本模块,每层所述第一基本模块与所述第二基本模块对应,所述第二损失参数由多层眼部特征差异参数确定,每层所述眼部特征差异参数由对应层所述第一基本模块输出的第一眼部图像特征和对应层所述第二基本模块输出的第二眼部图像特征确定。
  10. 根据权利要求2-8中任一项所述的方法,其中,当所述预测视线信息还包括预测视线差参数时,所述待构建视线预测模型还包括第二特征拼接网络和第二差分预测网络,所述第二特征拼接网络用于基于所述第二眼部图像特征确定第二视线拼接特征,所述第二差分预测网络用于基于所述第二视线拼接特征确定所述预测视线差参数。
  11. 一种视线预测方法,所述方法包括:
    获取目标眼部图像;
    将目标眼部图像输入至视线预测模型,获得目标眼部图像对应的视线参数,所述视线预测模型由根据权利要求1~10任一项所述的方法得到。
  12. 一种视线预测模型的构建装置,所述装置包括:
    训练模块,被配置为将第一眼部样本图像和第二眼部样本图像输入辅助预测模型,获得视线差描述信息,所述视线差描述信息用于表征所述第一眼部样本图像和所述第二眼部样本图像的视线差;将所述第一眼部样本图像和所述第二眼部样本图像输入待构建视线预测模型,获得预测视线信息,所述预测视线信息包括所述第一眼部样本图像的第一预测视线参数和所述第二眼部样本图像的第二预测视线参数;基于所述视线差描述信息和所述预测视线信息确定所述待构建视线预测模型的损失;在所述待构建视线预测模型的损失未满足收敛条件的情况下,基于所述视线差描述信息和所述预测视线信息更新所述待构建视线预测模型的模型参数;
    确定模块,被配置为在所述待构建视线预测模型的损失满足收敛条件的情况下,确定所述待构建视线预测模型为视线预测模型。
  13. 根据权利要求12所述的视线预测模型的构建装置,其中,所述训练模块还被配置为:
    基于包含在视线差描述信息中的参考视线差参数、以及标注视线差参数,更新所述辅助预测模型的模型参数,所述标注视线差参数由所述第一眼部样本图像的第一标注视线方向和所述第二眼部样本图像的第二标注视线方向确定。
  14. 根据权利要求12所述的视线预测模型的构建装置,其中,所述训练模块还被配置为:基于所述第一眼部样本图像的第一标注视线方向和所述第二眼部样本图像的第二标注视线方向确定标注视线差参数;并且,所述装置还包括:
    校准模块,被配置为基于所述标注视线差参数对包含在视线差描述信息中的参考视线差参数进行校准。
  15. 一种视线预测装置,所述装置包括:
    获取模块,被配置为获取目标眼部图像;
    预测模块,用于将目标眼部图像输入至视线预测模型,获得目标眼部图像对应的视线,所述视线预测模型由权利要求12~14中任一项所述的视线预测模型的构建装置训练。
  16. 一种电子设备,包括:
    处理器;以及,
    存储程序和/或指令的存储器;
    其中,所述程序和/或指令在由所述处理器执行时使所述处理器执行根据权利要求1~11任一项所述的方法。
  17. 一种存储有计算机指令的非瞬时计算机可读存储介质,所述计算机指令在由计算机执行时使得所述计算机执行根据权利要求1~11中任一项所述的方法。
  18. 一种计算机程序产品,包含指令,该指令在由处理器执行时使得处理器实现根据权利要求1~11中任一项所述的方法。
  19. 一种计算机程序,包括程序代码,该程序代码在由处理器执行时导致实现根据权利要求1~11中任一项所述的方法。
PCT/CN2024/096691 2023-06-07 2024-05-31 视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质 Ceased WO2024251044A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310668262.4 2023-06-07
CN202310668262.4A CN119107513A (zh) 2023-06-07 2023-06-07 视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质

Publications (1)

Publication Number Publication Date
WO2024251044A1 true WO2024251044A1 (zh) 2024-12-12

Family

ID=93720347

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/096691 Ceased WO2024251044A1 (zh) 2023-06-07 2024-05-31 视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质

Country Status (2)

Country Link
CN (1) CN119107513A (zh)
WO (1) WO2024251044A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120071422A (zh) * 2025-01-13 2025-05-30 西安电子科技大学 一种基于最小化多支路网络旋转方差的跨域视线预测方法、装置、电子设备及存储介质

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111796681A (zh) * 2020-07-07 2020-10-20 重庆邮电大学 人机交互中基于差分卷积的自适应视线估计方法及介质
CN113506328A (zh) * 2021-07-16 2021-10-15 北京地平线信息技术有限公司 视线估计模型的生成方法和装置、视线估计方法和装置
CN113705550A (zh) * 2021-10-29 2021-11-26 北京世纪好未来教育科技有限公司 一种训练方法、视线检测方法、装置和电子设备

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111796681A (zh) * 2020-07-07 2020-10-20 重庆邮电大学 人机交互中基于差分卷积的自适应视线估计方法及介质
CN113506328A (zh) * 2021-07-16 2021-10-15 北京地平线信息技术有限公司 视线估计模型的生成方法和装置、视线估计方法和装置
CN113705550A (zh) * 2021-10-29 2021-11-26 北京世纪好未来教育科技有限公司 一种训练方法、视线检测方法、装置和电子设备

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120071422A (zh) * 2025-01-13 2025-05-30 西安电子科技大学 一种基于最小化多支路网络旋转方差的跨域视线预测方法、装置、电子设备及存储介质

Also Published As

Publication number Publication date
CN119107513A (zh) 2024-12-10

Similar Documents

Publication Publication Date Title
JP7331171B2 (ja) 画像認識モデルをトレーニングするための方法および装置、画像を認識するための方法および装置、電子機器、記憶媒体、並びにコンピュータプログラム
CN113707299A (zh) 基于问诊会话的辅助诊断方法、装置及计算机设备
CN112541122A (zh) 推荐模型的训练方法、装置、电子设备及存储介质
CN114037003B (zh) 问答模型的训练方法、装置及电子设备
WO2021135449A1 (zh) 基于深度强化学习的数据分类方法、装置、设备及介质
CN114443034A (zh) 优化界面布局的方法、装置、设备及介质
WO2025091924A1 (zh) 人机交互方法、装置、电子设备以及存储介质
CN115082920A (zh) 深度学习模型的训练方法、图像处理方法和装置
CN111859985B (zh) Ai客服模型测试方法、装置、电子设备及存储介质
CN114445682A (zh) 训练模型的方法、装置、电子设备、存储介质及产品
WO2024146468A1 (zh) 模型生成方法及装置、实体识别方法及装置、电子设备、存储介质
WO2024251044A1 (zh) 视线预测模型的构建方法、视线预测方法、装置及电子设备和存储介质
WO2021072864A1 (zh) 文本相似度获取方法、装置、电子设备及计算机可读存储介质
CN115719433A (zh) 图像分类模型的训练方法、装置及电子设备
CN114969543A (zh) 推广方法、系统、电子设备和存储介质
CN113032469B (zh) 文本结构化模型训练、医疗文本结构化方法及装置
CN118485819A (zh) 一种关键点检测方法、装置、设备及存储介质
US20240037410A1 (en) Method for model aggregation in federated learning, server, device, and storage medium
US20230122373A1 (en) Method for training depth estimation model, electronic device, and storage medium
CN118037418A (zh) 一种信用风险预测方法、装置及电子设备
CN112446192A (zh) 用于生成文本标注模型的方法、装置、电子设备和介质
WO2024007938A1 (zh) 一种多任务预测方法、装置、电子设备及存储介质
US20220004801A1 (en) Image processing and training for a neural network
CN116167846A (zh) 校准方法、装置、电子设备及计算机可读存储介质
CN115131709A (zh) 视频类别预测方法、视频类别预测模型的训练方法及装置

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24818576

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE