EP4720825A2 - Detecting reading state of a user based on visual attention information and semantic features of text - Google Patents

Detecting reading state of a user based on visual attention information and semantic features of text

Info

Publication number
EP4720825A2
EP4720825A2 EP24816404.8A EP24816404A EP4720825A2 EP 4720825 A2 EP4720825 A2 EP 4720825A2 EP 24816404 A EP24816404 A EP 24816404A EP 4720825 A2 EP4720825 A2 EP 4720825A2
Authority
EP
European Patent Office
Prior art keywords
reading
user
text
determining
data stream
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24816404.8A
Other languages
German (de)
French (fr)
Inventor
Robert P. Dick
Li Shang
Qin LV
Yuhu Chang
Yingying ZHAO
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of Michigan System
Original Assignee
University of Michigan System
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of Michigan System filed Critical University of Michigan System
Publication of EP4720825A2 publication Critical patent/EP4720825A2/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/013Eye tracking input arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • G06F3/012Head tracking input arrangements
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/174Facial expression recognition

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Multimedia (AREA)
  • User Interface Of Digital Computer (AREA)
  • Image Analysis (AREA)

Abstract

A method for detecting a reading state of a user includes obtaining a first data stream indicative of one or both of eye movement and facial expression of the user as the user is reading a text, determining, based on the first data stream, visual information indicative of one or both of eye movements and facial expressions of the user during reading of the text, obtaining a second data stream indicative of content of the text, determining, based on the second data stream, semantic information from the text, determining reading state information based on analyzing the visual information determined based on the first data stream and the semantic information determined based on the second data stream, performing, by the processor, an operation with respect to the reading state information, wherein the reading state information tracking the reading state of the user during reading of the text.

Description

Atty. Docket No.010110-23010A DETECTING READING STATE OF A USER BASED ON VISUAL ATTENTION INFORMATION AND SEMANTIC FEATURES OF TEXT CROSS-REFERENCE TO RELATED APPLICATION [0001] This application claims the benefit of U.S. provisional application entitled “Detecting Reading State of a User Based on Visual Attention Information and Semantic Features of Text,” filed on May 30, 2023, and assigned Serial No.63/469,706, the entire disclosure of which is hereby expressly incorporated by reference. BACKGROUND OF THE DISCLOSURE Field of the Disclosure [0002] The disclosure relates generally to detecting reading states of a user as the user is reading a text. Brief Description of Related Technology [0003] The science of reading has attracted decades of interest in human-computer interaction (HCI), cognitive science, psychology, educational psychology, cognition and neuroscience, pedagogy, and brain science. Reading is a cognitive process and understanding it benefits numerous research communities. For example, reading is a fundamental approach to learning, through which people can expand their vocabulary, gain knowledge, and develop skills. Research has shown a positive relationship between reading and learning; for example, the more people read, the more effectively they improve vocabulary, knowledge levels, and cognitive skills. In fact, reading has long been considered the most important path to lifelong learning, and lifelong readers are generally more successful, both personally and professionally. [0004] Eye tracking has been used to explore eye-mind relationships. Eye-tracking technology can acquire real-time eye movements in a non-intrusive manner. Such eye movement data has been utilized to probe the reading process, as the reading process initiates visual input and operates as an interactive eye-mind cognition process. Analysis of eye movement data obtained during reading has been used to understand the reading cognitive process and provide reading assistance. However, eye-tracking technologies suffer from a number of shortcomings. The error of typical eye-tracking technologies ranges from 1 to 4 degrees. Under reading scenarios, this angular accuracy translates to a spatial tracking resolution of about 1.4–2.6 cm. Considering a computerized-reading task where the distance from eye to screen is 40–50 cm, this means that the resolution of the eye tracker is about 3 to 4 lines for a single- Atty. Docket No.010110-23010A spaced document and about 1 to 3 words in the horizontal direction. Such low spatial resolution makes it difficult or impossible to track reading states during word-by-word and line-by-line reading because the words and lines cannot be located with sufficient accuracy. This problem is typically tackled by using an unrealistic setting with a very wide line spacing (e.g., triple-spaced), leaving them unsuitable for use with normally spaced text. In addition, eye-tracking techniques are subject to the inherent transient jitter of human gaze and vertical drift, which require constant calibration. SUMMARY OF THE DISCLOSURE [0005] In accordance with one aspect of the disclosure, a method for detecting a reading state of a user includes obtaining, by a processor, a first data stream indicative of one or both of i) eye movement and ii) facial expression of the user as the user is reading a text, determining, by the processor based on the first data stream, visual information indicative of one or both of i) eye movements and ii) facial expressions of the during reading of the text, obtaining, by the processor, a second data stream indicative of content of the text, determining, by the processor based on the second data stream, semantic information from the text, determining, by the processor, reading state information based on analyzing the visual information determined based on the first data stream and the semantic information determined based on the second data stream, and performing, by the processor, an operation with respect to the reading state information, wherein the reading state information tracking the reading state of the user during reading of the text. [0006] In accordance with another aspect of the disclosure, a system comprises a first sensor configured to generate a first data stream indicative of one or both of i) eye movement and ii) facial expression of a user as the user is reading a text, a second sensor configured to generate a second data stream indicative of content of the text, and a reading cognition analysis engine implemented on one or more integrated circuits, the reading cognition analysis engine configured to obtain the first data stream from the first sensor, determine, based on the first data stream, visual information indicative of one or both of i) eye movements and ii) facial expressions of the user during reading of the text, obtain the second data stream from the second sensor, determine, based on the second data stream, semantic information from the text, determine reading state information based on analyzing the visual information determined based on the first data stream and the semantic information determined based on the second Atty. Docket No.010110-23010A data stream, wherein the reading state information tracking the reading state of the user during reading of the text, and perform an operation with respect to the reading state information. [0007] In connection with any one of the aforementioned aspects, the devices and/or methods described herein may alternatively or additionally include or involve any combination of one or more of the following aspects or features. [0008] Determining the visual information includes determining, based on the first data stream, a gaze point sequence indicative of a direction of gaze of the user as the user is reading the text, and extracting visual attention features from the gaze point sequence. Determining the gaze point sequence includes determining, based on the first data stream, an initial gaze point sequence, and performing one or both of i) median filtering and ii) mean filtering of the initial gaze point sequence to generate a smoothed gaze point sequence. Extracting the visual attention features comprises extracting the visual attention features from the smoothed gaze point sequence. Extracting the visual attention features includes determining, based on the gaze point sequence, visual attention features indicative of one or both of i) word-level processing state of the user and iii) segment-level processing state of the user. Extracting the visual attention features includes performing normalization of the visual attention features based at least in part on statistical data determined based on visual attention information obtained while the user is reading the text. Determining the visual attention features indicative of the segment level processing state of the user includes determining visual attention features indicative of one or more of i) phrase-level processing state of the user, ii) sentence-level processing state of the user, ii) paragraph-level processing state of the user, iv) page-level processing state of the user, or v) document-level processing state of the user. Determining the visual attention features indicative of the word-level processing state of the user includes, determining, for each word among at least some words of the text, one or more of i) fixation duration on the word, ii) number of fixations on the word, or iii) number of repeated readings of the word. Determining the visual attention features indicative of the segment-level processing state of the user includes determining, for each of at least some text segments of the text being read by the user, i) dwell time of the user on the text segment, ii) forward saccade time exhibited by the user as the user is reading the text segment and iii) backward saccade time exhibited by the user as the user is reading the text segment. Determining the visual attention features further includes using a deep neural network (DNN) to extract deep features from the gaze point sequence. Determining the visual information includes determining Atty. Docket No.010110-23010A facial features indicative of emotional state of the user as the user is reading the text. Determining the semantic information includes generating a first feature vector to include encodings of respective words of the text, generating a second feature vector to include probabilities of the respective words of the text describing content of the text, generating a third feature vector to include difficulty scores for the respective words of the text, and determining semantic features based on a combined sematic feature vector that includes concatenation of the first feature vector, the second feature vector, and the third feature vector. Determining the semantic information further includes concatenating the semantic features with the visual attention features to generate a concatenated feature vector, including, when fusing the semantic features with the visual attention features at a time t, identifying a first set of words being processed at the time t, and zero-padding semantic features of a second set of words that has not yet been processed at the time t, and determining the reading state information based on analyzing the concatenated feature vector. Determining the reading state information comprises determining one or more of i) that the user is having difficulty comprehending a particular word in the text, ii) that the user is having difficulty comprehending a particular text segment in the text, or iii) that mind of the user is wandering as the user is reading the text. Performing the operation with respect to the reading state comprises one or more of i) providing an after-reading summary to the user or ii) providing an in-time intervention to the user. Providing the in-time intervention to the user comprises one or more of i) providing a definition of a word, with which the user is having difficulty, to be displayed in real-time on a screen as the user is reading the text on the screen, ii) providing an explanation of a sentence, with which the user is having difficulty, to be displayed in real-time on a screen as the user is reading the text on the screen, or iii) providing an indication that the user’s mind is wandering to be displayed in real-time on a screen as the user is reading the text on the screen. Obtaining the first data stream comprises obtaining one or more images depicting an eye region of a face of the user, and obtaining the second data stream comprises obtaining one or more images depicting the text. BRIEF DESCRIPTION OF THE DRAWING FIGURES [0009] For a more complete understanding of the disclosure, reference should be made to the following detailed description and accompanying drawing figures, in which like reference numerals identify like elements in the figures. Atty. Docket No.010110-23010A [0010] Fig.1 is a block diagram of a reading state detection system in accordance with an example. [0011] Fig.2 is a diagram depicting an example implementation of the reading state detection system of Fig.1 in accordance with an example. [0012] Figs.3A-C illustrate several examples of user intervention assistance provided in the system of Fig.1 in accordance with examples. [0013] Fig.4 is a block diagram of an example architecture of a reading state detection and intervention system in accordance with an example. [0014] Fig.5 depicts example smart glasses used with a reading state detection system in accordance with an example. [0015] Figs.6A-C illustrate a data collection method in accordance with an example. [0016] Fig.7 depicts reading state recognition performance comparison between two baseline methods and a method in accordance with an example. [0017] Figs.8A-C are plots illustrating Area Under the Curve (AUC) for two baseline methods and a method in accordance with an example. [0018] Fig.9 illustrates visual attention for a user reading a sentence as determined by reading state detection in accordance with an example. [0019] Fig.10 is a bar chart illustrating a profile of word-level comprehension failures for a number of users in accordance with an example. [0020] Fig.11 is a diagram illustrating an example in which two users have difficulty understanding a word in accordance with an example. [0021] Fig.12 is a diagram illustrating an example in which two users are reading the same sentence but with different visual attention in accordance with an example. [0022] Fig.13 is a diagram illustrating an example of two readers having different “reread” behaviors when in the same reading state: encountering comprehension difficulties on the sentence in accordance with an example. [0023] Fig.14 is a diagram illustrating an example in which two users are reading the same sentence Atty. Docket No.010110-23010A twice in accordance with an example. [0024] Fig.15 is a bar chart illustrating a profile of word-level comprehension failures for a number of users in accordance with an example. [0025] Fig.16 depicts a method for reading state detection in accordance with an example. [0026] Fig.17 is a block diagram of a computing system with which aspects of the disclosure may be practiced. [0027] The aspects of the disclosed systems and methods may assume various forms. Specific embodiments are illustrated in the drawing and hereafter described with the understanding that the disclosure is intended to be illustrative. The disclosure is not intended to limit the invention to the specific embodiments described and illustrated herein. DETAILED DESCRIPTION OF THE DISCLOSURE [0028] Systems and methods are provided that detect reading states of a user based on both semantic context of a text being read by the user and visual attention features obtained based on tracking of user’s eye movements and/or facial expressions as the user is reading the text. In an example, the semantic content of the text may be obtained from one or more images of the text. For example, natural language processing (NLP) techniques may be used to process the text to obtain semantic features of the text. The semantic features may include encoded words of the text, probabilities of respective words describing entirety of the text, word difficulty scores, or the like. Eye-tracking and/or facial expression detection may be performed based on images or other recordings of the user’s eyes and/or features of the user’s face as the user is reading the text. Eye tracking may include obtaining points of gaze of the user as the user is reading the text, and extracting various levels of visual attention features from the points of gaze of the user. The visual attention features may include word-level visual attention features, such as duration of fixation on a word, number of fixations on a word, number of repeated readings of a word, or the like. The visual attention features may also include segment-level features corresponding to text segments, such as phrases, sentences, passages, etc. that include multiple words. Segment-level visual attention features may include dwell time on a segment, saccade times exhibited by the user’s eyes as the user is reading the segment, forward saccade times exhibited by the user’s eyes as the user Atty. Docket No.010110-23010A is reading the segment, backward saccade times exhibited by the user’s eyes as the user is reading the segment, or the like. In examples, the word-level and segment-level features may be normalized based on statistical data to reduce or eliminate person-to-person reading variations as well as individual during- reading variations as the user is reading the text. Additionally or alternatively, deep features extracted from the points of gaze of the user may be used. Facial features may include features indicative of emotional state of the user as the user is reading the text, such as whether the user is calm, engaged, interested, disinterested, confused, frustrated, etc. As described in more detail below, in at least some examples, using both semantic context of a text being read by the user and visual attention features obtained based on tracking of user’s eye movements and/or facial expressions as the user is reading the text improves accuracy of reading state detection and allows the disclosed systems and methods to operate on relatively closely spaced texts such as texts typically encountered on computer screens or the like. [0029] In aspects, the disclosed systems and methods may time-align and fuse together the semantic features extracted from the text and the visual attention and/or facial expression features of the user as the user is reading the text. In an example, time alignment of the semantic features and the visual attention and/or facial expression features may include identifying a set of words being processed at a time t and focusing on the sematic features of the set of words while de-emphasizing or zero-padding sematic features of the other words. Based on the time-aligned and fused semantic features and visual attention and/or facial expression features, the disclosed systems and methods may detect reading states of the user as the user is reading the text. For example, the disclosed systems and methods may detect word-level and/or segment-level processing difficulties exhibited by the user as the user is reading various words and segments of the text. The disclosed systems and methods may also detect whether the user’s mind is wandering as the user is reading the text. In aspects, the disclosed systems and methods also provide interventions that may assist the user as the user is reading the text. In examples, interventions may include providing definitions or explanations of words with which the user is having difficulties, providing explanations or summarizations of text passages with which the user is struggling, signaling to the user that the user’s mind is wandering, and/or the like. In an example, the interventions may be provided in real-time as the reader is reading the text. For example, words and/or passages with which the user is having difficulties may be highlighted in real-time on a screen from which the user is reading the text. As another examples, word definitions, text explanations, indications of a wandering Atty. Docket No.010110-23010A mind, or the like may be displayed on the screen at times at which the user is encountering the corresponding difficulties as the user is reading the text on the screen. In other examples, the explanations and interventions may be provided in other suitable manners. For example, a reading state and/or processing difficulty summary, word definitions, passage explanations, etc. may be provided to a user after the user has finished reading the text. [0030] In aspects, a reading state detection system in accordance with an example may be implemented in eyewear. The use of the system may thus not interrupt the reading process, which may reduce subjective bias, in at least some examples. In aspects, the real-time association between visual and semantic information enables the interactions between visual attention and semantic context to be better interpreted and explained. In aspects, an interactive reading assistance system is also provided. For example, the assistance system may provide just-in-time intervention when readers encounter processing difficulties, thus promoting self-awareness of the cognitive process involved in reading and helping the readers to develop more effective reading habits. [0031] In various examples, the reading state detection system is equipped with a deep neural network that extracts features pertaining to the visual attention history and text semantic content. In aspects, the system may fuse the two types of features via a shared convolutional filter mechanism based on temporal convolutional network (TCN) technology to enable accurate reading state estimation at various granularities. [0032] In an example, an interactive reading assistant system that utilizes the reading state detection system is provided. In-field studies have been conducted to demonstrate that the reading state detection system integrated with the interactive assistance system enables helpful interventions for readers, thus improving self-awareness in the reading process and helping readers adopt more effective reading habits. [0033] The disclosed systems and methods are suitable for use in wearable and/or other battery- powered and/or embedded systems, but not limited thereto. Disclosed systems and methods may be implemented at least partially locally in a wearable device, such as a camera system integrated with eyewear. In some examples, any one or more aspects of the data processing described herein may be implemented remotely, for example at a remote server. The number, location, arrangement, configuration, and other characteristics of the processor(s) of the disclosed systems, or the processor(s) Atty. Docket No.010110-23010A used to implement the disclosed methods, may vary accordingly. [0034] In various examples, the reading state detection system may probe and explain human cognitive processes while a reader is reading a text. Reading sate detection as described herein may thus be used in the study of reading and learning to read. In some examples, reading sate detection as described herein may be used in HCI and educational application investigations on improving reading productivity. More generally, the disclosed systems and methods may be used in various reading applications. For example, the disclosed systems and methods may be utilized in studying how people understand the semantics and syntax of text, which may, in turn, aid in understanding natural language representation and processing, which are key functionalities of human-level intelligence. Understanding the reading process can also advance the theory of human behavior, thus benefiting the domains of applied psychology, pedagogy, and educational psychology. For instance, human cognitive abilities such as verbal working memory capacity, inhibitory control ability, perceptual speed, and immediate and delayed effects on reading processes may be scrutinized. Furthermore, understanding how people read sheds light on reading patterns and strategies, potentially helping readers achieve metacognitive awareness and read more efficiently. For example, human reading efficiency, reading proficiency, reading skills, reading comprehension performance, and reading outcomes may be enhanced. [0035] Fig.1 illustrates a cognition-aware smart eyewear system 100, sometimes referred herein as “CASES.” The system 100 includes a first sensor 102 and a second sensor 104. The first sensor 102 and/or the second sensor 104 may comprise a single sensor or may comprise multiple sensors, such as multiple sensors of different types. The first sensor 102 may be configured to capture information indicative of eye movement and gaze direction of the user. In some examples, the first sensor 102 may be additionally configured to capture information of an expression of one or both eyes of the user and/or facial expression of the user. In various examples, the first sensor 102 may comprise one or more of i) a camera, such as a visible light camera, an infrared (IR) camera, etc. that may be configured to capture images or videos depicting one or both eyes of the user, ii) an infrared sensor configured to capture eye movement, eye gaze direction and/or eye or facial expression information based on active IR illumination of one or both eyes of the user, iii) a camera configured to passively capture appearance of or one or both eyes of the user, etc. In some examples, the first sensor 102 may comprise one or more wearable position and/or orientation sensor devices, such as an accelerometer, a gyroscope, a magnetometer, Atty. Docket No.010110-23010A etc., that may be attached to the user (e.g., user’s head, user’s body, etc.), or to a wearable device (e.g., eyewear) that may be worn by the user, and may be configured to detect position and/or orientation of the user (e.g., user’s head and/or body) relative to the scene being viewed by the user, such as a screen displaying a text being read by the user. In an example, the orientation and/or position of the user relative to the scene being viewed by the user may be indicative of the eye movement and/or gaze direction of the user relative to the scene. In other examples, the first sensor 102 may additionally or alternatively comprise other suitable sensor devices that may be configured to capture or otherwise generate information indicative of eye movement, eye gaze direction and/or eye or facial expression of the user. [0036] The second sensor 104 may be configured to capture information indicative of the content of text being read by the user. For example, the second sensor 104 may be configured to capture image data, video data, etc. capturing the text read by the user. In various examples, the second sensor 104 may comprise one or more of i) a camera, such as a visible light camera, an infrared camera, etc., ii) a camcorder, iii) a video recorder, etc. In other examples, the second sensor 104 may additionally or alternatively comprise other suitable sensor devices that may be configured to capture or otherwise generate data, such as image or video data, indicative of visual content in the field of view of the user. [0037] In an example, the first sensor 102 and the second sensor 104 are mounted on eyewear, such as glasses or goggles, that may be worn by the user, with the first sensor 102 (sometimes referred to herein as “eye camera”) configured as an inward-facing sensor, e.g., facing the eyes of the user, and the second sensor 104 (sometimes referred to herein as “scene camera”) configured as a forward-facing sensor with respect to field of view of the user. In other examples, instead of being attached to a user or to a device worn by the user, the first sensor 102 and/or the second sensor 104 may be located at a suitable distance from the user. For example, the first sensor 102 and/or the second sensor 104 may be a distance sensor (e.g., distance camera) positioned in the vicinity of the user. As just an example, the first sensor 102 may be a web camera, or webcam, that may generally be facing the user as the user is reading text on a screen. [0038] In an example, the second sensor 104 is configured to capture images (e.g., video) or other data indicative of the text being read by the user, whereas the first sensor 102 is configured to concurrently capture images or other data indicative of the user’s eye movement, eye gaze and/or eye or facial expression of the user as the user is reading the text. The cognition-aware smart eyewear system Atty. Docket No.010110-23010A 100 may thus be configured to collect high-quality bi-modal data using the sensors 102 and 104 to record two modalities: the sensor 104 capturing text and the sensor 102 tracking gaze points and/or user’s facial expressions as the user is reading the text. In the illustrated example, the cognition-aware smart eyewear system 100 is implemented in the form of eyewear. The cognition-aware smart eyewear system 100 may thus avoid interfering with the reading process when collecting data. In some examples in which reader surveys are needed (e.g., in data collection examples described in more detail below), the surveys are deferred until after a reading task is completed, also to avoid interference. [0039] The cognition-aware smart eyewear system 100 may include a reading cognition analysis engine 120 (sometimes referred to herein as “Cognition-Aware Smart Eyewear System Network” or “CASES-Net”). The reading cognition analysis engine 120 may include a visual feature extractor 122 and a sematic feature extractor 124. The visual feature extractor 120 may be configured to extract expression-related information indicative of eye or facial expression of the user from images and/or other information obtained by the sensor 104. The sematic feature extractor 124 may be configured to extract semantic attributes of the text read by the user from images and/or other information obtained by the sensor 104. The reading cognition analysis engine 120 may be configured to determine a reading state of the user based on i) semantic attributes of the text read by the user extracted by the sematic feature extractor 124 from images and/or other information obtained by the sensor 104 and ii) expression-related information indicative of eye or facial expression of the user extracted by the visual feature extract 122 from images and/or other information obtained by the sensor 102. In an example, the reading cognition analysis engine 120 comprises a bi-modal multi-task network. The bi-modal multi-task network may be provided with the bi-modal data, i.e., the eye-tracking and reading text data as inputs, and may estimate, based on the bi-modal data, cognitive reading states in real-time at two granularities: word and sentence level. [0040] Reading science has attracted decades of interest in various research communities, e.g., HCI, pedagogy, and educational psychology. These studies primarily deal with the outcomes of reading and reading comprehension. For example, reading patterns and strategies that improve the efficiency of reading, e.g., interactive reading systems that detect mind wandering during reading have been explored. They mitigated the negative effect of mind wandering on reading comprehension using just-in-time interventions. Other methods detect words readers do not know and provide appropriate help. In Atty. Docket No.010110-23010A psychology, applied psychology, and educational psychology, it has been studied how texts are read and comprehended. For example, a blueprint of reading has been developed. The blueprint consists of the visual process, representation process that converts visual perception into a linguistic representation, and operation process on the representation. In cognition science, neuroscience, and brain science, extensive reading studies focus on developing computational theories of cognition. One important branch studies the representations and processing of natural languages by the human brain. For example, a theoretical framework to explain how verbal working memory supports sentence processing has been provided. It has been studied, for example, how the global and local information in texts impact sentence processing. Language representation and processing may be jointly considered because discovering language representation can help answer questions about computation, and vice versa. Computationally explicit evidence has shown that language comprehension mechanisms in human brains are fundamentally shaped by predictive processing through an integrative modeling approach. [0041] Generally, reading is a multi-level interactive eye-mind cognitive process. In the short term, readers visually perceive each word, encode it, and mentally assign semantics. In the long term, readers visually perceive a sentence and mentally associate it with context and domain knowledge. Reading can be viewed as a sequence of numerous time-varying states. For instance, some studies explored the state of mind wandering, to detect whether a reader is cognitively engaged or decoupled from the current reading task. Furthermore, the state of having difficulty processing unfamiliar words has been explored. However, processing difficulties can present at multiple granularities, e.g., readers may encounter difficulties at the level of a single word, a sentence, or a paragraph. In an example, the reading cognition analysis engine 120 is configured analyze the reading cognitive process to detect and explain multiple states at word and sentence levels. For example, the reading cognition analysis engine 120 may determine whether a reader’s mind is wandering, whether the reader is positively engaged, and when comprehension is delayed due to word-or sentence-level processing difficulties. In other examples, the reading cognition analysis engine 120 is configured to analyze the reading cognitive process to detect and explain reading states at granularities other than word and sentence levels. [0042] Eye movements are good indicators to infer the cognitive process. This is based on the eye mind hypothesis, which states that there is a close relationship between where the eyes look and where the mind is engaged. Eye tracking data is available and can be accessed to explore eye-mind Atty. Docket No.010110-23010A relationships. The relationship between eye movements and cognitive processes has been studied. An eye-tracking reading system may track the participants’ eye movements in a non-intrusive way. Hand- engineered eye movement features have been used to probe the reading cognitive process. [0043] However, as discussed above, eye-tracking technologies suffer from a number of shortcomings. The error of commercially available eye-tracking technologies typically ranges from 1 to 4 degrees. Under reading scenarios, this angular accuracy translates to a spatial tracking resolution of about 1.4–2.6 cm. Considering a computerized reading task where the distance from eye to screen is 40–50 cm, this means that the resolution of the eye tracker is about 3 to 4 lines for a single-spaced document and about 1 to 3 words in the horizontal direction. Such low spatial resolution makes it difficult or impossible to track reading states during word-by-word and line-by-line reading because we cannot locate the words and lines accurately. Typical reading systems often use an unrealistic setting with a very wide line spacing (e.g., triple-spaced), leaving them unsuitable for use with normally spaced text. In addition, as also discussed above, eye-tracking techniques are subject to the inherent transient jitter of human gaze and vertical drift, which require constant calibration. Eye-tracking techniques suited to real-world scenarios have the potential to advance the study of reading. [0044] Furthermore, given the same reading context and motivations, the factors influencing reading states mainly pertain to the reading material’s domain and subject’s domain knowledge about the content. For example, a good reader may cross-reference previously read text to assist in understanding new and unfamiliar text. In such cases, the high reading frequencies of the earlier text do not necessarily imply that they are difficult. To correctly estimate the current reading state, the reading cognition analysis engine 120 may use semantic meaning of the current text, the cross-referenced text, their semantic correlations, and real-time eye gaze patterns. The reading cognition analysis engine 120 may thus be configured to perform the non-trivial task to properly fuse the semantics of reading text and eye movements and learn from them in progressive reading scenarios, and infer semantic explanations for reading state time-series. [0045] The reading cognition analysis engine 120 may comprise a neural network 126 (e.g., a deep neural network) and a multi-task hierarchical classifier 128. The neural network 126 may be trained to predict the reading state based on the two types of sequential modalities: the eye-tracking and the reading text content. For example, the neural network 126 may comprise a four-layer temporal Atty. Docket No.010110-23010A convolutional network (TCN) based module to fuse the two types of sequential modalities. The reading text content modality may be informed by semantic information extracted from a pre-trained NLP model, for example. The multi-task hierarchical classifier 128 may be configured to perform estimation of reading state at two granularities as two distinct but related tasks and utilize a shared convolutional filter mechanism within the TCN to learn the characteristics of the two tasks and their commonalities. In some example, as described in more detail below, a multi-task and hierarchical loss function may be employed to guide reading state estimation. [0046] The reading cognition analysis engine 120 may provide accurate estimations and semantic explanations for reading state time-series to support research and outreach efforts in the field of reading science. In an example, the reading cognition analysis engine 120 is designed to consider several research questions (RS) and posit the corresponding hypotheses. [0047] A first research question (RS1) is: Do readers in the same reading states show different visual attention distributions on the reading text? [0048] A first hypothesis is that readers in the same reading state will show varying visual attention histories, e.g., different total fixation duration, reading times, scanning paths, etc. That is, the visual attention histories of readers in the same reading state differ from each other. [0049] A second research question (RS2) is: When readers are in the same reading states, e.g., encountering difficulty progressing, how does reader visual attention interact with semantic cues in the text? [0050] As indicated by previous studies, readers’ cognitive effort in processing text is positively related to the difficulty of the text. A second hypotheses is that readers can overcome reading difficulties by fetching contextual semantic cues from the surrounding text. When progress is blocked, easy text that is semantically related to difficult text also receives more visual attention and cognitive effort. [0051] Generally, the semantic context of text has a direct impact on the multi-level interactive eye- brain cognitive reading process. Leveraging the rich semantic information about reading materials, which may be extracted by advanced natural language processing (NLP) techniques, the reading cognition analysis engine 120 can improve estimation accuracy and provide semantic interpretation of reading states. The semantic information is generally relatively high-resolution because NLP models can provide Atty. Docket No.010110-23010A semantics at the word level. The inherent hierarchical structure of the semantic information can also be inferred by summarizing the semantics of words to a sentence level. The high-resolution semantic information can compensate for the low-resolution eye movements for more accurate reading state tracking. The real-time interaction of eye movements and semantic context can provide semantic explanations for the ongoing reading states. [0052] In examples, in addition to eye-tracking, the reading cognition analysis engine 120 may consider how individual readers perceive and process the text in real-time. The reading cognition analysis engine 120 may thus enable use of context information obtained from texts in the study of reading cognitive processes, in at least some examples. Eye-tracking technology may acquire real-time eye movements in a non-intrusive manner. It is natural to utilize eye movement data to probe the reading process, as the reading process initiates visual input and operates as an interactive eye-mind cognition process. Eye movement data has been analyzed to understand the reading cognitive process and provide reading assistance. For example, a gaze-aware reading assistance system may provide help at the right time without interrupting the reader’s thoughts. As another example, a social reading system may be used to share eye gaze annotations generated by experts to promote reading comprehension for non-experts. The reading cognition analysis engine 120 may thus detect reading states or behaviors, such as mind wandering or encountering difficulties in comprehending unfamiliar words. [0053] The reading cognition analysis engine 120 may obtain the semantic contextual information from texts using natural language processing (NLP), for example. NLP uses computational techniques to represent and analyze human languages. NLP may be classified into two categories: natural language under- standing and natural language generation. Such natural language understanding techniques may provide generic models for NLP downstream tasks, such as analyzing the association among text components, extracting keywords, and analyzing syntax. For example, given targeted syntax supervision, a long short-term memory (LSTM) network may learn syntax information. Furthermore, the NLP neural network may provide good representations of text; for example, the bidirectional encoder representations from transformers (BERT) model, which is based on transformers, may be utilized to obtain state-of-the- art results on several NLP tasks by providing high-quality language representations. Considering the dependency between the masked positions and the discrepancy from pretrain-finetune that BERT neglects, in some examples, a generalized autoregressive pretraining method may be used to overcome Atty. Docket No.010110-23010A the limitations of BERT. A pre-trained model, such as XLNet, that outperforms BERT on various tasks may be utilized. Pre-trained NLP models may be used by the reading cognition analysis engine 120 to facilitate understanding of the reading cognitive process. [0054] Eye movement patterns can reveal reading strategies and are vital to understanding the reading cognitive process. Reading generally consists of a series of pauses and rapid shifts in gaze locations. The pauses are called fixation and the shifts are called saccades. These patterns reflect the low-level oculomotor characteristics during reading, typically determined by the physical properties of text, such as the positions or lengths of words. By exploring eye movement patterns, connections between low-level eye movement behaviors and higher-level cognitive processes during reading may be established. Generally, the direction and duration of eye fixation reveal how the cognitive process unfolds over time. More specifically, fixation locations indicate the attended content, while fixation duration suggests the level of cognitive effort invested by the reader, i.e., longer fixation suggests more effort. Also, the processing time-course of eye movement patterns may be used to reveal the temporally continuous reading process, which is often linked with comprehending or memorizing. For example, one common temporal reading activity is to move the gaze backward to review the already-read content. In this case, the informative eye movement patterns might be the reading and regression durations, which is also called the second pass. In examples, to alleviate the potential inter-person variations, global features or statistical features based on eye movement patterns may be used to access the reading process, such as the number of saccades, saccade frequencies, and variations in fixation duration. Given the potential ability of eye movement patterns in revealing reading cognitive processes, the reading cognition analysis engine 120 may employ such hand-engineered features as valuable indicators. Further, as described in more detail below, the reading cognition analysis engine 120 may distinguish the representing eye movement patterns at multiple granularities, such as at word and sentence levels, for example. [0055] There exists a strong correlation between eye movement patterns and attentional processing during reading. For instance, attention during reading generally moves from word to word continuously, and visual attention is generally linked to changes of focus in text processing. In aspects, the reading cognition analysis engine 120 tracks visual attention during reading by establishing the connection between eye movement patterns and the corresponding while-reading text components, such as words and sentences. In an example, the visual attention state may be defined to be the collection of eye Atty. Docket No.010110-23010A movement features on each text component. For example, when reading the sentence “They race to maturity, with the shortest generation time of any vertebrate”, the visual attention for the word “vertebrate” consists of fixation duration, reading times, number of fixations, etc. At the sentence level, the visual attention state is defined based on the total dwell time, saccade times, etc. [0056] The reading cognition analysis engine 120 may be configured to facilitate exploration of how the semantic meaning from text assists in estimating the time-series reading states and how they explain these states. From this perspective, the reading cognition analysis engine 120 may be configured to utilize a holistic semantic understanding of while-reading texts. The reading cognition analysis engine 120 may be configured to obtain such understanding should cover the semantic meaning of different grain sizes of texts, ranging from single words and sentences to passage levels. The semantics collection at various granularities is sometimes referred to herein as “semantic attention”. Semantic attention can hint at whether the while-reading text components are difficult. These difficult components may be unfamiliar or ambiguous words or sentences with complex syntax, which often delay reading. In this case, appropriately using such semantic meaning regarding the difficulty score can provide additional evidence in revealing the current reading state and deliver a reasonable interpretation regarding why the current text components block the reading. [0057] In some examples, the reading cognition analysis engine 120 may include or be integrated with an interactive assistance system that may provide interventions based on the detected reading state of the user. The interventions may be provided, for example, in-time and as-needed during the user’s reading of the text. For example, text components with which the user is having difficulty may be highlighted on a screen from which the user is reading the text, and the corresponding treatments may be shown on the screen (e.g., at the right top of the text content) in a pop-up window. Additionally or alternatively, the reading cognition analysis engine 120 may provide an after-reading summary that may include reading states and corresponding treatments for display to the user after the user has finished reading the text. [0058] Turning now to Fig.2, a diagram depicting an example implementation of a system 200 configured to determine reading states according to one example is provided. The system 200 generally corresponds to, or is utilized with, the system 100 of Fig.1 in some examples. For example, the system 200 (sometimes referred to herein as “CASES-Net”) corresponds to, or is included in, the reading Atty. Docket No.010110-23010A cognition analysis engine 120 of Fig.1. [0059] The system 200 includes a semantic attention extraction (SAE) engine, a visual attention extraction (VAE) engine, a cross-attention extraction (CAE) engine, and a reading state estimation/explanation engine, in the illustrated example. [0060] In an aspect, the system 200 operates to obtain a comprehensive semantic understanding of the text before the reading begins. This semantic meaning information may be used to compensate for the low-resolution eye-tracking data, thus enabling accurate reading state estimation. Semantic meaning also enables explanations during reading state detection tasks in later pipeline stages. To extract semantic meaning, the system 200 may turn on the outward-facing scene camera (e.g., the sensor 104) to obtain the text to be read. The SAE module may then run at least once on the text. The SAE module may utilize NLP techniques to extract the high-resolution semantic features and the inherent linguistic structure from the text, thus facilitating subsequent tasks. [0061] Texts contain rich semantic information, but for better individual reading state estimation, personalized visual attention data may also be necessary. In an aspect, to capture visual attention data, the VAE module may be triggered to obtain the online visual attention features corresponding with text components (e.g., while-gazing words or sentences). For example, the system 200 may sense (e.g., using the sensor 102) reader eye images to predict gaze sequences using continuous eye-tracking. Then, the VAE module may extract visual attention features from the sequential gaze data. In parallel, the scene camera (e.g., the sensor 104) may record time-aligned scene images to facilitate tracking of gaze positions. [0062] Because the obtained semantic meaning of the text and visual attention features are at different spatial resolutions, the CAE module may be configured to properly align them. For example, minimal units, such as words upon which sentences and global context depend, may be used to segment the visual attention features. The system 200 may be equipped with a TCN-based network configured to estimate the reading states at word and sentence levels, aiming to explore the task-specific features for the assistance of the multi-task output. A first feature, sometimes referred to herein as a “word-level task,” may represent the binary determination of whether a reader has difficulty processing a word. The second task may comprise hierarchical multi-label classification at the sentence level, which may include a first task (Task I) estimating whether a reader is having sentence-level processing difficulty and if so, a Atty. Docket No.010110-23010A second task (Task II) estimating whether the reader is facing comprehension challenges, the reader’s mind is wandering, or both. In an aspect, a multi-task and hierarchical loss function for training guides CASES-Net. Reasons for the predicted reading states may be qualitatively determined by visualizing the learned semantic attention and visual attention features. [0063] The SAE engine may be configured to understand the high-resolution semantic meaning of the document ^, ranging from the word level to the document level. For example, at least three types of semantic features may be extracted by utilizing various advanced pre-trained NLP models: encoded word features, keyword features, and word difficulty features. [0064] Each word in ^ may be encoded as a 768-dimensional vector by a suitable model, such as XLNet model. The semantic meaning of the document may thus be obtained by processing the whole text passage once. To lower the potential adverse effect incurred by the high dimensionality, the obtained features (e.g., the XLNet features) may be reduced to 64 dimensions via a fully-connected (FC) layer. The dimensionally reduced encoded features may be denoted as ^^ = {^^ ^ }^ ^^^ , where ^^ ^ ∈ ^^^ and ^ is the total number of words. [0065] To understand the keyword information in the document, the probability of each word describing the whole document may be calculated using a suitable model, such as the YAKE model. The keyword features may be denoted [0066] Word difficulty may be used to assist in the task of identifying the reading state. In an aspect, the word difficulty may be described using the length of the word, number of syllables, and familiarity scored by the MRC psycholinguistic database. The difficulty of words may be denoted by ^^ = where ^^ ^ = ^^^ , ^^ , ^^^ ∈ ^^. [0067] Each word in the document may represented by the concatenation of the three feature vectors; that is ^^ = ^^^ ^ , ^^ ^ , ^^ ^^ ∈ ^^^ (^ = 1,2, … , ^). The semantics regarding more coarse levels (e.g., sentence- and passage- level) may be generalized from that of the word level, as words are inherently structured and semantically connected — a passage consists of multiple sentences and a sentence of multiple words. [0068] A reliable gaze sequence is the foundation for accurate visual attention feature extraction. However, the raw gaze points are noisy due to difficult-to-avoid human motion and limited eye-tracking Atty. Docket No.010110-23010A resolution. In an aspect, to alleviate this issue, a filtering algorithm may be utilized to smooth the raw gaze points, leveraging their sequential characteristics. In an aspect, eye-tracking may first be used to estimate the points of gazes (PoGs) and record the initial PoGs sequencies as ! = where & is the total number of timestamps considered, and the initial PoGs sequencies may then be filtered using a designed filtering method. The designed filtering method may, for example, first apply median filtering to discard outliers due to gaze jitter. Then, mean filtering may be applied to stabilize the fluctuations of sequential PoGs due to the limited eye-tracking resolution. Filtered or smoothed PoGs ! = {" #}% $^^ may thus be obtained. Each word and sentence may be segmented using !and then sent to the next step for visual attention extraction. [0069] Generally, the number of PoGs increases rapidly during reading. To reduce the size of PoGs, the system 200 may use a number of representative features reflecting how people comprehend characters during reading or whether they are disengaged from reading. For example, the following features may be used to describe word-level processing state while reading: fixation duration, number of fixations, and number of repeated word readings. A variation in these features from person-to-person may exist. The system 200 may be configured to normalize personal data to reduce or eliminate person- to-person variation. It is noted, however, that these three features may vary not only person-to-person but also for a particular person during reading. Such variation may significantly affect estimation performance. In an aspect, the system 200 may be configured to use local information to reduce or eliminate the during reading variation. For example, every ( seconds, the system 200 may add statistical features, including mean and variations, for ! = {" #}) $^^ to describe the visual attention for each word. In total, a 9-dimensional feature may be obtained for each word. Further, the system 200 may normalize obtained visual features, such as 4 sentence-level representative visual features, including dwell time, saccade times, forward saccade times, and backward saccade times, using the sentence length, so these features better describe the local variation. With * words being segmented during (, each word may be represented using = 1,2, … , *). There may be 9 word-level features and 4 sentence-level features that are identical to the words in the same sentence. [0070] In an aspect, the deep features of the sequential gaze data may be used for eye movement classification. For example, deep neural networks (DNN) that have been shown to be effective on eye movement classification may be used. In an example, feature extractor based on the 1D-CNN with Atty. Docket No.010110-23010A BLSTM backbone (denoted may be adapted to extract 8-dimensional deep features during time duration ^, i.e., ^^ ^ 22 ∈ ^^ (^ = 1,2, … , *), with the classifier discarded. [0071] To facilitate downstream multi-task learning, the cross-attention extraction (CAE) engine may fuse the two modalities to predict the reading states at different granularities and the distinct task-specific information. [0072] Before fusing the two modalities, the system 200 may perform time synchronization of the two modalities. For example, for each smoothed PoGs sequence " # , the * words being processed at time 3 may be identifies. The two feature vectors may then be concatenated to obtain 4^ # = ∈ as the overall representation of the two modalities. For all other words ^6 that have been visually processed till the time 3, the semantic attention feature vector may be padded, e.g., ^^ 7 with a zero vector, i.e., 4^7 # = In this way, the word being processed at time 3 may be properly described semantically with its corresponding visual attention features. In contrast, the unread words padded with zeros are given less attention. [0073] The CAE engine may use a Temporal Convolutional Network (TCN) model, which can capture temporal dependencies. [0074] For example, the CAE engine may use temporal convolutional filters/kernels to process input sequences. Each filter may calculate a weighted average in the time domain. The parameters of the filters may be learned to optimize the objective function. In an aspect, each TCN layer consists of temporal convolutions, a non-linear ReLU activation function, and a max pooling function or an upsampling function. [0075] The CAE engine may comprise four TCN layers. To learn different tasks more efficiently, the filters of the last layer may be divided into task-specific filters, namely word-level filters/sentence-level filters, and task-shared filters, namely common filters. The features extracted by word-level filters and common filters may be used for word-level tasks. The features extracted by sentence-level filters and common filters may be used for sentence-level tasks. [0076] Based on the obtained cross-attention features, the system 200 may detect the reading state of “processing difficulty”. Detection may comprise the following three tasks. (1) Word-level binary-class Atty. Docket No.010110-23010A classification task &^<=> : The word-level features may be fed to a fully connected layer to predict whether the reader finds the word being processed difficult. Sentence-level and word-level tasks differ. Because mind wandering may co-occur with reading difficulty for a sentence the task at the sentence level may be formulated in the following hierarchical fashion. (2) Sentence-level binary-class classification task &?@A$,^ : With the sentence-level features, the system 200 may first determine whether the reader is in a normal reading state without any processing difficulties using a binary classifier. (3) Sentence-level multi-label classification task &?@A$,B : If the reader enters into an abnormal state, the reader can be either mind wandering or processing difficulty, or both; This is a multi-label classification task, where multi labels can be assigned simultaneously; label 1 is mind wandering and label 2 is processing difficulty. [0077] To train the network, the following loss function reflecting the performances of all tasks may be utilized: Equation 1 [0078] Binary Cross Entropy (BCE) loss may be used for &^<=> . ℒD&^<=> E may be obtained as follows: Equation 2 where ^ denotes the number of words; I^ ^<=> denotes the label of word ^, I^ ^<=> = 0 indicates the reader finds the word ^ easy, I^ ^<=> = 1 indicates the reader finds the word ^ difficult; M^ ^<=> is the word- level estimation result given by the network O^<=> . [0079] BCE loss may also be used for &?@A$,^ . ℒ+&?@A$,^. may be obtained as follows: ?@A$,^ ?@A$,^ ?@A log M? + D1 − I? E logD1 − M? $,^E Equation 3 where Q denotes the number of sentences, I? ?@A$,^ denotes the binary classification label of the ^th sentence, I? ?@A$,^ = 0 indicates the reader is in a normal reading state without any processing difficulties for sentence ^, I? ?@A$,^ = 1 indicates is in an abnormal reading state, M? ?@A$,^ is the sentence-level binary classification estimation results given by the network O?@A$,^. [0080] For sentences with I? ?@A$,^ = 1 to solve the multi-label problem, BCE loss may be used for each Atty. Docket No.010110-23010A label separately. The loss of &?@A$,B may be obtained as follows: ?@A$ logD1 − M?,U ,BEE Equation 4 where W = 2 denotes the number of labels, i.e. label 1 as mind wandering and label 2 as processing difficulty; I? ? ,T @A$,^ denotes the supervised information of the ^th label for sentence ^, I? ? ,@ T A$,^ = 1 indicates sentence ^ has the ^th label, I? ? ,@ T A$,^ = 0 indicates sentence ^ does not have the ^th label; 1+⋅. is an indicator function, ?@A$,^ ?@A$,^ ?@A$,B = 1E = 1 when I? = 1, = 1E = 0 when I? = 0; M?,T is the sentence-level multi-label estimation results given by the network O?@A$,B. [0081] In an aspect, a real-time reading state detection and intervention system (sometime referred to herein EYEReader) is provided. The real-time reading state detection and intervention system may be configured to (e.g., using the CASES system) determine reading state series that influence reading fluency and mitigate the negative effects of reading processing difficulties. For the convenience of readers, EYEReader may be implemented in the form of a website plugin browser, enabling cross- platform compatibility. In various examples, such reading intervention or assistance system may contribute to educational applications, HCI studies, etc. [0082] In an aspect, the text materials to be used with the EYEReader, as selected to include various topics to facilitate text-agnostic intervention. For example, a number (e.g., 36) reading comprehension materials with diverse topics may be selected from an English qualification test to match users’ reading comprehension ability. Each article may have around 450 words on average. Users may log in to the system, select their preferred articles from existing materials, and start reading by simply clicking a button. [0083] Because the detection system used (e.g., the system 200) may not have limited resolution issues when eye-tracking is used during reading scenarios, the interface of text presentation of EYEReader may be similar to common computerized reading settings. In an example, articles may be divided divided into several different pages (e.g., around 240 words per page) with a regular line height, Atty. Docket No.010110-23010A approximately single-spaced. An 18-point default font typeface may be utilized. [0084] During the reading process, readers wear the prototype eyeglass and sit in front of the computer to read. The pre-trained CASES-Net model may be always-on to detect potential abnormal reading states, i.e., whether the user is struggling with difficult words or complex sentences, or their mind is wandering. When abnormal events that affect reading are detected, the system may trigger interventions. The text components may be highlighted, and the corresponding treatments may be shown on the screen (e.g., at the right top of the text content) in a pop-up window. This way, the user-system interaction cost may be minimal. [0085] Figs.3A-C illustrate screenshots of examples of the system’s user intervention assistance that may be provided by the system 100 of Fig 1. In particular, Figs.3A-C illustrates screenshots of three intervention examples. Detection and Interventions: at the word level (Fig.3A), at the sentence level (Fig. 3B), and on mind wandering (Fig.3C). As illustrated in Fig.3A, word-level intervention may be provided by highlighting (e.g., boldfacing or otherwise highlighting) a word with which the user is having difficulty and displaying a translation, a definition, or other explanation of the word in a box 302 on the screen on which the user is reading the text. Similarly, as illustrated in Fig.3B, sentence-level intervention may be provided by highlighting (e.g., boldfacing or otherwise highlighting) a sentence with which the user is having difficulty and displaying a translation or other explanation of the sentence in a box 352 on the screen on which the user is reading the text. As illustrated in Fig.3C, a mind wandering intervention may be provided by highlighting (e.g., boldfacing or otherwise highlighting) a sentence on which the user’s mind may be wandering and displaying an alert indicating the mind wandering state to the user in a box 372 on the screen on which the user is reading the text. In some examples, interventions in Figs.3A-C are provided in a native language of a user which may be different from the language of the text. For example, the interventions include translations of words or segments (e.g., phrases, sentences, etc.) into the native language of the user and/or text explanations or other interventions in the native language of the user. In other examples, interventions may be provided in a same language as the language of the text. [0086] Fig.4 illustrates example architecture of EYEReader. In an example, Vue.js framework may be used to develop the front-end website. A Python framework (e.g., Django) may be used for the back-end of the website. Django offers a variety of third-party tools for building communication between the front- Atty. Docket No.010110-23010A end and back-end efficiently following the REST API specification. To store and manage the data on the server, open-source database management system (e.g., MySQL) may be adapted. In addition, Pupil Capture and Pupil Service may be used to handle the real-time eye image data collected, which is based on the IPC Backbone provided by Pupil Labs. [0087] In an example, the overall operation workflow of the intervention system that provides just-in- time interventions for users encountering reading processing difficulties may include six steps as follows. [0088] Step 1: During system operation, EYEReader loads the pre-trained CASES-Net from the server when receiving the requests from the front end. [0089] Step 2: The recorded eye/scene images are used for eye-tracking using the Pupil service. [0090] Step 3: The tracked gaze points are sent to the server for further visual attention feature extraction. [0091] Step 4: The server loads the historical eye-tracking data, visual attention features, and texts to decide whether it is the right time to intervene. [0092] Step 5: Once processing difficulties are detected, the estimation results are returned to the front end for triggering interventions. The corresponding treatment is shown at the front end to facilitate the current reading. [0093] Step 6: After that, the current interventions and all other data are saved in Database. [0094] In an example, CASES-Net may be integrated into eyewear, as eyewear is a natural way to be used in various reading scenarios. The eyewear may be suitably designed for use with the EYEReader system. Fig.5 illustrates example eyewear hardware in accordance with an aspect. Eyewear may be well-suited for use in various reading scenarios. To facilitate adaptation of eyewear, a stand-alone scheme may be sued to integrate the computing components and power supply into the headset frame. [0095] The eye-tracker may follow the Pupil 1 with several adjustments. In an example, Qualcomm Snapdragon 865 platform may be directly integrated into the left leg of the eyewear. The eye camera and scene camera modules may be replaced with 20 MegaPixels (MP) Samsung S5K3T2 and 64 MP Samsung S5KGW1, respectively. The eye camera may be used to record eye videos to perform eye tracking. The scene camera may sense scene videos to capture the text being read. The 3D eyeglass Atty. Docket No.010110-23010A frame may be designed to fit the two cameras into the left leg of the mounting frame. To balance the weight of the headset, the battery may be integrated into the right leg of the eyewear. [0096] Experiments to evaluate CASES, the cognition-aware eyewear system for estimating reading states, have been conducted as described below. We first detail the experimental setup, data collection, and evaluation measures. We present results and quantify the technical capabilities of CASES. [0097] Experimental setup included recruiting 20 participants by posting a questionnaire at our university campus. A summary of the participant demographics follows. Age: 22–28 years old with an average age of 23.7. Gender ratio: 18 males (90.0%) and 2 females (10.0%). [0098] Figs.6A-C illustrate a data collection method in accordance with an example. As shown in Fig. 6A, the participant wears eyeglasses and sits in front of the computer to read. While reading, videos were recorded using the eye camera and time-aligned videos using the scene camera. [0099] Texts covered a wide range of subjects so readers can enter multiple reading states. Moreover, each text was relatively short, allowing participants to read several texts.36 articles were selected with the following three subjects: [00100] Subject matter 1 - One-minute BBC world news: 10 articles with approximately 300 words per article on average. [00101] Subject matter 2 - English qualification tests: 16 articles containing reading comprehension materials with approximately 450 words per article on average. [00102] Subject matter 3 - Philosophy related: 10 articles with approximately 500 words per article on average. [00103] The first two of these provide challenging words and sentences. The third provides mundane subject that may lead to mind wandering. It was anticipated that most participants are unfamiliar with the third subject matter, and it is hard to understand the content without prior knowledge. [00104] The CASES requires time-aligned eye gaze data and text data (i.e., the words or sentences being read) to detect reading states. In addition, the synchronized data should capture continuous reading, during which users may encounter various reading states. In an aspect, an online system to collect data meeting the requirements is provided. Atty. Docket No.010110-23010A [00105] First, each article may be divided into pages. There may be around 240 words per page in single-spaced 18-point typeface. Then, articles may be randomly selected from each topic for the participants to ensure that they cover all three subject matters. This design allows most participants to encounter numerous reading states. Each article may be read by at least two participants. Participants may be instructed (e.g., verbally) on how to use the data-collection system, such as navigating to the next/previous page. Each participant may then read the texts. Fig.6B shows a screenshot from an example page. Reading one article may take approximately six minutes. [00106] After completing an article, the participant may be immediately instructed to label their reading states. A labeling tool with a GUI window may be used to facilitate labeling. Participants may review each page of the article. On each page, the participants may use a single click to label the words they cannot comprehend and use a double-click to label sentences they do not comprehend. A button may be provided for each sentence (e.g., at the right top of each sentence) for users to mark whether their minds wandered when reading it. The annotated words and sentences may be boldfaced as shown in Fig.6C and/or may be highlighted in different colors so users can quickly double-check their annotations. Annotating one article may take around three minutes. In total, data collection (including annotation collection) may take approximately seven days. [00107] The collected dataset may be randomly split into training (80%) and test (20%) sets per participant/article. The total numbers of labels for “word-level processing difficulties”/“sentence-level processing difficulties”/“mind wandering” may be 841/177/186. [00108] Because the disclosed system is hierarchical and multi-task, appropriate measures to evaluate each task may be adopted. The first task is binary classification of whether a reader is facing difficulty processing a word. Performance of this task may be evaluated using accuracy and the receiver operating characteristics (ROC) curve. The second task is hierarchical multi-label classification at the sentence level, which may include sentence-level Task I and Task II described above. Task II is multi-labeled. Accordingly, the multilabel-based macro-averaging metric may be used, e.g., averaged-accuracy and ROC curve, to evaluate it. [00109] The CASES system has been evaluated in real-world contexts. A study involving 20 participants has been conducted. In the study, the CASES system showed superior reading state detection to baseline methods. In aspects, encoding text semantic content facilitates learning from Atty. Docket No.010110-23010A context cues and improves reading state estimation accuracy. Compared with the conventional eye- tracking-only method, the disclosed reading state detection system improves accuracy by at least 19.70%. In at least some examples, the text semantic context enables quantitative explanations of reading (cognitive) states. [00110] Ablation studies were conducted to evaluate CASES, as there is no prior work solving the problem addressed in this work, thus making direct comparisons with prior work infeasible. The following three baseline methods were used for evaluation. [00111] Visual: Previous studies have demonstrated that some reading states, such as mind wandering, can be identified using gaze-relevant features, which are closely related to our work. To validate whether the eye-relevant features are sufficient for reading state recognition at multiple text element granularities (words and sentences), a baseline method leveraging 13 eye-relevant features (9 word-level features and 4 sentence-level features described above) was used to identify the state while reading. The support vector machine (SVM) method was used to conduct the three classification tasks: word-level task, sentence-level Task I, and sentence-level Task II. SVM was used because it has been successfully applied to various classification tasks, and is one of the widely used methods in similar tasks. For simplicity, this method is sometimes referred to herein as “Visual”. [00112] Visual+: Eye movement patterns are good indicators for reading state recognition. The CASES system may leverage deep neural network (DNN) to achieve accurate eye movement pattern identification. The deep features (e.g., 8-dimensional deep features) extracted from a deep neural network (e.g., 1D-CNN with BLSTM) to improve the accuracy of reading state estimation. To make a fair comparison, the extracted deep features were concatenated with the 13 expert-designed features and sent to the CAE module to estimate reading state. This baseline method is an improved version of the Visual method sometimes referred to herein as” Visual+”. [00113] NLP: Visual and Visual+ identify reading states based solely on visual attention features. To verify the classification performance based on the semantic content of texts, an NLP method is used. As in the Visual+ method, semantic features are first extracted using the SAE module and then the extracted features are sent to the CAE module to infer reading states. [00114] Fig.7 illustrates the reading state recognition performance of the disclosed method and three Atty. Docket No.010110-23010A baseline methods. The disclosed method achieves the best performance. Compared with the Visual method, i.e., conventional eye-tracking only, CASES improves the accuracy by 28.13%, 58.70%, 19.70% for the word-level task and the sentence-level Task I and Task II. Furthermore, compared with the baseline method Visual+ and NLP, CASES has superior reading state estimation. For example, the word- level detection accuracy of CASES is 82.11% while it is 78.80% or lower for the baseline methods. Accordingly, it can be seen that using context derived from text improves reading state estimation. [00115] The Receiver Operating Characteristic (ROC) of different methods. Figs.8A-C illustrate plots of Receiver Operating Characteristic (ROC) of different methods. It can be seen that CASES outperforms the baseline methods in Area Under the Curve (AUC), which is one of the most widely used performance measures in classification or retrieval problems. [00116] Accordingly, as further explained in more detail below, CASES outperforms the baseline methods and offers semantic explanations of the predicted reading states, in at least some examples. [00117] To study progression through cognitive states while reading to assist understanding of the reading process, conducted in-field pilot studies using CASES have been conducted. The findings around the designed two RS and hypotheses using CASES are described in more detail below. The capability of EYEReader is demonstrated to make helpful real-time interventions when reading difficulties are encountered. [00118] Ten volunteers were recruited to participate in the pilot study. The average age was 24.0 years (SD=1.6, min=22, max=28), with n=1 (10%) female and n=9 (90%) males. All participants were English as Second Language readers with self-reported normal or corrected-to-normal vision. All participants also reported that they passed a standard English Qualification Test. [00119] During the pilot study, participants were encouraged to use the system whenever they read. The whole pilot study lasted several months and consisted of two stages. During the first stage, participants were required to label the words and sentences they encounter difficulty processing. These labels were treated as ground truth. Based on the qualitative evaluations made by the participants, several findings on how people read at different granularities, i.e., single words and sentences, were made and six patterns to discuss were summarized. The second stage focused on applying EYEReader in practice. At the end of the pilot study, each participant completed a survey of their opinions on the Atty. Docket No.010110-23010A usability and value of EYEReader. [00120] The following three observations on how users read at the single-word level were made. [00121] Observation I: Users comprehend the lexical meanings of words by directing their gazes more frequently toward material they find difficult to process. When users encounter difficulty processing a word, they usually gaze at it longer, and more times than typical. Fig.9 illustrates one example of this observation, where participant P6 has difficulty comprehending the meaning of “mitigate” and “debris”. Text read by participant P6 is shown in the top box of Fig.9 for legibility. Circles over the text in the bottom box of Fig.9 represent gaze points. Shading in the bottom box of Fig.9 corresponds to text fixations as the reader is reading the text, with lighter shading corresponding to longer fixations of the user. P6 fixates “mitigate” (fixation label 13) and “debris” (with fixation label 19) for a long time and reads them more than two times. In particular, P6 has the longest fixation duration on the word “debris” and he has the most reading times on the word “debris” and “mitigate”. [00122] Observation II: When a user encounters difficulty processing a word, the user first directs their gaze to the word and then to other words to examine the semantic context. Readers generally avoid breaking their chain of thinking by stopping when a difficult word is encountered, especially when the word does not affect their understanding of the text. However, when readers consider a difficult word to be highly topic-relevant or meaningful for subsequent text comprehension, they tend to interrupt their reading and attempt to deduce the semantic meaning of the word from its semantic context. This observation differs from previous studies and the next observation complements it. [00123] Observation III: When users examine the semantic context of a difficult-to-process word, they gather semantic clues by shifting their gazes to different locations even when considering the same difficult word, from the same text, under similar reading conditions. Readers typically attempt to find an appropriate location in the text to help comprehend the current difficult-to-process word. The text at the location should reveal the relevant information about the difficult word. Also, that location varies from person to person, depending on their current cognitive states about the context. [00124] Fig.10 illustrates the proportion of the three above observations for each participant by summarizing their past experienced processing difficult words. All ten participants experience Observation I in most cases (around 87.47% cases on average). The ten users fall into Observation II & Atty. Docket No.010110-23010A Observation III in fewer times, i.e., around 12.53% on average. [00125] Fig.11 illustrates an exemplary case to provide further insights on Observation II and Observation III. Text read by participants P2 and P5 is shown in the top box of Fig.11 for legibility. Circles over the text in the bottom box of Fig.11 represent gaze points. Shading in the bottom box of Fig.11 corresponds to text fixations as the reader is reading the text, with lighter shading corresponding to longer fixations of the user. Here two readers, P2 and P5, face the same reading difficulty in comprehending the word “liberation” when they read the same sentence from the same article. In can be seen that the two participants first direct their visual attention to the target word, “liberation” where the fixation labels are 14 and 13 for P4 and P5, respectively. They then shift their gazes. Participant P4 gazes back at the previously read word, “pleasant”, while Participant P5 gazes forward to the word, “promised”. Both of these words are semantically relevant to the difficult word, “liberation”, as shown in Fig.11 (top row). [00126] Observations at Sentence Level. This section focuses on two modes of comprehending sentences: interpretive (semantic) and structural (syntactic). [00127] Observation IV: People incrementally comprehend the semantics of a sentence as they read each word, while with different gaze time series. Fig.12 illustrates the inter-reader differences in gaze time series for the same sentence. Text read by participants P1 and P4 is shown in the top box of Fig.12 for legibility. Circles over the text in the bottom box of Fig.12 represent gaze points. As can be seen in Fig.12, P1 focuses on the first parts of sentences (with more fixation, labels 0–12) while P4 focuses on other parts of sentences (fixation labels 9–11). [00128] Observation V: Readers enter the “rereading” or “reanalysis” state at different times when having difficulty with the same sentence, as illustrated in Fig.13. Text read by participants P8 and P3 is shown in the top box of Fig.13 for legibility. Circles over the text in the bottom box of Fig.13 represent gaze points. As can be seen in Fig.13, P8 backtracks 3–4 words (with a fixation label starting from 28) when reading the middle of the sentence, and then continues reading the sentence; while P3 rereads the sentence from the beginning when reading the middle of the sentence (fixation label 12). [00129] Observation VI: Different people “reread” the same sentence with different reading states. Fig. 14 shows two participants reading the same sentence twice. Text read by participants P1 and P4 is shown in the top box of Fig.14 for legibility. Circles over the text in the bottom box of Fig.14 represent Atty. Docket No.010110-23010A gaze points. As can be seen in Fig.14, P1 gets distracted (i.e., enters the mind wandering state) during the first reading of the sentence (typical fixation labels 2, 6, and 14); therefore, P1 spends more time and has more fixations on the sentence in the second reading (fixation labels 21, 24, 26, 28) than in the first pass. In contrast, P4 spends more time when reading the sentence the first time (fixation labels 3, 17, and 18), but he quickly skims it the second time (fixation labels 26 and 37). [00130] In various examples, CASES-Net accurately detects reading states. In various examples, EYEReader may promote reading comprehension by detecting reading states implying processing difficulties and making real-time interventions. [00131] To make quantitative assessment of EYEReader, reading comprehension improvement as may be defined as (^MY^3−^M^Z^Z[3 )/^MY^3 , where ^M^Z^Z[3 and ^MY^3 denote the number of challenging words or sentences at present and in the past, respectively. The higher the (^MY^3−^M^Z^Z[3 )/^MY^3 , the higher the reading comprehension improvement. This definition is used to identify challenging words and sentences. After pilot studies, participants were asked to indicate whether they still face challenges in comprehending these words and sentences. Fig.15 illustrates the results. All ten participants have positive reading gains, which means that, in at least some examples, EYEReader is effective in helping users to overcome unfamiliar words and complex sentences. [00132] Several open-ended questionnaires to qualitatively evaluate EYEReader were designed. Ten questionnaires were sent to participants, eight of which were returned. Among them, 7/8 of the participants positively commented on word-level intervention. They believe that fine-grained intervention at the word level can precisely pinpoint the reading difficulties they are experiencing. There are 6/8 of the participants found sentence-level intervention helpful. In particular, when facing challenging sentences with complex syntactic structures, it was difficult to comprehend the sentence even though they were familiar with all the words. In this case, EYEReader helped them overcome this reading difficulty by highlighting the sentence and explaining it. In addition, 6/8 of the participants found EYEReader valuable in reminding them when their minds wandered; these participants stated that they usually do not realize when they are distracted. Timely reminders can make their reading more focused and efficient. [00133] CASES may accurately estimate and provide semantic explanations of reading states over time, which can facilitate the scientific study of reading by enabling a deeper understanding of the cognitive processes involved in learning to read, disentangling the complex combination of cognitive skills Atty. Docket No.010110-23010A and their impact on reading fluency, and measuring the efficacy of methods for teaching reading and beneficial reading habits, in at least some examples. [00134] The disclosed systems and methods may facilitate investigation of the human cognitive reading process by exploring the complementarity of eye movements and text. In some examples, the disclosed systems and methods may further use illustration information to understand how people read. Text-diagram instructions may be used to improve reading comprehension. In examples, the disclosed systems and methods may integrate semantic information, including text and illustrations, with eye movements for more accurate reading state detection. Prediction of various reading states may be performed to provide a complete picture of the reading cognitive progress. In addition to, or instead of, determining the reading states at the word and sentence levels, the disclosed systems and methods may be used to measure how people read at other granularities, such as the entire passage level. This may deepen the understanding of how people summarize and reflect on learned knowledge during reading, in at least some examples. [00135] As described herein, a cognition-aware smart eyewear system or CASES may recognize (cognitive) reading state time-series using eye tracking based visual attention and text semantic context. Ablation studies demonstrate that CASES significantly improves the accuracy of reading state recognition over the conventional approach using only eye tracking. Furthermore, in-field studies enabled several observations about how individual reading state time-series are related to text semantic context at different granularities. In at least some examples, the ability to track semantic context cues enables better understanding of progressive reading states. In aspects, an interactive reading assistant system may be equipped with CASES and may provide just-in-time interventions when reading difficulty is encountered. The interactive reading assistant system may promote self-awareness of cognitive processes while reading and facilitate improvement of reading habits. In various examples, CASES may be used in the scientific study of reading, cognition, and human-computer interfaces. [00136] Fig.16 depicts a method 1600 for detecting reading states, in accordance with one example. The method 1600 may be implemented by one or more of the processors described herein. For instance, the method 1600 may be implemented by an image signal processor implementing an image processing system, such as the system 100 of Fig.1 or the system 200 of Fig.2. Additional and/or alternative processors may be used. For instance, one or more acts of the method 1600 may be implemented by an Atty. Docket No.010110-23010A application processor, such as a processor configured to execute a computer vision task. [00137] The method 1600 includes an act 1602 in which one or more procedures may be implemented to obtain a first data stream. The first data stream may be indicative of one or both of i) eye movements and facial expression of the user as the user is reading a text. The first data stream may include, for example, one or more images or video frames depicting an eye region of a face of the user, eye movement of one or both eyes of the user, eye gaze of the user, etc. as the user is reading the text. The first image steam may be obtained from a first sensor, such as an inward-facing camera that may be attached to a smart eyewear frame worn by the user, or other suitable sensor. In an example, the first data stream is obtained from the first sensor 102 of Fig.1. [00138] In an act 1604, one or more procedures may be implemented to determine, based on the first data stream, visual information indicative of one or both i) eye movements and ii) facial expressions of the user during reading of the text. The visual information indicative of eye movements of the user may be indicative of, and/or may be used to determine, gaze direction of the user as the user is reading the text. Act 1604 may include an act 1606 in which one or more procedures may be implemented to determine a gaze point sequence based on the first data stream. The gaze point sequence, or points of gaze (PoGs), may be indicative of a direction of gaze of the user as the user is reading the text. In some aspects, the determined PoGs may be filtered to generate smoothed PoGs. Filtering may include median filtering to discard outliers due to gaze jitter and/or mean filtering to stabilize the fluctuations of sequential PoGs due to the limited eye-tracking resolution, for example. In an example, determining the gaze point sequence at act 1604 may include determining, based on the first data stream, an initial gaze point sequence, and performing one or both of i) median filtering and ii) mean filtering of the initial gaze point sequence to generate a smoothed gaze point sequence. [00139] Act 1604 may include an act 1608 in which one or more procedures may be implemented to extract, from the gaze point sequence (or the smoothed gaze point sequence), visual attention features indicative of one or both of i) word-level processing state of the user and ii) segment-level processing state of the user. Determining the visual attention features indicative of segment level processing state of the user may include determining visual attention features indicative of one or more of i) phrase-level processing state of the user, ii) sentence-level processing state of the user, ii) paragraph-level processing state of the user, iv) page-level processing state of the user, or v) document-level processing state of the Atty. Docket No.010110-23010A user. [00140] Determining the visual attention features indicative of the word-level processing state of the user may include, determining, for each word among at least some words of the text, one or more of i) fixation duration on the word, ii) number of fixations on the word, or iii) number of repeated readings of the word. Determining the visual attention features indicative of the segment-level processing state of the user includes determining, for each of at least some text segments of the text being read by the user, i) dwell time of the user on the text segment, ii) forward saccade time exhibited by the user as the user is reading the text segment and iii) backward saccade time exhibited by the user as the user is reading the text segment. Determining the visual features may also include using a deep neural network (DNN) to extract deep features from the gaze point sequence. [00141] In some aspects, act 1604 may also include an act 1610 in which one or more procedures may be implemented to normalize of the visual attention features based at least in part on statistical data determined from the visual attention information obtained while the user is reading the text. For example, word-level visual attention feature normalization may be performed based on statistical history of the user’s reading to reduce or eliminate during-reading variation in eye movements exhibited by the user. As another example, segment-level visual attention feature normalization may be performed based on a length of the segment such that the segment-level features better describe the local variation. [00142] In some examples, determining the visual information in the act 1604 may also include determining facial features indicative of emotional state of the user as the user is reading the text, such as whether the user is calm, engaged, interested, disinterested, confused, frustrated, etc. [00143] In an act 1612, one or more procedures may be implemented to obtain a second data stream. The second data stream may be indicative of the content of the text. The second data stream may include one or more images or video frames capturing the text prior to the user’s reading of the text and/or as the text is being read by the user, for example. The second data steam may be obtained from a second sensor, such as a forward-facing camera that may be attached to the smart eyewear frame worn by the user. In an example, the second data stream may be obtained from the second sensor 104 of Fig. 1. [00144] In an act 1614, one or more procedures may be implemented to determine sematic Atty. Docket No.010110-23010A information from the text based on the second data stream. For example, the text may be processed using NPL techniques to obtain sematic features of each of at least some of the words in the text. Act 1614 may include an act 1616 in which one or more procedures may be implemented to generate a first feature vector to include encodings of respective words of the text, a second feature vector to include probabilities of the respective words of the text describing content of the text, and a third feature vector to include difficulty scores for the respective words of the text. Act 1614 may also include an act 1618 in which one or more procedures may be implemented to generate a combined sematic feature vector may be generated to include concatenation of the first feature vector, the second feature vector, and the third feature vector. [00145] In an act 1620, one or more procedures may be implemented to determine reading state information based on analyzing the visual information determined based on the first data stream and the semantic information determined based on the second data stream. Determining the reading state information may include determining the reading state information comprises determining one or more of i) that the user is having difficulty comprehending a particular word in the text, ii) that the user is having difficulty comprehending a particular text segment in the text, or iii) that the user’s mind is wandering as the user is reading the text, for example. Act 1620 may include concatenating the semantic features with the visual attention features to generate a concatenated feature vector. Concatenating the semantic features with the visual attention features to generate a concatenated feature vector in the act 1620 may include an act 1622 in which, when concatenating the semantic features with the visual attention features at a time t, one or more procedures may be implemented to identify a first set of words being processed at the time t. Act 1620 may also include an act 1624 in which one or more procedures may be implemented to zero-pad semantic features of a second set of words that has not yet been processed at the time t. In an act 1626, one or more procedures may be implemented to determine reading state of the user at the time t based on the concatenated feature vector that includes the visual features vectors concatenated with the sematic feature vectors of the first set of words being processed at the time t and the zero-padded semantic feature vectors of the second set of words that has not yet been processed at the time t. Thus, in the determination of the reading state information in the state, the sematic features of the words being processed may be given a greater importance, while deemphasizing the semantic features of other words not yet processed in the text. Atty. Docket No.010110-23010A [00146] In an act 1630, one or more procedures may be implemented to perform an operation with respect to the reading state information. Act 1630 may include an act 1632 in which one or more procedures may be implemented to provide an after-reading summary to the user and/or to provide an in- time intervention to the user. Providing the in-time intervention to the user comprises one or more of i) providing a definition of a word, with which the user is having difficulty, to be displayed in real-time on a screen as the user is reading the text on the screen, ii) providing an explanation of a sentence, with which the user is having difficulty, to be displayed in real-time on a screen as the user is reading the text on the screen, or iii) providing an indication that the user’s mind is wandering to be displayed in real-time on a screen as the user is reading the text on the screen, for example. [00147] Fig.17 is a block diagram of a computing system 1700 with which aspects of the disclosure may be practiced. The computing system 1700 includes one or more processors 1702 (sometimes collectively referred to herein as simply “processor 1702”) and one or more memories 1704 (sometimes collectively referred to herein as simply “memory 1704”) coupled to the processor 1702. In some aspects, the computing system 1700 may also include a display 1706 and one or more storage devices 1708 (sometimes collectively referred to herein as simply “storage device 1708” or “memory 1708”). In other aspects, the system 1700 may omit the display 1706 and/or the storage device 1708. In some aspects, the display 1706 and/or the storage device 1708 may be remote from the computing system 1700, and may be communicatively coupled via a suitable network (e.g., comprising one or more wired and/or wireless networks) to the computing system 1700. The memory 1704 is used to store instructions or instruction sets to be executed on the processor 1702. In this example, training instructions 1710 and reading state analysis instructions 1720, which may include reading state prediction instructions 1722 and, in some cases, reading intervention instructions 1724, are stored on the memory 1704. The reading state prediction instructions 1722 may include instructions for implementing one or more engines, such as the semantic attention extraction (SAE) engine, the visual attention extraction (VAE) engine, the cross-attention extraction (CAE) engine, and the reading state estimation/explanation engine described above. The instructions or instruction sets may be integrated with one another to any desired extent. In an aspect, a set of machine-learned networks or other engines is stored on the storage device 1708. The set of trained machine-learned networks or other engines may include a complete predictor, a rationale generator, a rationale generator and/or a causal attention generator as described herein, for example. Atty. Docket No.010110-23010A [00148] The execution of the instructions by the processor 1702 may cause the processor 1702 to implement one or more of the methods described herein. In an example, the processor 1702 may be configured to execute the training instructions 1710 to train various neural networks, such as the temporal convolutional network (TCN) described above, by the computing system 1500. The processor 1702 may be configured to execute the reading state prediction instructions 1722 to detect reading states of a user as the user is reading a text on a screen. The processor 1702 may be configured to execute the reading intervention instructions 1724 to provide interventions, such as real-time interventions to be displayed on the screen, when reading difficulties are detected. [00149] The computing system 1700 may include fewer, additional, or alternative elements. For instance, the computing system 1700 may include one or more components directed to network or other communications between the computing system 1700 and other input data acquisition or computing components, such as sensors (e.g., an inward-facing camera and a forward-facing camera) that may be coupled to the computing system 1700 and may provide data streams for analysis by the computing system 1700. [00150] The term "about" is used herein in a manner to include deviations from a specified value that would be understood by one of ordinary skill in the art to effectively be the same as the specified value due to, for instance, the absence of appreciable, detectable, or otherwise effective difference in operation, outcome, characteristic, or other aspect of the disclosed methods and devices. [00151] The present disclosure has been described with reference to specific examples that are intended to be illustrative only and not to be limiting of the disclosure. Changes, additions and/or deletions may be made to the examples without departing from the spirit and scope of the disclosure. [00152] The foregoing description is given for clearness of understanding only, and no unnecessary limitations should be understood therefrom.

Claims

Atty. Docket No.010110-23010A What is Claimed is: 1. A method for detecting a reading state of a user, the method comprising: obtaining, by a processor, a first data stream indicative of one or both of i) eye movement and ii) facial expression of the user as the user is reading a text; determining, by the processor based on the first data stream, visual information indicative of one or both of i) eye movement and ii) facial expressions of the user during reading of the text; obtaining, by the processor, a second data stream indicative of content of the text; determining, by the processor based on the second data stream, semantic information from the text; determining, by the processor, reading state information based on analyzing the visual information determined based on the first data stream and the semantic information determined based on the second data stream; and performing, by the processor, an operation with respect to the reading state information, wherein the reading state information tracking the reading state of the user during reading of the text. 2. The method of claim 1, wherein determining the visual information includes: determining, based on the first data stream, a gaze point sequence indicative of a direction of gaze of the user as the user is reading the text, and extracting visual attention features from the gaze point sequence. 3. The method of claim 2, wherein: determining the gaze point sequence includes: determining, based on the first data stream, an initial gaze point sequence, and performing one or both of i) median filtering and ii) mean filtering of the initial gaze point sequence to generate a smoothed gaze point sequence, and extracting the visual attention features comprises extracting the visual attention features from the smoothed gaze point sequence. Atty. Docket No.010110-23010A 4. The method of claim 2, wherein extracting the visual attention features includes determining, based on the gaze point sequence, visual attention features indicative of one or both of i) word-level processing state of the user and iii) segment-level processing state of the user. 5. The method of claim 4, wherein extracting the visual attention features includes performing normalization of the visual attention features based at least in part on statistical data determined based on visual attention information obtained while the user is reading the text. 6. The method of claim 4, wherein determining the visual attention features indicative of the segment-level processing state of the user includes determining visual attention features indicative of one or more of i) phrase-level processing state of the user, ii) sentence-level processing state of the user, ii) paragraph-level processing state of the user, iv) page-level processing state of the user, or v) document- level processing state of the user. 7. The method of claim 4, wherein determining the visual attention features indicative of the word- level processing state of the user includes, determining, for each word among at least some words of the text, one or more of i) fixation duration on the word, ii) number of fixations on the word, or iii) number of repeated readings of the word. 8. The method of claim 4, wherein determining the visual attention features indicative of the segment-level processing state of the user includes determining, for each of at least some text segments of the text being read by the user, i) dwell time of the user on the text segment, ii) forward saccade time exhibited by the user as the user is reading the text segment and iii) backward saccade time exhibited by the user as the user is reading the text segment. 9. The method of claim 4, wherein determining the visual attention features further includes using a deep neural network (DNN) to extract deep features from the gaze point sequence. 10. The method of claim 4, wherein determining the visual information includes determining facial features indicative of emotional state of the user as the user is reading the text. 11. The method of claim 4, wherein determining the semantic information includes: generating a first feature vector to include encodings of respective words of the text; Atty. Docket No.010110-23010A generating a second feature vector to include probabilities of the respective words of the text describing content of the text; generating a third feature vector to include difficulty scores for the respective words of the text; and determining semantic features based on a combined sematic feature vector that includes concatenation of the first feature vector, the second feature vector, and the third feature vector. 12. The method of claim 11, further comprising: concatenating the semantic features with the visual attention features to generate a concatenated feature vector, including, when fusing the semantic features with the visual attention features at a time t, identifying a first set of words being processed at the time t, and zero-padding semantic features of a second set of words that has not yet been processed at the time t, and determining the reading state information based on analyzing the concatenated feature vector. 13. The method of claim 1, wherein determining the reading state information comprises determining one or more of i) that the user is having difficulty comprehending a particular word in the text, ii) that the user is having difficulty comprehending a particular text segment in the text, or iii) that mind of the user is wandering as the user is reading the text. 14. The method of claim 1, wherein performing the operation with respect to the reading state comprises one or more of i) providing an after-reading summary to the user or ii) providing an in-time intervention to the user. 15. The method of claim 14, wherein providing the in-time intervention to the user comprises one or more of i) providing a definition of a word, with which the user is having difficulty, to be displayed in real- time on a screen as the user is reading the text on the screen, ii) providing an explanation of a sentence, with which the user is having difficulty, to be displayed in real-time on a screen as the user is reading the text on the screen, or iii) providing an indication that mind of the user is wandering to be displayed in real- time on a screen as the user is reading the text on the screen. 16. The method of claim 1, wherein: obtaining the first data stream comprises obtaining one or more images depicting an eye region Atty. Docket No.010110-23010A of a face of the user, and obtaining the second data stream comprises obtained one or more images depicting the text. 17. A system comprising: a first sensor configured to generate a first data stream indicative of one or both of i) eye movement and ii) facial expression of a user as the user is reading a text; a second sensor configured to generate a second data stream indicative of content of the text; and a reading cognition analysis engine implemented on one or more integrated circuits, the reading cognition analysis engine configured to: obtain the first data stream from the first sensor; determine, based on the first data stream, visual information indicative of one or both of i) eye movements and ii) facial expressions of the user during reading of the text; obtain the second data stream from the second sensor; determine, based on the second data stream, semantic information from the text; determine reading state information based on analyzing the visual information determined based on the first data stream and the semantic information determined based on the second data stream, wherein the reading state information tracks a reading state of the user during reading of the text; and perform an operation with respect to the reading state information. 18. The system of claim 17, wherein the reading cognition analysis engine is configured to determine the reading state information at least by determining one or more of i) that the user is having difficulty comprehending a particular word in the text, ii) that the user is having difficulty comprehending a particular text segment in the text, or iii) that mind of the user is wandering as the user is reading the text. 19. The system of claim 17, wherein the reading cognition analysis engine is configured to perform the operation with respect to the reading state information at least by performing one or more of i) providing an after-reading summary to the user or ii) providing an in-time intervention to the user. Atty. Docket No.010110-23010A 20. The system of claim 19, wherein the reading cognition analysis engine is configured to provide the in-time intervention to the user at least by performing one or more of i) causing a definition of a word, with which the user is having difficulty, to be displayed in real-time on a screen as the user is reading the text on the screen, ii) causing an explanation of a sentence, with which the user is having difficulty, to be displayed in real-time on a screen as the user is reading the text on the screen, or iii) casing an indication that mind of the user is wandering to be displayed in real-time on a screen as the user is reading the text on the screen.
EP24816404.8A 2023-05-30 2024-05-30 Detecting reading state of a user based on visual attention information and semantic features of text Pending EP4720825A2 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202363469706P 2023-05-30 2023-05-30
PCT/US2024/031633 WO2024249609A2 (en) 2023-05-30 2024-05-30 Detecting reading state of a user based on visual attention information and semantic features of text

Publications (1)

Publication Number Publication Date
EP4720825A2 true EP4720825A2 (en) 2026-04-08

Family

ID=93658836

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24816404.8A Pending EP4720825A2 (en) 2023-05-30 2024-05-30 Detecting reading state of a user based on visual attention information and semantic features of text

Country Status (2)

Country Link
EP (1) EP4720825A2 (en)
WO (1) WO2024249609A2 (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119854019B (en) * 2025-01-15 2025-09-30 哈尔滨工业大学 Industrial control intrusion detection method based on ISAE self-encoder and AFF feature fusion

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11604832B2 (en) * 2019-01-03 2023-03-14 Lucomm Technologies, Inc. System for physical-virtual environment fusion

Also Published As

Publication number Publication date
WO2024249609A3 (en) 2025-01-16
WO2024249609A2 (en) 2024-12-05

Similar Documents

Publication Publication Date Title
Nerušil et al. Eye tracking based dyslexia detection using a holistic approach
US10002311B1 (en) Generating an enriched knowledge base from annotated images
KR20210019266A (en) Apparatus and method for diagnosis of reading ability based on machine learning using eye tracking
Prabhu et al. Harnessing emotions for depression detection
Palliya Guruge et al. Advances in multimodal behavioral analytics for early dementia diagnosis: A review
Qi et al. Cases: A cognition-aware smart eyewear system for understanding how people read
Harisinghani et al. Classification of alzheimer's using deep-learning methods on webcam-based gaze data
Benabderrahmane et al. A novel multi-modal model to assist the diagnosis of autism spectrum disorder using eye-tracking data
Xue et al. Enhancing online learning: A multimodal approach for cognitive load assessment
Baray et al. Eog-based reading detection in the wild using spectrograms and nested classification approach
Hijazi et al. Dynamically predicting comprehension difficulties through physiological data and intelligent wearables
Svaricek et al. INSIGHT: Combining Fixation Visualisations and Residual Neural Networks for Dyslexia Classification From Eye‐Tracking Data
EP4720825A2 (en) Detecting reading state of a user based on visual attention information and semantic features of text
Ranjana et al. ADET MODEL: Real time autism detection via eye tracking model using retinal scan images
Bottos et al. An approach to track reading progression using eye-gaze fixation points
CN120732420A (en) Evaluation and feedback method and system combining eye movement and visual and audio-visual algorithm
Hollenstein Leveraging cognitive processing signals for natural language understanding
CN120613119A (en) Method, device, medium and program product for emotion assessment
Cavicchio et al. Multimodal corpora annotation: Validation methods to assess coding scheme reliability
Guo et al. A visually grounded language model for fetal ultrasound understanding
McTear et al. Affective conversational interfaces
Malathi et al. Automated Detection of Language Disorders in Children Using NLP and Machine Learning
Rathod et al. Towards Smarter E-Learning: A Machine Learning Approach to Emotion Recognition and Text Analysis
Kugapriya et al. UNWIND–a mobile application that provides emotional support for working women
Shangareev et al. Reading progress tracking: A novel autoencoder model approach

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251219

AK Designated contracting states

Kind code of ref document: A2

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR