EP4637548A1 - Establishing optimal aggregation of data in signals generated in free-living scenarios - Google Patents
Establishing optimal aggregation of data in signals generated in free-living scenariosInfo
- Publication number
- EP4637548A1 EP4637548A1 EP23908182.1A EP23908182A EP4637548A1 EP 4637548 A1 EP4637548 A1 EP 4637548A1 EP 23908182 A EP23908182 A EP 23908182A EP 4637548 A1 EP4637548 A1 EP 4637548A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- signal
- data
- aggregated
- values
- candidate
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/48—Other medical applications
- A61B5/486—Biofeedback
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/103—Measuring devices for testing the shape, pattern, colour, size or movement of the body or parts thereof, for diagnostic purposes
- A61B5/11—Measuring movement of the entire body or parts thereof, e.g. head or hand tremor or mobility of a limb
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61B—DIAGNOSIS; SURGERY; IDENTIFICATION
- A61B5/00—Measuring for diagnostic purposes; Identification of persons
- A61B5/72—Signal processing specially adapted for physiological signals or for diagnostic purposes
- A61B5/7235—Details of waveform analysis
- A61B5/7264—Classification of physiological signals or data, e.g. using neural networks, statistical classifiers, expert systems or fuzzy systems
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H10/00—ICT specially adapted for the handling or processing of patient-related medical or healthcare data
- G16H10/20—ICT specially adapted for the handling or processing of patient-related medical or healthcare data for electronic clinical trials or questionnaires
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/30—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/70—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients
Definitions
- Various embodiments concern computer programs and associated computer-implemented techniques for determining the reliability of digital biomarkers derived from information collected in free-living settings.
- biomarker a portmanteau of “biological” and “marker” - is commonly used to refer to a measurable indicator of a physiological disease (also called a “physiological condition” or “physiological ailment”).
- biomarkers were evaluated through analysis of blood, urine, or soft tissue taken from a living body, generally for the purpose of predicting the onset of a physiological disease, monitoring the progression of a physiological disease, or determining the pharmacologic response to a therapeutic intervention for a physiological disease.
- biomarkers An emerging field of biomarkers relies on the analysis of measurements generated by sensors that monitor physiological aspects of living bodies. These biomarkers are commonly called “digital biomarkers,” since insights into the health of a given individual can be surfaced through analysis of measurements generated by sensors embedded in digital computing devices (or simply “computing devices”). Because computing devices with embedded sensors are now ubiquitous in society, digital biomarkers can serve as a reliable tool for advancing targeted treatment and guidance in health care. BRIEF DESCRIPTION OF THE DRAWINGS
- Figure 1 illustrates a network environment that includes a diagnostic platform that is executed by a computing device.
- Figure 2 illustrates an example of a computing device that is able to implement a diagnostic platform designed to establish digital biomarkers through analysis of signals and verify the reliability of those digital biomarkers.
- Figure 3 depicts an example of a communication environment that includes a diagnostic platform that is configured to receive several types of data.
- Figure 4 depicts another example of a communication environment that includes a diagnostic platform that is configured to obtain data from one or more sources.
- Figure 5 includes a schematic diagram of an approach to establishing reliability of digital biomarkers determined through analysis of data collected during free-living scenarios.
- Figure 6 illustrates how a given measurement (i.e., step count) can be aggregated across two different candidate aggregation scopes (i.e., one day and two days).
- Figure 7 includes an example of a plot in which reliability values (here, intraclass correlation coefficients or “ICCs”) are plotted against candidate aggregation scopes.
- reliability values here, intraclass correlation coefficients or “ICCs”
- Figure 8 includes an example of an aggregated value A given an optimal aggregation scope of three days.
- Figures 9A-C illustrate several different methods that can be employed to generate paired aggregated values.
- Figure 10 shows an example in which test-retest reliability values are calculated for different types of motor measurements based on signals output by sensors.
- Figure 11 includes a flow diagram of a process for establishing an optimal scope for aggregation of measurements derived from data included in a signal generated over time.
- Figure 12 includes a flow diagram of a process for estimating the reliability of a signal that includes values generated by an inertial measurement unit.
- Figure 13 includes a block diagram of a processing system in which at least some operations described herein can be implemented.
- Digital biomarkers are quantifiable measures that are established through analysis of measurements collected by sensors. As further discussed below, these sensors are generally included in computing devices that are worn on, implanted in, or digested by a living body. Over time, each of these sensors collects measurements, and these data can be used to explain, influence, or predict health-related outcomes.
- Digital biomarkers also allow for more targeted health care, as individualized digital biomarkers can be established and recorded to create personal baselines for health. Because of the important role that digital biomarkers can play in providing targeted health care, the measurements collected by the sensors must be reliable. There are no standard approaches for establishing test-retest reliability of measurements generated by sensors embedded in computing devices in free-living scenarios, however. This is likely due to the limited availability of sufficiently longitudinal data in studies or the lack of infrastructure for processing large amounts of data generated by sensors in free-living scenarios. The lack of a standard approach means that reliability of a continuous signal - from which digital biomarkers can be derived - is difficult to consistently verify.
- a computer program that is designed to establish the optimal scope for aggregating data collected in free-living scenarios and evaluate the test-retest reliability of the data, as well as associated computing systems (or simply “systems”) for implementing the same.
- these data may include measurements generated by a sensor, questionnaire responses provided by individuals, and the like.
- the computer program offers an end-to-end system for measuring test-retest reliability of digital measurements that are collected in free-living settings, informing the appropriate level of aggregation for a given measurement to determine accurate reliability.
- Reliability is generally considered a quantitative measurement of data, indicating whether additional data of a similar nature will remain consistent.
- a metric that is indicative of reliability may be used to convey the likelihood that a measurement generated by a sensor that monitors a characteristic of an individual will be consistent with repeated measurements by the sensor.
- Reliability has historically been difficult to quantify in free-living scenarios, as there is greater variability in the data than more controlled scenarios, resulting in less consistency.
- Free-living scenarios also pose a challenge in developing scalable approaches to analyzing large volumes of data, especially when collected at a high frequency, with the presence of noisiness and missingness. For example, data collected during one interval of time may be noisier than data collected during another interval of time, or data may simply not be collected during an interval of time. Approaches to appropriately aggregating data generated during free-living scenarios are important, so that useful insights can be retrieved, derived, or inferred from the data.
- Intraclass correlation - or the intraclass correlation coefficient (“ICC”) - is a descriptive statistic that can be used when quantitative measurements are made on units that are organized into groups. The ICC describes how strongly units in the same group resemble one another.
- the computer program may default to a bootstrapping approach; however, the complexity of construction and execution of random effects models makes bootstrapping a less ideal candidate because of the computational cost of solving for the maximum likelihood estimate for each bootstrapped sample.
- statistical inference from random effects models may be based on asymptotic assumptions.
- random effects models are not well suited for establishing test-retest reliability on longitudinal data, because those models overblend points in time that are further apart into calculation of the reliability metric. This can inflate within-individual variability (also called “intra-patient variability”), and therefore may unjustly penalize reliability estimates.
- Embodiments may be described in the context of executable instructions for the purpose of illustration. However, those skilled in the art will recognize that aspects of the technology could be implemented via hardware, firmware, or software.
- a computer program that is representative of a diagnostic platform may be executed by the processor of a computing device.
- the computer program may interface, directly or indirectly, with hardware, firmware, or other software implemented on the computing devices.
- the computer program may obtain data that is collected, computed, or otherwise obtained by a sensor included in the computing device.
- the computer program may obtain data that is collected, computed, or otherwise obtained by a sensor included in another computing device. This data may be processed in accordance with the approaches described herein, so as to establish its reliability (and by extension, reliability of digital biomarkers that are based on this data).
- references in the present disclosure to “an embodiment” or “some embodiments” means that the feature, function, structure, or characteristic being described is included in at least one embodiment. Occurrences of such phrases do not necessarily refer to the same embodiment, nor do they necessarily refer to alternative embodiments that are mutually exclusive of one another.
- connection or coupling can be physical, logical, or a combination thereof.
- elements may be electrically or communicatively connected to one another despite not sharing a physical connection.
- module may refer broadly to software, firmware, hardware, or combinations thereof. Modules are typically functional components that generate one or more outputs based on one or more inputs.
- a computer program may include or utilize one or more modules. For example, a computer program may utilize multiple modules that are responsible for completing different tasks, or a computer program may utilize a single module that is responsible for completing multiple tasks.
- the terms "longitudinal signal” and “longitudinal data” may be used to refer to a series of values that are periodically generated over an interval of time rather than on an ad hoc basis.
- a parameter may be measured by a sensor multiple times per second (e.g., at a sampling rate of 25, 50, or 100 hertz) in an ongoing manner, or a parameter may be measured once per minute for several hours.
- the length of time over which the series of values extend may depend on the nature of the parameter being measured.
- the term “parameter” is generally used to describe the unit of measure of the underlying signal.
- An example of a parameter is x-axis acceleration.
- a “measurement” may be used to refer to data that are either collected directly from a computing device or derived from a raw signal output by a sensor included in a computing device.
- a “measurement” may be a patient-reported outcome provided via a symptom survey, or a “measurement” may be step count as derived from values output by an accelerometer that is indicative of acceleration along one or more axes.
- step count may become a digital biomarker of disease progression (also called a “progression biomarker”) if it can be proven that step count is sensitive to worsening of mobility of the corresponding patient. However, before this relationship is established, step count may simply be a “measurement.”
- free-living scenario may be used to refer to a scenario in which an individual under observation is permitted to live her life as if not under observation.
- fitness tracking devices also called “fitness trackers”
- Observation in free-living scenarios is generally less prone to the intentional alteration of habits that commonly occurs in more controlled environments.
- Figure 1 illustrates a network environment 100 that includes a diagnostic platform 102 that is executed by a computing device 104.
- An individual also referred to as a “user” can interact with the diagnostic platform 102 via interfaces 106.
- a patient may be able to access an interface through which information regarding a physiological disease, such as digital biomarkers or analyses of digital markers, can be reviewed.
- a healthcare professional may be able to access an interface through which information regarding patients, such as digital biomarkers or analyses or digital markers, can be reviewed.
- the interfaces 106 may allow for the review of physiological data, examination of outputs produced by the diagnostic platform 102, and management of preferences.
- Some interfaces may be configured to facilitate interactions between patients and healthcare professionals, while other interfaces may be configured to serve as informative dashboards for patients or healthcare professionals.
- the physiological data obtained by the diagnostic platform 102 could be associated with the individual accessing the interfaces 106 or some other person.
- the interfaces 106 may enable a person diagnosed with a physiological disease to view her own physiological data.
- the interfaces may enable an individual to view physiological data associated with another person.
- the individual may be a healthcare professional who is responsible for monitoring, managing, or treating the other person. Examples of healthcare professionals include physicians, nurses, dietitians, and the like.
- the diagnostic platform 102 can reside in a network environment 100.
- the computing device 104 on which the diagnostic platform 102 resides can be connected to one or more networks 108A-B.
- the computing device 104 could be connected to a personal area network (“PAN”), local area network (“LAN”), wide area network (“WAN”), metropolitan area network (“MAN”), or cellular network.
- PAN personal area network
- LAN local area network
- WAN wide area network
- MAN metropolitan area network
- cellular network cellular network.
- the computing device 104 is a computer server
- the computing device 104 may be accessible to users via respective mobile phones that are connected to the Internet via LANs.
- the physiological data to be examined by the diagnostic platform 102 may be generated by the respective mobile phones (e.g., sensors included in the respective mobile phones) or acquired by the respective mobile phones.
- the computing device 104 may be connected to one or more other computing devices over a short-range wireless connectivity technology, such as Bluetooth®, Near Field Communication (“NFC”), Wi-Fi® Direct (also referred to as “Wi-Fi P2P”), and the like.
- the diagnostic platform 102 could be embodied as a mobile application that is executed by a mobile phone.
- the mobile phone may be communicative connected - via a wireless communication channel - to a source from which to acquire physiological data.
- the source could be a watch, fitness tracker, or another wearable computing device, for example.
- the physiological could alternatively be obtained from another computer program executing on the mobile phone.
- the physiological data could instead be acquired from another mobile application executing on the mobile phone or the operating system of the mobile phone.
- the interfaces 106 may be accessible via a web browser, desktop application, mobile application, or another form of computer program.
- a patient may be able to access interfaces through which information regarding her own health, such as digital biomarkers or analyses of digital biomarkers, is provided by a mobile application executing on a mobile phone.
- a healthcare professional may be able to access interfaces through which information regarding one or more patients can be reviewed via a web browser.
- the interfaces 106 generated by the diagnostic platform 102 may be accessible on various computing devices, including mobile phones, tablet computers, desktop computers, and the like.
- the diagnostic platform 102 is executed - at least partially - by a cloud computing service operated by, for example, Amazon Web Services®, Google Cloud PlatformTM, or Microsoft Azure®.
- the computing device 104 may be representative of a computer server that is part of a server system 1 10.
- the server system 110 is comprised of multiple computer servers. These computer servers can include different types of data (e.g., physiological data and information regarding patients, such as name, demographic information, disease classification, etc.), algorithms for processing incoming data, and other assets. Those skilled in the art will recognize that these data could also be distributed among the server system 1 10 and one or more computing devices. As an example, data that is input by, or related to, patients may be stored on, and processed by, their own computing devices for security or privacy purposes.
- FIG. 1 illustrates an example of a computing device 200 that is able to implement a diagnostic platform 212 designed to establish digital biomarkers through analysis of signals and verify the reliability of those digital biomarkers.
- the computing device 200 can include a processor 202, memory 204, display mechanism 206, communication module 208 and sensor suite 210. Each of these components is discussed in greater detail below.
- the computing device 200 may not include the display mechanism 206 or sensor suite 210.
- the computing device 200 may not include the display mechanism 206 and sensor suite 210.
- the computing device 200 on which the diagnostic platform 212 resides may not include sensors in some embodiments.
- the data examined by the diagnostic platform 212 could instead be generated by one or more sensors 222A-N that are external to the computing device 200.
- the computing device 200 may be a mobile phone, and the sensors 222A-N may be included in a watch or fitness tracker that is communicatively connected to the computing device 200.
- the processor 202 can have generic characteristics similar to general- purpose processors, or the processor 202 may be an application-specific integrated circuit (“ASIC”) that provides control functions to the computing device 200. As shown in Figure 2, the processor 202 can be coupled to all components of the computing device 200, either directly or indirectly, for communication purposes.
- ASIC application-specific integrated circuit
- the memory 204 can be comprised of any suitable type of storage medium, such as static random-access memory (“SRAM”), dynamic randomaccess memory (“DRAM”), electrically erasable programmable read-only memory (“EEPROM”), flash memory, or registers.
- SRAM static random-access memory
- DRAM dynamic randomaccess memory
- EEPROM electrically erasable programmable read-only memory
- flash memory or registers.
- the memory 204 can also store data generated by the processor 202 (e.g., when executing the modules of the diagnostic platform 212).
- the memory 204 is merely an abstract representation of a storage environment.
- the memory 204 could be comprised of actual integrated circuits (also called “chips”).
- the display mechanism 206 can be any mechanism that is operable to visually convey information to a user.
- the display mechanism 206 can be a panel that includes light-emitting diodes (“LEDs”), organic LEDs, liquid crystal elements, or electrophoretic elements.
- LEDs light-emitting diodes
- outputs produced by the diagnostic module 212 can be posted to the display mechanism 206 for review by a user of the computing device 200.
- the communication module 208 may be responsible for managing communications external to the computing device 200.
- the communication module 208 can be wireless communication circuitry that is able to establish wireless communication channels with other computing devices. Examples of wireless communication circuitry include 2.4 gigahertz (“GHz”) and 5 GHz chipsets compatible with Institute of Electrical and Electronics Engineers (“IEEE”) 802.11 - also referred to as “Wi-Fi chipsets.”
- the communication module 208 may be representative of a chipset configured for Bluetooth, NFC, and the like.
- Some computing devices - like mobile phones, tablet computers, and the like - are able to wirelessly communicate via separate channels, while other computing devices - like watches and fitness trackers - tend to wirelessly communicate via a single channel.
- the communication module 208 may be one of multiple communication modules implemented in the computing device 200, or the communication module 208 may be the only communication module implemented in the computing device 200.
- the nature, number, and type of communication channels established by the computing device 200 - and more specifically, the communication module 208 - can depend on (i) the sources from which data is received by the diagnostic platform 212 and (ii) the destinations to which data is transmitted by the diagnostic platform 212. Assume, for example, that the diagnostic platform 212 resides on a mobile phone in the form of a mobile application.
- the communication module 208 can communicate with sensors 222A-N external to the computing device 200 from which to obtain data.
- the communication module 208 may communicate with a server system (e.g., server system 1 10 of Figure 1 ) to which analyses of the data - or the data itself - are transmitted.
- the computing device 200 may include a motion sensor whose output is indicative of motion of the computing device 200 as a whole.
- motion sensors include accelerometers and gyroscopes.
- the motion sensor is implemented in an inertial measurement unit (“IMU”) that measures the force, angular rate, or orientation of the computing device 200.
- IMU inertial measurement unit
- the IMU may accomplish this through the use of one or more accelerometers, one or more gyroscopes, one or more magnetometers, or any combination thereof.
- the IMU could be a 6-axis IMU that draws low current and therefore is suitable for “always-on” applications in battery-driven computing devices, or the IMU could be a 3-axis IMU that includes logic-level shifting circuitry that can readily interface with a microcontroller.
- the computing device 200 may include an ambient light sensor whose output is indicative of the amount of light in the ambient environment.
- sensors 222A- N could also be acquired from sensors 222A- N that are external to the computing device 200.
- These sensors 222A-N could be included in another computing device, such as a watch or fitness tracker, that is directly connected to the computing device 200.
- these sensors 222A-N could be discrete sensing units that are directly or indirectly connected to the computing device 200.
- sensor 222A may be a pulse oximeter that monitors the oxygen saturation of the user of the computing device 200 or another person, for the purpose of creating a photoplethysmogram (“PPG”).
- PPG photoplethysmogram
- sensor 222A may be placed on a thin part of a living body, usually a fingertip or earlobe, and then pass two wavelengths of light through that body part toward a photodetector.
- the photodetector can measure the changing absorbance at each wavelength, allowing sensor 222A to determine the absorbances due to the pulsing of arterial blood through that body part.
- the diagnostic platform 212 is referred to as a computer program that resides within the memory 204.
- the diagnostic platform 212 could be comprised of software, firmware, or hardware that is implemented in, or accessible to, the computing device 200.
- the diagnostic platform 212 can include a processing module 214, computation module 216, aggregation module 218, and graphical user interface (“GUI”) module 220. These modules could be integral parts of the diagnostic platform 212, or these modules could be logically separate from the diagnostic platform 212 but operate “alongside” it. Together, these modules enable the diagnostic platform 212 to establish the reliability of data to be used to gain insights into the health of a given individual. As mentioned above, the given individual could be the user of the computing device 200 or another person.
- the processing module 214 can process data obtained by the diagnostic platform 212 into a format that is suitable for the other modules. For example, the processing module 214 can apply operations to sensor data obtained from the sensor suite 210 or sensors 222A-N in preparation for analysis by the other modules of the diagnostic platform 212. For example, the processing module 214 can filter or alter the sensor data, such that the sensor data can be more readily analyzed. As another example, the processing module 214 may parse the sensor data in order to temporally align the dataset obtained from each source. Accordingly, the processing module 214 may be responsible for ensuring that the appropriate sensor data is accessible to the other modules of the diagnostic platform 212.
- the computation module 216 may be responsible for deriving different measurements based on the processed data received from the processing module 214. As further discussed below, these different measurements may depend on the nature of the insights to be surfaced by the diagnostic platform 210. For example, these different measurements may be representative of walking-related measurements, running-related measurements, sleeping-related measurements, eating-related measurements, stress-related measurements, or any combination thereof.
- the computation module 216 can compute, infer, or otherwise establish measurements for different segments of processed data. These segments may be referred to as “windows” of processed data.
- the aggregation module 218 may implement an aggregation engine that, in operation, implements two programs in sequence.
- the first program may provide a method for programmatically determining the appropriate scope of aggregation for the processed data output by the processing module 412, while the second program may implement one of multiple methods for pairing the data for calculation of reliability values.
- the GUI module 220 may be responsible for generating interfaces that are viewable on the display mechanism 206. Various types of information can be presented on these interfaces. For example, digital biomarkers that are calculated, derived, or otherwise obtained by the computation module 216 may be presented on an interface for display to the user. Similarly, analyses of the digital biomarkers may be presented on the interface for display to the user. [0059] Figure 3 depicts an example of a communication environment 300 that includes a diagnostic platform 302 that is configured to receive several types of data.
- the diagnostic platform 302 receives response data 304, first sensor data 306 that is generated by a first sensor (e.g., sensor 222A of Figure 2), and second sensor data 308 that is generated by a second sensor (e.g., sensor 222B of Figure 2).
- first sensor e.g., sensor 222A of Figure 2
- second sensor data 308 that is generated by a second sensor (e.g., sensor 222B of Figure 2).
- treatment data e.g., including indicators of physiological disease, treatment regimens, etc.
- the response data 304 could be obtained directly on the computing device on which the diagnostic platform 302 is executing.
- the response data 304 may be representative of input provided by an individual through interfaces generated by the mobile application.
- the input may be representative of responses to queries included in a questionnaire that is designed to elicit responses indicative of health.
- the response data 304 could be obtained from another computing device.
- the mobile phone could obtain the response data 304 from a server system (e.g., server system 110 of Figure 1 ).
- the server system may manage a datastore in which responses provided by various individuals to questionnaires are documents.
- sensor data could be obtained from sensors included in the computing device that is responsible for executing the diagnostic platform 302, or sensor data could be obtained from sensors that are external to the computing device that is responsible for executing the diagnostic platform 302.
- the first and second sensor data 306, 308 may be generated by sensors included in the computing device on which the diagnostic platform 302 is executing, or the first and second sensor data 306, 308 may be generated by sensors that are external to the computing device on which the diagnostic platform 302 is executing.
- the first sensor data 306 may be generated by an IMU implemented in the computing device on which the diagnostic platform 302 is executing, while the second sensor data 308 may be generated by a pulse oximeter that is communicatively connected - either directly or indirectly - to the computing device on which the diagnostic platform 302 is executing.
- Figure 4 depicts another example of a communication environment 400 that includes a diagnostic platform 402 that is configured to obtain data from one or more sources.
- the diagnostic platform 402 may obtain data from a mobile phone 404, watch 406, laptop computer 408, or server system 410 (collectively referred to as the “networked devices”).
- the diagnostic platform 402 may obtain response data (e.g., response data 304 of Figure 3) from the laptop computer 408 or server system 410.
- the diagnostic platform 402 may obtain sensor data (e.g., sensor data 306, 308 of Figure 3) from the mobile phone 404 or watch 406.
- the networked devices can be connected to the diagnostic platform 402 via one or more networks. These networks can include PANs, LANs, WANs, MANs, cellular networks, the Internet, etc. Additionally or alternatively, the networked devices may communicate with one another over a short-range wireless connectivity technology. For example, if the diagnostic platform 402 resides on the mobile phone 404 in the form of a mobile application, data may be obtained from the watch 406 over a Bluetooth communication channel while data may be obtained from the server system 410 over the Internet via a Wi-Fi communication channel.
- Embodiments of the communication environment 400 may include a subset of the networked devices.
- the communication environment 400 may include a diagnostic platform 402 that obtains, in real time, data from the mobile phone 404 and watch 406 as that data is generated over the course of a free-living scenario. Additional data could be obtained from the server system 410 on a periodic basis (e.g., daily or weekly).
- Figure 5 includes a schematic diagram of an approach to establishing reliability of digital biomarkers determined through analysis of data collected during free-living scenarios. At a high level, the approach includes five steps, each of which is discussed in greater detail below.
- a diagnostic platform can collect signals from one or more sources (step 501 ).
- the diagnostic platform may collect, from a server system (e.g., server system 110 of Figure 1 ), a signal that is representative of responses to a questionnaire that is accessible via interfaces generated by the diagnostic platform.
- the diagnostic platform may collect, from a watch, signals that are representative of measurements generated by sensors included in the watch.
- the signal collected by the diagnostic platform is representative of a series of measurements generated by a sensor in temporal order
- the signal is preferably a longitudinal signal with values generated over an interval of time.
- the signal could include values that are generated at a predetermined cadence by an IMll.
- the signal could include values that are generated at a predetermined cadence by a pulse oximeter.
- the duration of the signal - and the cadence of its values - may depend on the nature of the sensor.
- data collected by the diagnostic platform may not be limited to “raw signals” generated by sensors.
- the diagnostic platform could collect patient-reported outcomes that are provided, for example, through interfaces generated by the diagnostic platform.
- the diagnostic platform may surface, through an interface, a questionnaire that solicits responses to questions from patients whose health is being monitored.
- patient-reported outcomes could be collected using “open” queries (e.g., where each patient is permitted to answer as she sees fit) or “closed” queries. Closed queries may be posed using a rating scale such as Likert scale, so that responses can be readily scaled.
- topics that are generally well suited for open queries vital signs and symptoms being experienced while examples of topics that are generally well suited for closed queries include current pain level and current ability to perform regular activities.
- the signal is preferably collected by the diagnostic platform - or at least values are generated by the source - at a relatively high frequency. How frequently data is collected may depend on not only its type but also its intended application (e.g., the digital biomarkers that are based on such data). For example, data may be collected every second by sensors included in a watch or fitness tracker, whereas patient-reported outcomes may be collected on a daily or weekly basis (e.g., via interfaces generated by the diagnostic platform).
- the data that is represented by, and conveyed via, the signal collected by the diagnostic platform can be stored in a data structure by the diagnostic platform.
- the data may be populated into one or more data structures (e.g., tables) that are accessible to the diagnostic platform.
- the diagnostic platform can include a data processing component - namely, a processing module (e.g., processing module 214 of Figure 2) - that may be responsible for applying a set of preprocessing operations to the signal. Said another way, the processing module may be responsible for examining the data included in the signal (step 502).
- the preprocessing operations that are applied to the signal may depend on the type of data contained therein. Assume, for example, that the signal includes measurements generated by a sensor. In such a scenario, the processing module can apply standard preprocessing operations such as resampling, bias removal, and noise removal using a low-pass filter or high-pass filter.
- the processing module may code the data into a numeric format and then filter the data (e.g., to remove outlier answers, non-responsive answers, incomprehensible answers, etc.). Moreover, the processing module may perform a binning operation, in which the numerically coded data is sorted into categories or “bins.” As an example, patient-reported outcomes may be sorted into different age ranges (e.g., 0-19, 18-30, 30-45, etc.), geographical locations (e.g., by country, region, state, or county), disease classification (e.g., critical, severe, moderate, and mild), races (e.g., White, Asian, Hispanic, Black, American Indian, Native Hawaiian and Other Pacific Islander), gender (e.g., male, female, and non-binary), and the like.
- age ranges e.g., 0-19, 18-30, 30-45, etc.
- geographical locations e.g., by country, region, state, or county
- disease classification e.g., critical, severe,
- the preprocessed data can then be provided to a data computing component - namely, a computation module (e.g., computation module 216 of Figure 2) - that may be responsible for deriving different measurements based on the preprocessed data (step 503).
- a computation module e.g., computation module 216 of Figure 2
- no further derivation may be necessary.
- the computation module can apply heuristics and/or machine-learning models (or simply “models”) to derive the measurements using the preprocessed data. For example, to derive a set of walking-related measurements, the computation module may apply one or more activity classifiers can be applied to windows of preprocessed data that includes values generated by an IMU, such that each window is classified as either ambulatory or non-ambulatory. A similar approach could be used to classify each window as sleeping or non-sleeping, stressed or non-stressed, etc. Each window may be representative of a segment of preprocessed data of predetermined length (e.g., 5, 10, or 20 seconds).
- predetermined length e.g., 5, 10, or 20 seconds
- each activity classifier may be representative of a model that when applied to a window of preprocessed data, classifies that window based on an analysis of the preprocessed data contained therein.
- the computation module can then identify walking bouts by joining consecutive ambulatory windows and then labeling ambulatory windows of sufficient length (e.g., 10, 20, or 40 seconds) as walking bouts. For each walking bout, the computation module can derive one or more measurements. Examples of such measurements include step count, cadence, bout duration, arm swing magnitude, arm swing velocity, arm swing acceleration, and arm swing range of motion.
- the nature and number of measurements derived from the preprocessed data will depend on various factors, including the nature of the preprocessed data and intended application of measurements derived therefrom.
- the computation module may rely on the same preprocessed data, namely, values generated by an IMU.
- the measurements may differ. Examples of walking-related measurements include step count and cadence as discussed above, while examples of sleeping-related measurements include total sleep duration and wake count.
- an aggregation module e.g., aggregation module 218 of Figure 2
- the aggregation module may implement an analytics engine 500 in order to perform aggregation.
- the analytics engine 500 (or simply “engine”) is representative of the core logic of the aggregation module that, in operation, takes an input (e.g., the preprocessed data), performs at least one operation, and then produces an output (e.g., an indication of an optimal window for aggregation). More specifically, the analytics engine 500 may be responsible for exploring optimal points of aggregation scope and conducting reliability assessments.
- the analytics engine 500 can include two programs that, in operation, are performed in sequence.
- the first program may provide a method for programmatically determining the appropriate scope of aggregation for the data included in the signal collected as input. As further discussed below, the first program can determine the optimal window scope for aggregation based on analysis of a plot.
- the second program may provide multiple methods for pairing the data for calculation of reliability values, as further discussed below.
- the aggregation scope can be defined such that it contains a sufficient number of datapoints to enable a stable measure of the underlying computing device, while also allowing enough flexibility to accurately understand change.
- the aggregation scope may be described as the number of datapoints per patient, and each datapoint may be representative of a periodic (e.g., every second, minute, hour, day, or week) measurement collected for that patient.
- aggregation scopes can - and often do - correspond to intervals of time having different lengths. Assume, for example, that the aggregation scope is defined as 20.
- the aggregation scope may correspond to a 20-minute interval of time, while for a sensor that generates measurements hourly or daily, the aggregation scope may correspond to a 20-hour or 20-day interval of time, respectively.
- the first program may provide a method for programmatically determining the optimal aggregation scope.
- the first program can initially define a range of candidate aggregation scopes.
- these candidate aggregation scopes correspond to, or are based on, the frequency of the data included in the signal received by the diagnostic platform as input.
- the set of candidate aggregation scopes may be defined as 1 , 2, 5, 10, and 20 days.
- the set of candidate aggregation scopes may be defined as 1 , 2, 4, 8, 16, and 24 minutes.
- a set of candidate aggregation scopes is normally defined such that even the shortest candidate aggregation scope includes at least a predetermined number of datapoints.
- an aggregated measurement can be calculated within the candidate scope timeframe and then compared to adjacent aggregation windows.
- the set of candidate aggregation scopes is defined as 1 , 2, 5, 10, and 20 days.
- a given measurement e.g., step count
- the analytics engine 500 may produce, as output, a set of data structures, each of which corresponds to a different one of the set of candidate aggregation scopes.
- Each data structure may include a plurality of values that are representative of the given measurement as computed across the signal collected as input on a rolling basis. Examples of these data structures are provided in Figure 6. Specifically, Figure 6 illustrates how a given measurement (i.e., step count) can be aggregated across two different candidate aggregation scopes (i.e., one day and two days).
- reliability values can be calculated based on the aggregated measures, such that each candidate aggregation scope is associated with a single reliability value.
- the number of aggregated measurements may depend on the number of windows that can be formed in accordance with the corresponding candidate aggregation scope. In Figure 6, for example, step count is determined using non-overlapping windows in accordance with the approach further described below with reference to Figure 9A. Generally, as the candidate aggregation scope becomes larger, the number of windows will decrease - and therefore, the number of aggregated measurements will decrease. However, only one reliability value may be calculated for each candidate aggregation scope. Thus, a single reliability value may be produced for each candidate aggregation scope of a given measurement.
- the reliability values may be ICCs, for example.
- FIG. 7 includes an example of a plot in which reliability values (here, ICCs) are plotted against candidate aggregation scopes.
- reliability values are plotted along the y-axis while different candidate aggregation scopes are plotted along the x-axis, and each line visually illustrates the change in reliability value of a measurement (e.g., step count) along different aggregation scopes.
- the analytics engine 500 can identify the minimum acceptable aggregation scope for the measure, which will allow for maximization of the measure’s ability to resolve change.
- the analytics engine 500 may examine the plot shown in Figure 7 to find the number of days in which the change in reliability value is smaller than a threshold (e.g., 0.05).
- a threshold e.g., 0.05.
- the second program can aggregate the data in accordance with the optimal aggregation scope via different mathematical operators in step 505.
- mathematical operators include the max operator, min operator, X th - percentile operator (where Xis an integer that ranges from 0 to 100, such as 75, 95, etc.), and the like.
- Such action will result in a set of aggregated values within each window defined in accordance with the optimal aggregation scope.
- Figure 8 includes an example of an aggregated value A given an optimal aggregation scope of three days. Referring to Figure 8 as an example, if the mathematical operator is a max operator, then A will be the maximum value of a measurement within three days’ worth of data.
- a first method all of the aggregated values are calculated based on non-overlapping windows, as shown in Figure 9A.
- One advantage of this first method is that the second program can avoid deflating the reliability values.
- the point estimate may be less accurate when the signal is “short” and doesn’t include a large number of datapoints.
- a second method all of the aggregated values are calculated based on overlapping windows, as shown in Figure 9B.
- A' and A" can be calculated by shifting the beginning of the window.
- A' and A" are calculated by shifting the starting day.
- a table of paired aggregated values can be created by having all of the original values paired and shifted values paired (e.g., A-B, A'- B', B'- C , etc.).
- One advantage of this second method is the ability to stabilize the point estimate, while one disadvantage is that statistical inference is impacted and confidence interval (“Cl”) calculation is affected.
- a third method all of the aggregated values are calculated based on non-overlapping windows, as shown in Figure 9C.
- One advantage of this third method is the ability to stabilize the point estimate, while one disadvantage is that implementation can be computationally complicated, with the Cis potentially being smaller than reality.
- the choice of which method to use may be determined based on the intended application.
- the aggregation module may select one of the aforementioned methods based on the intended application of the aggregated values.
- the choice of which method to use may depend on the interest of the user. Accordingly, the user may specify which method to use, for example, through an interface generated by the diagnostic platform.
- the aggregation module can estimate the test-retest reliability of the measurements derived by the computation module.
- Fisher’s original formulation of ICC may be used to estimate test-retest reliability as set forth below in Eq. 1 -3.
- the data may need to be organized into a pair of vectors ( x n l , x n 2 ) having columnar form, as shown in Figures 9A-C. These vectors can be constructed as set forth above.
- both the point estimate and Cl of the ICC may be provided.
- the point estimate can be obtained by computing the ICC from the observed data, while the Cl can be computed using a bootstrapping method.
- the aggregation module can resample the observed data with replacement being performed m times, where m is an integer value (e.g., between 100 and 2,500, and preferably ⁇ 1 ,000).
- m is an integer value (e.g., between 100 and 2,500, and preferably ⁇ 1 ,000).
- the value of the ICC for that given sample can be obtained, and in total, m number of ICCs may be obtained.
- a (100 - a) percent Cl can be obtained by taking the a/2 and (100 - a/2) percentiles of the resultant m ICCs.
- Figure 10 shows an example in which test-retest reliability values are calculated for different types of motor measurements based on digital measures derived from signals output by sensors of a watch.
- the aggregation module may be able to determine different ways of aggregation for longitudinal data generated during free-living scenarios, as discussed above.
- the ability to collect data at high frequency (e.g., multiple datapoints per second, minute, or hour) while offering the appropriate level of aggregation can lead to improved reliability in measurements that are computed based on the data.
- the framework described above can inform the appropriate level of aggregation for a given measurement to determine or maximize reliability.
- Figure 11 includes a flow diagram of a process 1 100 for establishing an optimal scope for aggregation of measurements derived from data included in a signal generated over time.
- a diagnostic platform can obtain a plurality of measurements that are derived via analysis of the signal (step 1101 ).
- Each of the plurality of measurements may be associated with a corresponding one of a plurality of segments of the signal.
- the plurality of measurements may be computed, derived, or otherwise established by a computation module as discussed above.
- the nature of each measurement may depend on the nature of the data. For example, if the data is representative of discrete motion values generated by an IMU, then the measurement could be step count or arm swing magnitude. As another example, if the data is representative of discrete absorbance values generated by a pulse oximeter, then the measurement could be heart rate or blood oxygen level.
- the diagnostic platform can define a plurality of candidate scopes for aggregation (step 1102).
- Each of the plurality of candidate scopes corresponds to a different duration.
- the plurality of candidate scopes may be 1 , 2, 5, 10, and 20 days, or the plurality of candidate scopes may be 1 , 2, 4, 8, 16, and 24 minutes.
- the durations of the candidate scopes are normally defined such that even the shortest candidate scope includes at least a predetermined number of datapoints. Therefore, the durations of the plurality of candidate scopes may be defined based on the frequency of datapoints in the signal. Note that the number of candidate scopes can vary.
- the plurality of candidate scopes may include (i) a shortest candidate scope that is defined such that a sufficient number of datapoints (e.g., the predetermined number) are included to enable stable measure and (ii) a longest candidate scope that is defined such that change over time can be understood. If the shortest candidate scope is too short, the aggregated measurements may be too variable (e.g., because some scopes will not include sufficient data). If the longest candidate scope is too long, then it becomes difficult to distinguish between aggregated measurements. The number of candidate scopes between the shortest and longest candidate scopes may vary.
- the diagnostic platform may be programmed to define at least a predetermined number of candidate scopes (e.g., 2, 3, 5) between the shortest and longest candidate scopes, or the diagnostic platform may be programmed to define intervening candidate scopes at a predetermined cadence (e.g., that is determined based on the shortest candidate scope or longest candidate scope).
- a predetermined number of candidate scopes e.g., 2, 3, 5
- intervening candidate scopes e.g., that is determined based on the shortest candidate scope or longest candidate scope.
- the diagnostic platform can calculate, based on the plurality of measurements, an aggregated measurement within that candidate scope (step 1103). Said another way, the diagnostic platform can calculate an aggregated measurement for each of the plurality of candidate scopes, such that a plurality of aggregated measurements are calculated.
- the diagnostic platform can also determine a reliability metric based on the aggregated measurement (step 1 104). Specifically, the diagnostic platform may determine, for each of the plurality of candidate scopes, a reliability metric based on a comparison of the aggregated measurement to aggregated measurements calculated for the next highest candidate scope and/or next lowest candidate scope. For example, the diagnostic platform may compute an ICC for each of the plurality of candidate scopes. At a high level, each ICC may indicate how strongly the corresponding candidate scope resembles nearby candidate scopes.
- the diagnostic platform can identify an optimal scope from among the plurality of candidate scopes based on an analysis of the plurality of reliability metrics determined for the plurality of candidate scopes (step 1 105). For example, the diagnostic platform may plot the plurality of reliability metrics against the plurality of candidate scopes as shown in Figure 7, and then the diagnostic platform may determine the optimal aggregation scope (and therefore, window size) based on an analysis of the plot.
- identifying the optimal aggregation scope may involve consultation with partners (e.g., pharmaceutical manufacturers, healthcare systems, insurers) to establish the windows size of interest. Assume, for example, that a partner is interested in weekly measurement reliability, the diagnostic platform can provide the appropriate weekly measurement. Alternatively, a partner may be interested in conducting aggregation within a window size of 14 days. In this scenario, the diagnostic may use a plot - similar to the one shown in Figure 7 - to infer (i) what reliability this aggregation scope will provide and (ii) whether an inflection point is reached (and therefore, whether a larger window size would help improve reliability).
- partners e.g., pharmaceutical manufacturers, healthcare systems, insurers
- the diagnostic platform may aggregate, via a mathematical operator, data included in the signal, such that a plurality of aggregated values are produced for a plurality of windows defined across the signal in accordance with the optimal scope (step 1 106).
- mathematical operators include the max operator, min operator, and 95 th - percentile operator. Note that in some embodiments, the mathematical operator is one of multiple mathematical operators that are applied to the data included in the signal.
- the diagnostic platform can populate the plurality of aggregated values into a data structure (step 1 107). Specifically, the diagnostic platform may populate the plurality of aggregated values into a tabular data structure, such that the plurality of aggregated values are paired together. There are several different methods that could be employed by the diagnostic platform to generate the paired aggregated values, as further discussed above with reference to Figures 9A-C. [00103] Moreover, the diagnostic platform may estimate test-retest reliability of the signal based on an analysis of the data structure and provide statistical inference of the estimate (step 1 108). For example, Fisher’s original formulation of ICC may be used to estimate the test-retest reliability as set forth above in Eq. 1 -3.
- the diagnostic platform may establish the point estimate and/or Cl of the ICC. Accordingly, estimation of test-retest reliability may be based on point estimates of the ICCs, Cis of the ICCs, or a combination thereof. This estimate of the test-retest reliability can be used in various ways. For example, the diagnostic platform may determine, based on the estimate of the test-retest reliability, reliability of a digital biomarker that is computed using the signal.
- the process 1 100 may vary depending on the nature of the data included in the signal.
- the data is representative of discrete measurements output by a sensor in temporal order.
- the diagnostic platform may apply a classifier to the signal to obtain a plurality of outputs, each of which is representative of a classification of a different portion of the signal.
- the classifier may be representative of a model that when applied to datapoints included in a portion of the signal, classifies that portion based on an analysis of the datapoints.
- the classifier may be designed and/or trained to distinguish between ambulatory periods and non-ambulatory periods, sleeping periods and nonsleeping periods, stressful periods and non-stressful periods, etc.
- the diagnostic platform can join consecutive portions that have the same classification.
- Superset portions that are assigned a given classification (e.g., ambulating) and exceed a given length (e.g., two or more portions) may be identified as bouts.
- the diagnostic platform may derive a measurement for each superset portion.
- the diagnostic platform may determine the step count during an ambulatory bout by summing the steps taken during the portions included in the corresponding superset portion.
- the data is representative of inputs provided by one or more individuals.
- the data may include values that are representative responses to queries by a single person or multiple people.
- the diagnostic platform may codify the signal (and more specifically, its values) into a numerical format. Such an approach allows the inputs to be more readily compared across different portions of the signal.
- the diagnostic platform may filter the coded values included in the signal to remove outlier inputs and/or bin the coded values such that inputs are assigned to categories, for example, representing characteristics of the individual(s).
- the categories may correspond to different age ranges, geographical locations, disease classifications, disease severity classifications, and the like.
- Figure 12 includes a flow diagram of a process 1200 for estimating the reliability of a signal that includes values generated by an IMU. Note that the process 1200 is described in the context of values generated by an IMU for the purpose of illustration. The process 1200 may be similarly applicable to signals output by other types of sensors.
- a diagnostic platform can acquire a signal that includes values generated by an IMU over an interval of time (step 1201 ).
- the signal may include values generated over the course of a day, week, month, etc.
- the diagnostic platform is configured to apply a classifier to the signal to obtain a plurality of outputs, each of which is representative of a classification of a different one of a plurality of segments of the signal (step 1202).
- the diagnostic platform may apply an activity classifier to the signal in order to classify each segment as either ambulatory or non-ambulatory.
- the diagnostic platform may identify bouts of activity by joining consecutive ones of the plurality of segments that have the same classification (step 1203).
- the diagnostic platform may discover walking bouts by identifying instances where multiple ambulatory segments have been joined together.
- the diagnostic platform can then produce a plurality of measurements by deriving a separate measurement for each of the bouts of activity (step 1204).
- the diagnostic platform may derive measurements such as step count, cadence, bout duration, arm swing magnitude, and arm swing range of motion. In some embodiments, the diagnostic platform determines more than one of these measurements for each of the bouts of activity.
- the diagnostic platform can determine an optimal scope for aggregation (step 1205). As discussed above with reference to Figure 11 , the diagnostic platform can accomplish this by defining a plurality of candidate scopes corresponding to different durations and, for each of the plurality of candidate scopes, calculate an aggregated measurement within that candidate scope and then determine a reliability metric based on a comparison of the aggregated measurement to aggregated measurements calculated for adjacent candidate scopes. The diagnostic platform can plot the plurality of reliability metrics against the plurality of candidate scopes, as shown in Figure 7. Based on an analysis of the plot, the diagnostic platform can identify the optimal scope for aggregation.
- the diagnostic platform can aggregate the values included in the signal via a mathematical operator, such that a plurality of aggregated values are produced for a plurality of windows defined across the signal in accordance with the optimal scope (step 1206).
- mathematical operators include the max operator, min operator, and 95 th -percentile operator.
- the diagnostic platform can then populate the plurality of aggregated values into a data structure (e.g., a tabular data structure) such that the plurality of aggregated values are paired together (step 1207). Based on an analysis of the tabular data structure, the diagnostic platform can estimate the test-retest reliability of the signal and provide statistical inference of the estimate (step 1208).
- Figure 13 includes a block diagram of a processing system 1300 in which at least some operations described herein can be implemented.
- components of the processing system 1300 may be hosted on a computing device that includes a diagnostic platform.
- the processing system 1300 can include a processor 1302, main memory 1306, non-volatile memory 1310, network adapter 1312, video display 1318, input/output devices 1320, control device 1322 (e.g., a keyboard or pointing device such as a computer mouse or trackpad), drive unit 1324 including a storage medium 1326, and signal generation device 1330 that are communicatively connected to a bus 1316.
- the bus 1316 is illustrated as an abstraction that represents one or more physical buses or point-to-point connections that are connected by appropriate bridges, adapters, or controllers.
- the bus 1316 can include a system bus, a Peripheral Component Interconnect (“PCI”) bus or PCI-Express bus, a HyperTransport (“HT”) bus, an Industry Standard Architecture (“ISA”) bus, a Small Computer System Interface (“SCSI”) bus, a Universal Serial Bus (“USB”) data interface, an Inter- Integrated Circuit (“l 2 C”) bus, or a high-performance serial bus developed in accordance with Institute of Electrical and Electronics Engineers (“IEEE”) 1394.
- PCI Peripheral Component Interconnect
- HT HyperTransport
- ISA Industry Standard Architecture
- SCSI Small Computer System Interface
- USB Universal Serial Bus
- IEEE Inter- Integrated Circuit
- main memory 1306, non-volatile memory 1310, and storage medium 1326 are shown to be a single medium, the terms “machine-readable medium” and “storage medium” should be taken to include a single medium or multiple media (e.g., a centralized/distributed database and/or associated caches and servers) that store one or more sets of instructions 1328.
- the terms “machine-readable medium” and “storage medium” shall also be taken to include any medium that is capable of storing, encoding, or carrying a set of instructions for execution by the processing system 1300.
- routines executed to implement the embodiments of the disclosure can be implemented as part of an operating system or a specific application, component, program, object, module, or sequence of instructions (collectively referred to as “computer programs”).
- the computer programs typically comprise one or more instructions (e.g., instructions 1304, 1308, 1328) set at various times in various memory and storage devices in a computing device.
- the instruction(s) When read and executed by the processors 1302, the instruction(s) cause the processing system 1300 to perform operations to execute elements involving the various aspects of the present disclosure.
- machine- and computer-readable media include recordable-type media, such as volatile memory devices and non-volatile memory devices 1310, removable disks, hard disk drives, and optical disks (e.g., Compact Disk Read-Only Memory (“CD-ROMs”) and Digital Versatile Disks (“DVDs”)), and transmission-type media, such as digital and analog communication links.
- recordable-type media such as volatile memory devices and non-volatile memory devices 1310
- removable disks such as removable disks, hard disk drives, and optical disks (e.g., Compact Disk Read-Only Memory (“CD-ROMs”) and Digital Versatile Disks (“DVDs”)
- CD-ROMs Compact Disk Read-Only Memory
- DVDs Digital Versatile Disks
- the network adapter 1312 enables the processing system 1300 to mediate data in a network 1314 with an entity that is external to the processing system 1300 through any communication protocol supported by the processing system 1300 and the external entity.
- the network adapter 1312 can include a network adaptor card, a wireless network interface card, a router, an access point, a wireless router, a switch, a multilayer switch, a protocol converter, a gateway, a bridge, bridge router, a hub, a digital media receiver, a repeater, or any combination thereof.
Landscapes
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Public Health (AREA)
- Medical Informatics (AREA)
- General Health & Medical Sciences (AREA)
- Biomedical Technology (AREA)
- Pathology (AREA)
- Physics & Mathematics (AREA)
- Molecular Biology (AREA)
- Heart & Thoracic Surgery (AREA)
- Data Mining & Analysis (AREA)
- Surgery (AREA)
- Animal Behavior & Ethology (AREA)
- Biophysics (AREA)
- Primary Health Care (AREA)
- Veterinary Medicine (AREA)
- Epidemiology (AREA)
- Artificial Intelligence (AREA)
- Physiology (AREA)
- Databases & Information Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Mathematical Physics (AREA)
- Psychiatry (AREA)
- Signal Processing (AREA)
- Biodiversity & Conservation Biology (AREA)
- Fuzzy Systems (AREA)
- Evolutionary Computation (AREA)
- Dentistry (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Measuring And Recording Apparatus For Diagnosis (AREA)
- Measurement Of The Respiration, Hearing Ability, Form, And Blood Characteristics Of Living Organisms (AREA)
Abstract
Introduced here is a computer program and associated computer-implemented techniques for establishing the optimal scope for aggregating data collected in free-living scenarios and evaluating the test-retest reliability of the data. At a high level, the computer program can serve as an end-to-end system for measuring test-retest reliability of data that is collected in free-living settings, informing the appropriate level of aggregation of the data for a given digital measurement to determine accurate reliability.
Description
APPROACHES TO ESTABLISHING OPTIMAL AGGREGATION OF DATA IN SIGNALS GENERATED IN FREE-LIVING SCENARIOS AND DETERMINING RELIABILITY OF THOSE SIGNALS
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to US Provisional Application No. 63/476,112 filed on December 19, 2022, which is incorporated by reference herein in its entirety.
TECHNICAL FIELD
[0002] Various embodiments concern computer programs and associated computer-implemented techniques for determining the reliability of digital biomarkers derived from information collected in free-living settings.
BACKGROUND
[0003] The term “biomarker” - a portmanteau of “biological” and “marker” - is commonly used to refer to a measurable indicator of a physiological disease (also called a “physiological condition” or “physiological ailment”). Traditionally, biomarkers were evaluated through analysis of blood, urine, or soft tissue taken from a living body, generally for the purpose of predicting the onset of a physiological disease, monitoring the progression of a physiological disease, or determining the pharmacologic response to a therapeutic intervention for a physiological disease.
[0004] An emerging field of biomarkers relies on the analysis of measurements generated by sensors that monitor physiological aspects of living bodies. These biomarkers are commonly called “digital biomarkers,” since insights into the health of a given individual can be surfaced through analysis of measurements generated by sensors embedded in digital computing devices (or simply “computing devices”). Because computing devices with embedded sensors are now ubiquitous in society, digital biomarkers can serve as a reliable tool for advancing targeted treatment and guidance in health care.
BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Figure 1 illustrates a network environment that includes a diagnostic platform that is executed by a computing device.
[0006] Figure 2 illustrates an example of a computing device that is able to implement a diagnostic platform designed to establish digital biomarkers through analysis of signals and verify the reliability of those digital biomarkers.
[0007] Figure 3 depicts an example of a communication environment that includes a diagnostic platform that is configured to receive several types of data.
[0008] Figure 4 depicts another example of a communication environment that includes a diagnostic platform that is configured to obtain data from one or more sources.
[0009] Figure 5 includes a schematic diagram of an approach to establishing reliability of digital biomarkers determined through analysis of data collected during free-living scenarios.
[0010] Figure 6 illustrates how a given measurement (i.e., step count) can be aggregated across two different candidate aggregation scopes (i.e., one day and two days).
[0011] Figure 7 includes an example of a plot in which reliability values (here, intraclass correlation coefficients or “ICCs”) are plotted against candidate aggregation scopes.
[0012] Figure 8 includes an example of an aggregated value A given an optimal aggregation scope of three days.
[0013] Figures 9A-C illustrate several different methods that can be employed to generate paired aggregated values.
[0014] Figure 10 shows an example in which test-retest reliability values are calculated for different types of motor measurements based on signals output by sensors.
[0015] Figure 11 includes a flow diagram of a process for establishing an optimal scope for aggregation of measurements derived from data included in a signal generated over time.
[0016] Figure 12 includes a flow diagram of a process for estimating the reliability of a signal that includes values generated by an inertial measurement unit.
[0017] Figure 13 includes a block diagram of a processing system in which at least some operations described herein can be implemented.
[0018] Various embodiments are shown in the drawings for the purpose of illustration. However, those skilled in the art will recognize that alternative embodiments may be employed without departing from the principles of the present disclosure. Accordingly, while certain embodiments are shown in shown in the drawings, the technologies described herein are amenable to various modifications.
DETAILED DESCRIPTION
[0019] The development of digital biomarkers has led to significant developments in detecting and monitoring physiological diseases. As an example, digital biomarkers have impacted the field of neurology where better approaches to reliably and consistently tracking motor function in a noninvasive manner are needed. Digital biomarkers are quantifiable measures that are established through analysis of measurements collected by sensors. As further discussed below, these sensors are generally included in computing devices that are worn on, implanted in, or digested by a living body. Over time, each of these sensors collects measurements, and these data can be used to explain, influence, or predict health-related outcomes.
[0020] Digital biomarkers also allow for more targeted health care, as individualized digital biomarkers can be established and recorded to create personal baselines for health. Because of the important role that digital biomarkers can play in providing targeted health care, the measurements collected by the sensors must be reliable. There are no standard approaches for establishing test-retest reliability of measurements generated by sensors embedded in computing devices in free-living scenarios, however. This is likely due to the limited availability of sufficiently longitudinal data in studies or the lack of infrastructure for processing large amounts of data generated by sensors in free-living scenarios. The lack of a standard approach means that reliability of a continuous signal - from which digital biomarkers can be derived - is difficult to consistently verify.
[0021] Conventional tools for establishing test-retest reliability rely on random effects models that are computationally complex and costly not only during the construction stage but also the execution stage (also called the “implementation stage”). At a high level, a random effects model assumes that explanatory variables have fixed relationships with the response variable across all observations but that these fixed relationships may vary from one observation to
another. To solve a random effect model, an algorithm is employed that, through optimization, finds the maximum likelihood estimate. Execution of this optimization algorithm also tends to have high computational costs during the execution stage.
[0022] Introduced here is a computer program that is designed to establish the optimal scope for aggregating data collected in free-living scenarios and evaluate the test-retest reliability of the data, as well as associated computing systems (or simply “systems”) for implementing the same. As further discussed below, these data may include measurements generated by a sensor, questionnaire responses provided by individuals, and the like. At a high level, the computer program offers an end-to-end system for measuring test-retest reliability of digital measurements that are collected in free-living settings, informing the appropriate level of aggregation for a given measurement to determine accurate reliability.
[0023] Reliability is generally considered a quantitative measurement of data, indicating whether additional data of a similar nature will remain consistent. As an example, a metric that is indicative of reliability may be used to convey the likelihood that a measurement generated by a sensor that monitors a characteristic of an individual will be consistent with repeated measurements by the sensor. Reliability has historically been difficult to quantify in free-living scenarios, as there is greater variability in the data than more controlled scenarios, resulting in less consistency.
[0024] Free-living scenarios also pose a challenge in developing scalable approaches to analyzing large volumes of data, especially when collected at a high frequency, with the presence of noisiness and missingness. For example, data collected during one interval of time may be noisier than data collected during another interval of time, or data may simply not be collected during an interval of time. Approaches to appropriately aggregating data generated during
free-living scenarios are important, so that useful insights can be retrieved, derived, or inferred from the data.
[0025] As mentioned above, conventional tools for establishing test-retest reliability rely on random effects models that rely on optimization algorithms for finding maximum likelihood estimates. Conversely, the approach described herein may rely on a closed-form formula that is based on intraclass correlation as originally proposed by Ronald Fisher. Intraclass correlation - or the intraclass correlation coefficient (“ICC”) - is a descriptive statistic that can be used when quantitative measurements are made on units that are organized into groups. The ICC describes how strongly units in the same group resemble one another.
[0026] With respect to statistical inference, the computer program may default to a bootstrapping approach; however, the complexity of construction and execution of random effects models makes bootstrapping a less ideal candidate because of the computational cost of solving for the maximum likelihood estimate for each bootstrapped sample. As a result, statistical inference from random effects models may be based on asymptotic assumptions. Additionally, random effects models are not well suited for establishing test-retest reliability on longitudinal data, because those models overblend points in time that are further apart into calculation of the reliability metric. This can inflate within-individual variability (also called “intra-patient variability”), and therefore may unjustly penalize reliability estimates. Last but not least, conventional tools (e.g., the irr package for R Project) are generally incapable of handling data streams with missing entries. This makes it difficult, if not impossible to use such conventional tools to establish test-retest reliability of data collected during free-living scenarios, as collection tends to be inconsistent. As an example, if an individual opts not to wear a watch with sensors embedded thereon, data may not be available when the watch is not worn by the individual.
[0027] Embodiments may be described in the context of executable instructions for the purpose of illustration. However, those skilled in the art will
recognize that aspects of the technology could be implemented via hardware, firmware, or software. As an example, a computer program that is representative of a diagnostic platform may be executed by the processor of a computing device. The computer program may interface, directly or indirectly, with hardware, firmware, or other software implemented on the computing devices. For example, the computer program may obtain data that is collected, computed, or otherwise obtained by a sensor included in the computing device. Additionally or alternatively, the computer program may obtain data that is collected, computed, or otherwise obtained by a sensor included in another computing device. This data may be processed in accordance with the approaches described herein, so as to establish its reliability (and by extension, reliability of digital biomarkers that are based on this data).
Terminology
[0028] References in the present disclosure to “an embodiment” or “some embodiments” means that the feature, function, structure, or characteristic being described is included in at least one embodiment. Occurrences of such phrases do not necessarily refer to the same embodiment, nor do they necessarily refer to alternative embodiments that are mutually exclusive of one another.
[0029] The term “based on” is to be construed in an inclusive sense rather than an exclusive sense. That is, in the sense of “including but not limited to.” Thus, the term “based on” is intended to mean “based at least in part on” unless otherwise noted.
[0030] The terms “connected,” “coupled,” and variants thereof are intended to include any connection or coupling between two or more elements, either direct or indirect. The connection or coupling can be physical, logical, or a combination thereof. For example, elements may be electrically or communicatively connected to one another despite not sharing a physical connection.
[0031] The term “module” may refer broadly to software, firmware, hardware, or combinations thereof. Modules are typically functional components that
generate one or more outputs based on one or more inputs. A computer program may include or utilize one or more modules. For example, a computer program may utilize multiple modules that are responsible for completing different tasks, or a computer program may utilize a single module that is responsible for completing multiple tasks.
[0032] When used in reference to a list of items, the word “or” is intended to cover all of the following interpretations: any of the items in the list, all of the items in the list, and any combination of items in the list.
[0033] The terms "longitudinal signal" and “longitudinal data” may be used to refer to a series of values that are periodically generated over an interval of time rather than on an ad hoc basis. For example, a parameter may be measured by a sensor multiple times per second (e.g., at a sampling rate of 25, 50, or 100 hertz) in an ongoing manner, or a parameter may be measured once per minute for several hours. The length of time over which the series of values extend may depend on the nature of the parameter being measured. The term “parameter” is generally used to describe the unit of measure of the underlying signal. An example of a parameter is x-axis acceleration.
[0034] The terms “digital measurement” and “measurement” may be used to refer to data that are either collected directly from a computing device or derived from a raw signal output by a sensor included in a computing device. For example, a “measurement” may be a patient-reported outcome provided via a symptom survey, or a “measurement” may be step count as derived from values output by an accelerometer that is indicative of acceleration along one or more axes.
[0035] The term “digital biomarker” may be used to refer to a subset of measurements that are deemed fit for a diagnostic purpose and validated as diagnostically relevant. For example, step count may become a digital biomarker of disease progression (also called a “progression biomarker”) if it can be proven that step count is sensitive to worsening of mobility of the corresponding patient.
However, before this relationship is established, step count may simply be a “measurement.”
[0036] The term "free-living scenario" may be used to refer to a scenario in which an individual under observation is permitted to live her life as if not under observation. As an example, fitness tracking devices (also called “fitness trackers”) may monitor movement of individuals by measuring acceleration of the wrist in free-living scenarios. Observation in free-living scenarios is generally less prone to the intentional alteration of habits that commonly occurs in more controlled environments.
Overview of Diagnostic Platform
[0037] Figure 1 illustrates a network environment 100 that includes a diagnostic platform 102 that is executed by a computing device 104. An individual (also referred to as a “user”) can interact with the diagnostic platform 102 via interfaces 106. For example, a patient may be able to access an interface through which information regarding a physiological disease, such as digital biomarkers or analyses of digital markers, can be reviewed. As another example, a healthcare professional may be able to access an interface through which information regarding patients, such as digital biomarkers or analyses or digital markers, can be reviewed. Depending on the nature of the individual accessing the interfaces 106, the interfaces 106 may allow for the review of physiological data, examination of outputs produced by the diagnostic platform 102, and management of preferences. Some interfaces may be configured to facilitate interactions between patients and healthcare professionals, while other interfaces may be configured to serve as informative dashboards for patients or healthcare professionals.
[0038] The physiological data obtained by the diagnostic platform 102 could be associated with the individual accessing the interfaces 106 or some other person. For example, the interfaces 106 may enable a person diagnosed with a physiological disease to view her own physiological data. Alternatively, the
interfaces may enable an individual to view physiological data associated with another person. In such embodiments, the individual may be a healthcare professional who is responsible for monitoring, managing, or treating the other person. Examples of healthcare professionals include physicians, nurses, dietitians, and the like.
[0039] As shown in Figure 1 , the diagnostic platform 102 can reside in a network environment 100. Thus, the computing device 104 on which the diagnostic platform 102 resides can be connected to one or more networks 108A-B. Depending on its nature, the computing device 104 could be connected to a personal area network (“PAN”), local area network (“LAN”), wide area network (“WAN”), metropolitan area network (“MAN”), or cellular network. For example, if the computing device 104 is a computer server, then the computing device 104 may be accessible to users via respective mobile phones that are connected to the Internet via LANs. The physiological data to be examined by the diagnostic platform 102 may be generated by the respective mobile phones (e.g., sensors included in the respective mobile phones) or acquired by the respective mobile phones.
[0040] Additionally or alternatively, the computing device 104 may be connected to one or more other computing devices over a short-range wireless connectivity technology, such as Bluetooth®, Near Field Communication (“NFC”), Wi-Fi® Direct (also referred to as “Wi-Fi P2P”), and the like. As an example, the diagnostic platform 102 could be embodied as a mobile application that is executed by a mobile phone. In such embodiments, the mobile phone may be communicative connected - via a wireless communication channel - to a source from which to acquire physiological data. The source could be a watch, fitness tracker, or another wearable computing device, for example. The physiological could alternatively be obtained from another computer program executing on the mobile phone. For example, the physiological data could instead be acquired from another mobile application executing on the mobile phone or the operating system of the mobile phone.
[0041] The interfaces 106 may be accessible via a web browser, desktop application, mobile application, or another form of computer program. For example, a patient may be able to access interfaces through which information regarding her own health, such as digital biomarkers or analyses of digital biomarkers, is provided by a mobile application executing on a mobile phone. As another example, a healthcare professional may be able to access interfaces through which information regarding one or more patients can be reviewed via a web browser. Accordingly, the interfaces 106 generated by the diagnostic platform 102 may be accessible on various computing devices, including mobile phones, tablet computers, desktop computers, and the like.
[0042] Generally, the diagnostic platform 102 is executed - at least partially - by a cloud computing service operated by, for example, Amazon Web Services®, Google Cloud Platform™, or Microsoft Azure®. Thus, the computing device 104 may be representative of a computer server that is part of a server system 1 10. Often, the server system 110 is comprised of multiple computer servers. These computer servers can include different types of data (e.g., physiological data and information regarding patients, such as name, demographic information, disease classification, etc.), algorithms for processing incoming data, and other assets. Those skilled in the art will recognize that these data could also be distributed among the server system 1 10 and one or more computing devices. As an example, data that is input by, or related to, patients may be stored on, and processed by, their own computing devices for security or privacy purposes.
[0043] Components of the diagnostic platform 102 could also be hosted locally. That is, part of the diagnostic platform 102 may reside on the computing device used to access one of the interfaces 106. For example, the diagnostic platform 102 may be embodied as a mobile application executing on a mobile phone as mentioned above. Note, however, that the mobile application may be communicatively connected to the server system 1 10 on which other components of the diagnostic platform 102 are hosted.
[0044] Figure 2 illustrates an example of a computing device 200 that is able to implement a diagnostic platform 212 designed to establish digital biomarkers through analysis of signals and verify the reliability of those digital biomarkers. As shown in Figure 2, the computing device 200 can include a processor 202, memory 204, display mechanism 206, communication module 208 and sensor suite 210. Each of these components is discussed in greater detail below.
[0045] Those skilled in the art will recognize that different combinations of these components may be present depending on the nature of the computing device 200. For example, if the computing device 200 is a computer server that is part of a server system (e.g., server system 1 10 of Figure 1 ), then the computing device 200 may not include the display mechanism 206 or sensor suite 210. Conversely, if the computing device 200 is a mobile phone, then the computing device 200 can include the display mechanism 206 and sensor suite 210.
[0046] As further discussed below, the computing device 200 on which the diagnostic platform 212 resides may not include sensors in some embodiments. In such embodiments, the data examined by the diagnostic platform 212 could instead be generated by one or more sensors 222A-N that are external to the computing device 200. For example, the computing device 200 may be a mobile phone, and the sensors 222A-N may be included in a watch or fitness tracker that is communicatively connected to the computing device 200.
[0047] The processor 202 can have generic characteristics similar to general- purpose processors, or the processor 202 may be an application-specific integrated circuit (“ASIC”) that provides control functions to the computing device 200. As shown in Figure 2, the processor 202 can be coupled to all components of the computing device 200, either directly or indirectly, for communication purposes.
[0048] The memory 204 can be comprised of any suitable type of storage medium, such as static random-access memory (“SRAM”), dynamic randomaccess memory (“DRAM”), electrically erasable programmable read-only memory
(“EEPROM”), flash memory, or registers. In addition to storing instructions that can be executed by the processor 202, the memory 204 can also store data generated by the processor 202 (e.g., when executing the modules of the diagnostic platform 212). Note that the memory 204 is merely an abstract representation of a storage environment. The memory 204 could be comprised of actual integrated circuits (also called “chips”).
[0049] The display mechanism 206 can be any mechanism that is operable to visually convey information to a user. For example, the display mechanism 206 can be a panel that includes light-emitting diodes (“LEDs”), organic LEDs, liquid crystal elements, or electrophoretic elements. As further discussed below, outputs produced by the diagnostic module 212 (e.g., through execution of its modules) can be posted to the display mechanism 206 for review by a user of the computing device 200.
[0050] The communication module 208 may be responsible for managing communications external to the computing device 200. The communication module 208 can be wireless communication circuitry that is able to establish wireless communication channels with other computing devices. Examples of wireless communication circuitry include 2.4 gigahertz (“GHz”) and 5 GHz chipsets compatible with Institute of Electrical and Electronics Engineers (“IEEE”) 802.11 - also referred to as “Wi-Fi chipsets.” Alternatively, the communication module 208 may be representative of a chipset configured for Bluetooth, NFC, and the like. Some computing devices - like mobile phones, tablet computers, and the like - are able to wirelessly communicate via separate channels, while other computing devices - like watches and fitness trackers - tend to wirelessly communicate via a single channel. Accordingly, the communication module 208 may be one of multiple communication modules implemented in the computing device 200, or the communication module 208 may be the only communication module implemented in the computing device 200.
[0051] The nature, number, and type of communication channels established by the computing device 200 - and more specifically, the communication module 208 - can depend on (i) the sources from which data is received by the diagnostic platform 212 and (ii) the destinations to which data is transmitted by the diagnostic platform 212. Assume, for example, that the diagnostic platform 212 resides on a mobile phone in the form of a mobile application. In such embodiments, the communication module 208 can communicate with sensors 222A-N external to the computing device 200 from which to obtain data. Moreover, the communication module 208 may communicate with a server system (e.g., server system 1 10 of Figure 1 ) to which analyses of the data - or the data itself - are transmitted.
[0052] Often, various sensors are implemented in the computing device 200. Collectively, these sensors may be referred to as the “sensor suite” 210 of the computing device 200. For example, the computing device 200 may include a motion sensor whose output is indicative of motion of the computing device 200 as a whole. Examples of motion sensors include accelerometers and gyroscopes. In some embodiments, the motion sensor is implemented in an inertial measurement unit (“IMU”) that measures the force, angular rate, or orientation of the computing device 200. The IMU may accomplish this through the use of one or more accelerometers, one or more gyroscopes, one or more magnetometers, or any combination thereof. As specific examples, the IMU could be a 6-axis IMU that draws low current and therefore is suitable for “always-on” applications in battery-driven computing devices, or the IMU could be a 3-axis IMU that includes logic-level shifting circuitry that can readily interface with a microcontroller. As another example, the computing device 200 may include an ambient light sensor whose output is indicative of the amount of light in the ambient environment.
[0053] As mentioned above, data could also be acquired from sensors 222A- N that are external to the computing device 200. These sensors 222A-N could be included in another computing device, such as a watch or fitness tracker, that is
directly connected to the computing device 200. Alternatively, these sensors 222A-N could be discrete sensing units that are directly or indirectly connected to the computing device 200. For example, sensor 222A may be a pulse oximeter that monitors the oxygen saturation of the user of the computing device 200 or another person, for the purpose of creating a photoplethysmogram (“PPG”). To monitor the oxygen saturation, sensor 222A may be placed on a thin part of a living body, usually a fingertip or earlobe, and then pass two wavelengths of light through that body part toward a photodetector. The photodetector can measure the changing absorbance at each wavelength, allowing sensor 222A to determine the absorbances due to the pulsing of arterial blood through that body part.
[0054] For convenience, the diagnostic platform 212 is referred to as a computer program that resides within the memory 204. However, the diagnostic platform 212 could be comprised of software, firmware, or hardware that is implemented in, or accessible to, the computing device 200. In accordance with embodiments described herein, the diagnostic platform 212 can include a processing module 214, computation module 216, aggregation module 218, and graphical user interface (“GUI”) module 220. These modules could be integral parts of the diagnostic platform 212, or these modules could be logically separate from the diagnostic platform 212 but operate “alongside” it. Together, these modules enable the diagnostic platform 212 to establish the reliability of data to be used to gain insights into the health of a given individual. As mentioned above, the given individual could be the user of the computing device 200 or another person.
[0055] The processing module 214 can process data obtained by the diagnostic platform 212 into a format that is suitable for the other modules. For example, the processing module 214 can apply operations to sensor data obtained from the sensor suite 210 or sensors 222A-N in preparation for analysis by the other modules of the diagnostic platform 212. For example, the processing module 214 can filter or alter the sensor data, such that the sensor data can be
more readily analyzed. As another example, the processing module 214 may parse the sensor data in order to temporally align the dataset obtained from each source. Accordingly, the processing module 214 may be responsible for ensuring that the appropriate sensor data is accessible to the other modules of the diagnostic platform 212.
[0056] The computation module 216 (also called a “computation system”) may be responsible for deriving different measurements based on the processed data received from the processing module 214. As further discussed below, these different measurements may depend on the nature of the insights to be surfaced by the diagnostic platform 210. For example, these different measurements may be representative of walking-related measurements, running-related measurements, sleeping-related measurements, eating-related measurements, stress-related measurements, or any combination thereof. The computation module 216 can compute, infer, or otherwise establish measurements for different segments of processed data. These segments may be referred to as “windows” of processed data.
[0057] The aggregation module 218 (also called an “aggregation system”) may implement an aggregation engine that, in operation, implements two programs in sequence. As further discussed below, the first program may provide a method for programmatically determining the appropriate scope of aggregation for the processed data output by the processing module 412, while the second program may implement one of multiple methods for pairing the data for calculation of reliability values.
[0058] The GUI module 220 may be responsible for generating interfaces that are viewable on the display mechanism 206. Various types of information can be presented on these interfaces. For example, digital biomarkers that are calculated, derived, or otherwise obtained by the computation module 216 may be presented on an interface for display to the user. Similarly, analyses of the digital biomarkers may be presented on the interface for display to the user.
[0059] Figure 3 depicts an example of a communication environment 300 that includes a diagnostic platform 302 that is configured to receive several types of data. Here, for example, the diagnostic platform 302 receives response data 304, first sensor data 306 that is generated by a first sensor (e.g., sensor 222A of Figure 2), and second sensor data 308 that is generated by a second sensor (e.g., sensor 222B of Figure 2). Those skilled in the art will recognize that these data have been selected for the purpose of illustration. Other types of data, such as treatment data (e.g., including indicators of physiological disease, treatment regimens, etc.), could also be obtained by the diagnostic platform 302.
[0060] These data may be obtained from multiple sources.
[0061] Consider the response data 304, for example. The response data 304 could be obtained directly on the computing device on which the diagnostic platform 302 is executing. For example, if the diagnostic platform 302 is implemented on a mobile phone in the form of a mobile application, the response data 304 may be representative of input provided by an individual through interfaces generated by the mobile application. The input may be representative of responses to queries included in a questionnaire that is designed to elicit responses indicative of health. Alternatively, the response data 304 could be obtained from another computing device. Referring again to the aforementioned example in which the diagnostic platform 302 is implemented as a mobile application executing on a mobile phone, the mobile phone could obtain the response data 304 from a server system (e.g., server system 110 of Figure 1 ). The server system may manage a datastore in which responses provided by various individuals to questionnaires are documents.
[0062] As mentioned above, sensor data could be obtained from sensors included in the computing device that is responsible for executing the diagnostic platform 302, or sensor data could be obtained from sensors that are external to the computing device that is responsible for executing the diagnostic platform 302. Thus, the first and second sensor data 306, 308 may be generated by
sensors included in the computing device on which the diagnostic platform 302 is executing, or the first and second sensor data 306, 308 may be generated by sensors that are external to the computing device on which the diagnostic platform 302 is executing. As an example, the first sensor data 306 may be generated by an IMU implemented in the computing device on which the diagnostic platform 302 is executing, while the second sensor data 308 may be generated by a pulse oximeter that is communicatively connected - either directly or indirectly - to the computing device on which the diagnostic platform 302 is executing.
[0063] Figure 4 depicts another example of a communication environment 400 that includes a diagnostic platform 402 that is configured to obtain data from one or more sources. Here, the diagnostic platform 402 may obtain data from a mobile phone 404, watch 406, laptop computer 408, or server system 410 (collectively referred to as the “networked devices”). For example, the diagnostic platform 402 may obtain response data (e.g., response data 304 of Figure 3) from the laptop computer 408 or server system 410. As another example, the diagnostic platform 402 may obtain sensor data (e.g., sensor data 306, 308 of Figure 3) from the mobile phone 404 or watch 406.
[0064] The networked devices can be connected to the diagnostic platform 402 via one or more networks. These networks can include PANs, LANs, WANs, MANs, cellular networks, the Internet, etc. Additionally or alternatively, the networked devices may communicate with one another over a short-range wireless connectivity technology. For example, if the diagnostic platform 402 resides on the mobile phone 404 in the form of a mobile application, data may be obtained from the watch 406 over a Bluetooth communication channel while data may be obtained from the server system 410 over the Internet via a Wi-Fi communication channel.
[0065] Embodiments of the communication environment 400 may include a subset of the networked devices. For example, the communication environment
400 may include a diagnostic platform 402 that obtains, in real time, data from the mobile phone 404 and watch 406 as that data is generated over the course of a free-living scenario. Additional data could be obtained from the server system 410 on a periodic basis (e.g., daily or weekly).
Approaches to Establishing Data Reliability
[0066] Figure 5 includes a schematic diagram of an approach to establishing reliability of digital biomarkers determined through analysis of data collected during free-living scenarios. At a high level, the approach includes five steps, each of which is discussed in greater detail below.
A. _ Signal Collection
[0067] Initially, a diagnostic platform can collect signals from one or more sources (step 501 ). For example, the diagnostic platform may collect, from a server system (e.g., server system 110 of Figure 1 ), a signal that is representative of responses to a questionnaire that is accessible via interfaces generated by the diagnostic platform. As another example, the diagnostic platform may collect, from a watch, signals that are representative of measurements generated by sensors included in the watch.
[0068] In embodiments where the signal collected by the diagnostic platform is representative of a series of measurements generated by a sensor in temporal order, the signal is preferably a longitudinal signal with values generated over an interval of time. For example, the signal could include values that are generated at a predetermined cadence by an IMll. As another example, the signal could include values that are generated at a predetermined cadence by a pulse oximeter. The duration of the signal - and the cadence of its values - may depend on the nature of the sensor.
[0069] As mentioned above, data collected by the diagnostic platform may not be limited to “raw signals” generated by sensors. In addition to these raw signals (e.g., temporal series of IMll values, PPG signals, etc.), the diagnostic platform
could collect patient-reported outcomes that are provided, for example, through interfaces generated by the diagnostic platform. For example, the diagnostic platform may surface, through an interface, a questionnaire that solicits responses to questions from patients whose health is being monitored. Examples of patient-reported outcomes could be collected using “open” queries (e.g., where each patient is permitted to answer as she sees fit) or “closed” queries. Closed queries may be posed using a rating scale such as Likert scale, so that responses can be readily scaled. Examples of topics that are generally well suited for open queries vital signs and symptoms being experienced, while examples of topics that are generally well suited for closed queries include current pain level and current ability to perform regular activities.
[0070] The signal is preferably collected by the diagnostic platform - or at least values are generated by the source - at a relatively high frequency. How frequently data is collected may depend on not only its type but also its intended application (e.g., the digital biomarkers that are based on such data). For example, data may be collected every second by sensors included in a watch or fitness tracker, whereas patient-reported outcomes may be collected on a daily or weekly basis (e.g., via interfaces generated by the diagnostic platform).
[0071] The data that is represented by, and conveyed via, the signal collected by the diagnostic platform can be stored in a data structure by the diagnostic platform. For example, the data may be populated into one or more data structures (e.g., tables) that are accessible to the diagnostic platform.
B. _ Data Examination
[0072] As mentioned above, the diagnostic platform can include a data processing component - namely, a processing module (e.g., processing module 214 of Figure 2) - that may be responsible for applying a set of preprocessing operations to the signal. Said another way, the processing module may be responsible for examining the data included in the signal (step 502).
[0073] The preprocessing operations that are applied to the signal may depend on the type of data contained therein. Assume, for example, that the signal includes measurements generated by a sensor. In such a scenario, the processing module can apply standard preprocessing operations such as resampling, bias removal, and noise removal using a low-pass filter or high-pass filter. In the event that the signal includes data that is representative of patient- reported outcomes, the processing module may code the data into a numeric format and then filter the data (e.g., to remove outlier answers, non-responsive answers, incomprehensible answers, etc.). Moreover, the processing module may perform a binning operation, in which the numerically coded data is sorted into categories or “bins.” As an example, patient-reported outcomes may be sorted into different age ranges (e.g., 0-19, 18-30, 30-45, etc.), geographical locations (e.g., by country, region, state, or county), disease classification (e.g., critical, severe, moderate, and mild), races (e.g., White, Asian, Hispanic, Black, American Indian, Native Hawaiian and Other Pacific Islander), gender (e.g., male, female, and non-binary), and the like.
C. _ Measurement Derivation
[0074] The preprocessed data can then be provided to a data computing component - namely, a computation module (e.g., computation module 216 of Figure 2) - that may be responsible for deriving different measurements based on the preprocessed data (step 503). Note that for some patient-reported outcomes, such as vital signs and symptoms experienced, no further derivation may be necessary.
[0075] The computation module can apply heuristics and/or machine-learning models (or simply “models”) to derive the measurements using the preprocessed data. For example, to derive a set of walking-related measurements, the computation module may apply one or more activity classifiers can be applied to windows of preprocessed data that includes values generated by an IMU, such that each window is classified as either ambulatory or non-ambulatory. A similar
approach could be used to classify each window as sleeping or non-sleeping, stressed or non-stressed, etc. Each window may be representative of a segment of preprocessed data of predetermined length (e.g., 5, 10, or 20 seconds). Meanwhile, each activity classifier may be representative of a model that when applied to a window of preprocessed data, classifies that window based on an analysis of the preprocessed data contained therein. The computation module can then identify walking bouts by joining consecutive ambulatory windows and then labeling ambulatory windows of sufficient length (e.g., 10, 20, or 40 seconds) as walking bouts. For each walking bout, the computation module can derive one or more measurements. Examples of such measurements include step count, cadence, bout duration, arm swing magnitude, arm swing velocity, arm swing acceleration, and arm swing range of motion.
[0076] The nature and number of measurements derived from the preprocessed data will depend on various factors, including the nature of the preprocessed data and intended application of measurements derived therefrom. Consider, for example, walking-related measurements versus sleeping-related measurements. To compute these measurements, the computation module may rely on the same preprocessed data, namely, values generated by an IMU. However, the measurements may differ. Examples of walking-related measurements include step count and cadence as discussed above, while examples of sleeping-related measurements include total sleep duration and wake count.
D. _ Aggregation
[0077] Thereafter, another data computing component - namely, an aggregation module (e.g., aggregation module 218 of Figure 2) - can perform aggregation (step 504). As shown in Figure 5, the aggregation module may implement an analytics engine 500 in order to perform aggregation. At a high level, the analytics engine 500 (or simply “engine”) is representative of the core logic of the aggregation module that, in operation, takes an input (e.g., the
preprocessed data), performs at least one operation, and then produces an output (e.g., an indication of an optimal window for aggregation). More specifically, the analytics engine 500 may be responsible for exploring optimal points of aggregation scope and conducting reliability assessments.
[0078] The analytics engine 500 can include two programs that, in operation, are performed in sequence. The first program may provide a method for programmatically determining the appropriate scope of aggregation for the data included in the signal collected as input. As further discussed below, the first program can determine the optimal window scope for aggregation based on analysis of a plot. The second program may provide multiple methods for pairing the data for calculation of reliability values, as further discussed below.
L _ Determination of Optimal Aggregation Scope
[0079] For longitudinal measurements that are calculated from an aggregated ensemble of data, the aggregation scope can be defined such that it contains a sufficient number of datapoints to enable a stable measure of the underlying computing device, while also allowing enough flexibility to accurately understand change. For the purpose of illustration, the aggregation scope may be described as the number of datapoints per patient, and each datapoint may be representative of a periodic (e.g., every second, minute, hour, day, or week) measurement collected for that patient. In practice, aggregation scopes can - and often do - correspond to intervals of time having different lengths. Assume, for example, that the aggregation scope is defined as 20. For a sensor that generates measurements every minute, the aggregation scope may correspond to a 20-minute interval of time, while for a sensor that generates measurements hourly or daily, the aggregation scope may correspond to a 20-hour or 20-day interval of time, respectively.
[0080] As mentioned above, the first program may provide a method for programmatically determining the optimal aggregation scope. As a brief summary, the first program can initially define a range of candidate aggregation
scopes. Generally, these candidate aggregation scopes correspond to, or are based on, the frequency of the data included in the signal received by the diagnostic platform as input. For example, the set of candidate aggregation scopes may be defined as 1 , 2, 5, 10, and 20 days. As another example, the set of candidate aggregation scopes may be defined as 1 , 2, 4, 8, 16, and 24 minutes. A set of candidate aggregation scopes is normally defined such that even the shortest candidate aggregation scope includes at least a predetermined number of datapoints.
[0081] For each candidate aggregation scope, an aggregated measurement can be calculated within the candidate scope timeframe and then compared to adjacent aggregation windows. Assume, for example, that the set of candidate aggregation scopes is defined as 1 , 2, 5, 10, and 20 days. For each candidate aggregation scope, a given measurement (e.g., step count) can be aggregated across the corresponding aggregation windows. Therefore, the analytics engine 500 may produce, as output, a set of data structures, each of which corresponds to a different one of the set of candidate aggregation scopes. Each data structure may include a plurality of values that are representative of the given measurement as computed across the signal collected as input on a rolling basis. Examples of these data structures are provided in Figure 6. Specifically, Figure 6 illustrates how a given measurement (i.e., step count) can be aggregated across two different candidate aggregation scopes (i.e., one day and two days).
[0082] Then, reliability values can be calculated based on the aggregated measures, such that each candidate aggregation scope is associated with a single reliability value. The number of aggregated measurements may depend on the number of windows that can be formed in accordance with the corresponding candidate aggregation scope. In Figure 6, for example, step count is determined using non-overlapping windows in accordance with the approach further described below with reference to Figure 9A. Generally, as the candidate aggregation scope becomes larger, the number of windows will decrease - and therefore, the number of aggregated measurements will decrease. However, only
one reliability value may be calculated for each candidate aggregation scope. Thus, a single reliability value may be produced for each candidate aggregation scope of a given measurement. The reliability values may be ICCs, for example.
[0083] The reliability values can then be plotted against the candidate aggregation scopes to better understand the increase in reliability as a function of aggregation scope. Figure 7 includes an example of a plot in which reliability values (here, ICCs) are plotted against candidate aggregation scopes. In Figure 7, reliability values are plotted along the y-axis while different candidate aggregation scopes are plotted along the x-axis, and each line visually illustrates the change in reliability value of a measurement (e.g., step count) along different aggregation scopes. With the plot, the analytics engine 500 can identify the minimum acceptable aggregation scope for the measure, which will allow for maximization of the measure’s ability to resolve change. As an example, the analytics engine 500 may examine the plot shown in Figure 7 to find the number of days in which the change in reliability value is smaller than a threshold (e.g., 0.05). iL _ Longitudinal Aggregation Based On Optimal Aggregation Scope
[0084] After the optimal aggregation scope is identified by the first program in step 504, the second program can aggregate the data in accordance with the optimal aggregation scope via different mathematical operators in step 505. Examples of mathematical operators include the max operator, min operator, Xth- percentile operator (where Xis an integer that ranges from 0 to 100, such as 75, 95, etc.), and the like. Such action will result in a set of aggregated values within each window defined in accordance with the optimal aggregation scope. Figure 8 includes an example of an aggregated value A given an optimal aggregation scope of three days. Referring to Figure 8 as an example, if the mathematical operator is a max operator, then A will be the maximum value of a measurement within three days’ worth of data.
[0085] Given a mathematical operator for aggregation, all aggregated values can be calculated for all possible three-day windows, labeled as A, B, C, and D in Figure 9A. A data structure of paired aggregated values can then be created in preparation for calculating reliability. The data structure may be a table, as shown in Figure 9A. There are several different methods that can be employed by the second program to generate the paired aggregated values.
[0086] In a first method, all of the aggregated values are calculated based on non-overlapping windows, as shown in Figure 9A. There may be a max-distance parameter that controls whether two aggregated values should be paired. For example, since A and C are more than 14 days apart and the max-distance parameter is 14 days, A and C are not paired in the table of paired aggregated values. Only A and B are paired and C and D are paired. One advantage of this first method is that the second program can avoid deflating the reliability values. However, one disadvantage is that the point estimate may be less accurate when the signal is “short” and doesn’t include a large number of datapoints.
[0087] In a second method, all of the aggregated values are calculated based on overlapping windows, as shown in Figure 9B. In addition to A, A' and A" can be calculated by shifting the beginning of the window. Here, for example, A' and A" are calculated by shifting the starting day. A table of paired aggregated values can be created by having all of the original values paired and shifted values paired (e.g., A-B, A'- B', B'- C , etc.). One advantage of this second method is the ability to stabilize the point estimate, while one disadvantage is that statistical inference is impacted and confidence interval (“Cl”) calculation is affected.
[0088] In a third method, all of the aggregated values are calculated based on non-overlapping windows, as shown in Figure 9C. There may be a max-gap parameter that controls how many pairs should be created within a given gap size. For example, if the max-gap parameter is set for 14 days, then all of the aggregated values within the gap (i.e., 14 days) will be paired, resulting in a table of paired aggregated values of A-B, A- C, A-D, B- C, B-D, etc. One advantage of
this third method is the ability to stabilize the point estimate, while one disadvantage is that implementation can be computationally complicated, with the Cis potentially being smaller than reality.
[0089] The choice of which method to use may be determined based on the intended application. Thus, the aggregation module may select one of the aforementioned methods based on the intended application of the aggregated values. The choice of which method to use may depend on the interest of the user. Accordingly, the user may specify which method to use, for example, through an interface generated by the diagnostic platform.
E. _ Reliability Computation
[0090] Thereafter, the aggregation module can estimate the test-retest reliability of the measurements derived by the computation module. For example, Fisher’s original formulation of ICC may be used to estimate test-retest reliability as set forth below in Eq. 1 -3. To compute a given ICC, the data may need to be organized into a pair of vectors ( xn l, xn 2) having columnar form, as shown in Figures 9A-C. These vectors can be constructed as set forth above.
[0091] For each reliability value, both the point estimate and Cl of the ICC may be provided. The point estimate can be obtained by computing the ICC from the observed data, while the Cl can be computed using a bootstrapping method. Specifically, the aggregation module can resample the observed data with replacement being performed m times, where m is an integer value (e.g., between 100 and 2,500, and preferably ~1 ,000). For each set of resampled data, the value of the ICC for that given sample can be obtained, and in total, m
number of ICCs may be obtained. A (100 - a) percent Cl can be obtained by taking the a/2 and (100 - a/2) percentiles of the resultant m ICCs. Figure 10 shows an example in which test-retest reliability values are calculated for different types of motor measurements based on digital measures derived from signals output by sensors of a watch.
F. _ Technology Benefits
[0092] In operation, the aggregation module may be able to determine different ways of aggregation for longitudinal data generated during free-living scenarios, as discussed above. The ability to collect data at high frequency (e.g., multiple datapoints per second, minute, or hour) while offering the appropriate level of aggregation can lead to improved reliability in measurements that are computed based on the data. Thus, the framework described above can inform the appropriate level of aggregation for a given measurement to determine or maximize reliability.
Methodologies for Establishing Optimal Aggregation Scope
[0093] Several approaches to establishing the optimal aggregation scope are set forth below. These approaches are best understood when read in conjunction with the disclosure corresponding to Figures 5-10.
[0094] Figure 11 includes a flow diagram of a process 1 100 for establishing an optimal scope for aggregation of measurements derived from data included in a signal generated over time. Initially, a diagnostic platform can obtain a plurality of measurements that are derived via analysis of the signal (step 1101 ). Each of the plurality of measurements may be associated with a corresponding one of a plurality of segments of the signal. For example, the plurality of measurements may be computed, derived, or otherwise established by a computation module as discussed above. The nature of each measurement may depend on the nature of the data. For example, if the data is representative of discrete motion values generated by an IMU, then the measurement could be step count or arm swing magnitude. As another example, if the data is representative of discrete
absorbance values generated by a pulse oximeter, then the measurement could be heart rate or blood oxygen level.
[0095] Thereafter, the diagnostic platform can define a plurality of candidate scopes for aggregation (step 1102). Each of the plurality of candidate scopes corresponds to a different duration. For example, the plurality of candidate scopes may be 1 , 2, 5, 10, and 20 days, or the plurality of candidate scopes may be 1 , 2, 4, 8, 16, and 24 minutes. The durations of the candidate scopes are normally defined such that even the shortest candidate scope includes at least a predetermined number of datapoints. Therefore, the durations of the plurality of candidate scopes may be defined based on the frequency of datapoints in the signal. Note that the number of candidate scopes can vary.
[0096] Generally, it is desirable to have enough candidate scopes that a sufficiently large range of durations is covered. Thus, the plurality of candidate scopes may include (i) a shortest candidate scope that is defined such that a sufficient number of datapoints (e.g., the predetermined number) are included to enable stable measure and (ii) a longest candidate scope that is defined such that change over time can be understood. If the shortest candidate scope is too short, the aggregated measurements may be too variable (e.g., because some scopes will not include sufficient data). If the longest candidate scope is too long, then it becomes difficult to distinguish between aggregated measurements. The number of candidate scopes between the shortest and longest candidate scopes may vary. For example, the diagnostic platform may be programmed to define at least a predetermined number of candidate scopes (e.g., 2, 3, 5) between the shortest and longest candidate scopes, or the diagnostic platform may be programmed to define intervening candidate scopes at a predetermined cadence (e.g., that is determined based on the shortest candidate scope or longest candidate scope).
[0097] For each of the plurality of candidate scopes, the diagnostic platform can calculate, based on the plurality of measurements, an aggregated
measurement within that candidate scope (step 1103). Said another way, the diagnostic platform can calculate an aggregated measurement for each of the plurality of candidate scopes, such that a plurality of aggregated measurements are calculated. For each of the plurality of candidate scopes, the diagnostic platform can also determine a reliability metric based on the aggregated measurement (step 1 104). Specifically, the diagnostic platform may determine, for each of the plurality of candidate scopes, a reliability metric based on a comparison of the aggregated measurement to aggregated measurements calculated for the next highest candidate scope and/or next lowest candidate scope. For example, the diagnostic platform may compute an ICC for each of the plurality of candidate scopes. At a high level, each ICC may indicate how strongly the corresponding candidate scope resembles nearby candidate scopes.
[0098] Thereafter, the diagnostic platform can identify an optimal scope from among the plurality of candidate scopes based on an analysis of the plurality of reliability metrics determined for the plurality of candidate scopes (step 1 105). For example, the diagnostic platform may plot the plurality of reliability metrics against the plurality of candidate scopes as shown in Figure 7, and then the diagnostic platform may determine the optimal aggregation scope (and therefore, window size) based on an analysis of the plot. For example, using the plot, the diagnostic platform may be able to identify the minimum acceptable scope (which is approximately 14 days in Figure 7 because the improvement of the daily test- retest value is smaller than a user-defined cutoff value, a = 0.01, after 14 days), for example, by finding the duration for which the change in reliability metric is smaller than a threshold.
[0099] Note that the present disclosure concerns an analytical approach to identifying the optimal aggregation scope. In practice, identifying the optimal aggregation scope may involve consultation with partners (e.g., pharmaceutical manufacturers, healthcare systems, insurers) to establish the windows size of interest. Assume, for example, that a partner is interested in weekly measurement reliability, the diagnostic platform can provide the appropriate
weekly measurement. Alternatively, a partner may be interested in conducting aggregation within a window size of 14 days. In this scenario, the diagnostic may use a plot - similar to the one shown in Figure 7 - to infer (i) what reliability this aggregation scope will provide and (ii) whether an inflection point is reached (and therefore, whether a larger window size would help improve reliability).
[00100] In other words, there may be two approaches to determining the final aggregation scope (and therefore, window size). First, a qualitative approach in which a partner influences (e.g., chooses) selection of a preferred aggregation scope and then the diagnostic platform uses the plot to determine whether the preferred aggregation scope is a good choice. Second, a quantitative approach in which the diagnostic platform simply uses the graph and inflection point to determine the final aggregation scope.
[00101] After the optimal scope is established, additional steps could be performed. For example, the diagnostic platform may aggregate, via a mathematical operator, data included in the signal, such that a plurality of aggregated values are produced for a plurality of windows defined across the signal in accordance with the optimal scope (step 1 106). Examples of mathematical operators include the max operator, min operator, and 95th- percentile operator. Note that in some embodiments, the mathematical operator is one of multiple mathematical operators that are applied to the data included in the signal.
[00102] Then, the diagnostic platform can populate the plurality of aggregated values into a data structure (step 1 107). Specifically, the diagnostic platform may populate the plurality of aggregated values into a tabular data structure, such that the plurality of aggregated values are paired together. There are several different methods that could be employed by the diagnostic platform to generate the paired aggregated values, as further discussed above with reference to Figures 9A-C.
[00103] Moreover, the diagnostic platform may estimate test-retest reliability of the signal based on an analysis of the data structure and provide statistical inference of the estimate (step 1 108). For example, Fisher’s original formulation of ICC may be used to estimate the test-retest reliability as set forth above in Eq. 1 -3. For each reliability value, the diagnostic platform may establish the point estimate and/or Cl of the ICC. Accordingly, estimation of test-retest reliability may be based on point estimates of the ICCs, Cis of the ICCs, or a combination thereof. This estimate of the test-retest reliability can be used in various ways. For example, the diagnostic platform may determine, based on the estimate of the test-retest reliability, reliability of a digital biomarker that is computed using the signal.
[00104] As mentioned above, the process 1 100 may vary depending on the nature of the data included in the signal.
[00105] In some embodiments, the data is representative of discrete measurements output by a sensor in temporal order. In such embodiments, the diagnostic platform may apply a classifier to the signal to obtain a plurality of outputs, each of which is representative of a classification of a different portion of the signal. As mentioned above, the classifier may be representative of a model that when applied to datapoints included in a portion of the signal, classifies that portion based on an analysis of the datapoints. Depending on the nature of the data, the classifier may be designed and/or trained to distinguish between ambulatory periods and non-ambulatory periods, sleeping periods and nonsleeping periods, stressful periods and non-stressful periods, etc. To identify bouts of a given activity (e.g., ambulating, sleeping, etc.), the diagnostic platform can join consecutive portions that have the same classification. Superset portions that are assigned a given classification (e.g., ambulating) and exceed a given length (e.g., two or more portions) may be identified as bouts. As discussed above, the diagnostic platform may derive a measurement for each superset portion. As an example, the diagnostic platform may determine the step count
during an ambulatory bout by summing the steps taken during the portions included in the corresponding superset portion.
[00106] In other embodiments, the data is representative of inputs provided by one or more individuals. For example, the data may include values that are representative responses to queries by a single person or multiple people. In such embodiments, the diagnostic platform may codify the signal (and more specifically, its values) into a numerical format. Such an approach allows the inputs to be more readily compared across different portions of the signal. Moreover, the diagnostic platform may filter the coded values included in the signal to remove outlier inputs and/or bin the coded values such that inputs are assigned to categories, for example, representing characteristics of the individual(s). The categories may correspond to different age ranges, geographical locations, disease classifications, disease severity classifications, and the like.
[00107] Figure 12 includes a flow diagram of a process 1200 for estimating the reliability of a signal that includes values generated by an IMU. Note that the process 1200 is described in the context of values generated by an IMU for the purpose of illustration. The process 1200 may be similarly applicable to signals output by other types of sensors.
[00108] Initially, a diagnostic platform can acquire a signal that includes values generated by an IMU over an interval of time (step 1201 ). For example, the signal may include values generated over the course of a day, week, month, etc. In some embodiments, the diagnostic platform is configured to apply a classifier to the signal to obtain a plurality of outputs, each of which is representative of a classification of a different one of a plurality of segments of the signal (step 1202). For example, the diagnostic platform may apply an activity classifier to the signal in order to classify each segment as either ambulatory or non-ambulatory. In such embodiments, the diagnostic platform may identify bouts of activity by joining consecutive ones of the plurality of segments that have the same
classification (step 1203). Referring again to the aforementioned example, the diagnostic platform may discover walking bouts by identifying instances where multiple ambulatory segments have been joined together. The diagnostic platform can then produce a plurality of measurements by deriving a separate measurement for each of the bouts of activity (step 1204). For example, the diagnostic platform may derive measurements such as step count, cadence, bout duration, arm swing magnitude, and arm swing range of motion. In some embodiments, the diagnostic platform determines more than one of these measurements for each of the bouts of activity.
[00109] Thereafter, the diagnostic platform can determine an optimal scope for aggregation (step 1205). As discussed above with reference to Figure 11 , the diagnostic platform can accomplish this by defining a plurality of candidate scopes corresponding to different durations and, for each of the plurality of candidate scopes, calculate an aggregated measurement within that candidate scope and then determine a reliability metric based on a comparison of the aggregated measurement to aggregated measurements calculated for adjacent candidate scopes. The diagnostic platform can plot the plurality of reliability metrics against the plurality of candidate scopes, as shown in Figure 7. Based on an analysis of the plot, the diagnostic platform can identify the optimal scope for aggregation.
[00110] After the optimal scope is determined, the diagnostic platform can aggregate the values included in the signal via a mathematical operator, such that a plurality of aggregated values are produced for a plurality of windows defined across the signal in accordance with the optimal scope (step 1206). Examples of mathematical operators include the max operator, min operator, and 95th-percentile operator. The diagnostic platform can then populate the plurality of aggregated values into a data structure (e.g., a tabular data structure) such that the plurality of aggregated values are paired together (step 1207). Based on an analysis of the tabular data structure, the diagnostic platform can estimate the
test-retest reliability of the signal and provide statistical inference of the estimate (step 1208).
Processing System
[00111] Figure 13 includes a block diagram of a processing system 1300 in which at least some operations described herein can be implemented. For example, components of the processing system 1300 may be hosted on a computing device that includes a diagnostic platform.
[00112] The processing system 1300 can include a processor 1302, main memory 1306, non-volatile memory 1310, network adapter 1312, video display 1318, input/output devices 1320, control device 1322 (e.g., a keyboard or pointing device such as a computer mouse or trackpad), drive unit 1324 including a storage medium 1326, and signal generation device 1330 that are communicatively connected to a bus 1316. The bus 1316 is illustrated as an abstraction that represents one or more physical buses or point-to-point connections that are connected by appropriate bridges, adapters, or controllers. The bus 1316, therefore, can include a system bus, a Peripheral Component Interconnect (“PCI”) bus or PCI-Express bus, a HyperTransport (“HT”) bus, an Industry Standard Architecture (“ISA”) bus, a Small Computer System Interface (“SCSI”) bus, a Universal Serial Bus (“USB”) data interface, an Inter- Integrated Circuit (“l2C”) bus, or a high-performance serial bus developed in accordance with Institute of Electrical and Electronics Engineers (“IEEE”) 1394.
[00113] While the main memory 1306, non-volatile memory 1310, and storage medium 1326 are shown to be a single medium, the terms “machine-readable medium” and “storage medium” should be taken to include a single medium or multiple media (e.g., a centralized/distributed database and/or associated caches and servers) that store one or more sets of instructions 1328. The terms “machine-readable medium” and “storage medium” shall also be taken to include any medium that is capable of storing, encoding, or carrying a set of instructions for execution by the processing system 1300.
[00114] In general, the routines executed to implement the embodiments of the disclosure can be implemented as part of an operating system or a specific application, component, program, object, module, or sequence of instructions (collectively referred to as “computer programs”). The computer programs typically comprise one or more instructions (e.g., instructions 1304, 1308, 1328) set at various times in various memory and storage devices in a computing device. When read and executed by the processors 1302, the instruction(s) cause the processing system 1300 to perform operations to execute elements involving the various aspects of the present disclosure.
[00115] Further examples of machine- and computer-readable media include recordable-type media, such as volatile memory devices and non-volatile memory devices 1310, removable disks, hard disk drives, and optical disks (e.g., Compact Disk Read-Only Memory (“CD-ROMs”) and Digital Versatile Disks (“DVDs”)), and transmission-type media, such as digital and analog communication links.
[00116] The network adapter 1312 enables the processing system 1300 to mediate data in a network 1314 with an entity that is external to the processing system 1300 through any communication protocol supported by the processing system 1300 and the external entity. The network adapter 1312 can include a network adaptor card, a wireless network interface card, a router, an access point, a wireless router, a switch, a multilayer switch, a protocol converter, a gateway, a bridge, bridge router, a hub, a digital media receiver, a repeater, or any combination thereof.
Remarks
[00117] The foregoing description of various embodiments of the claimed subject matter has been provided for the purposes of illustration and description. It is not intended to be exhaustive or to limit the claimed subject matter to the precise forms disclosed. Many modifications and variations will be apparent to one skilled in the art. Embodiments were chosen and described in order to best
describe the principles of the invention and its practical applications, thereby enabling those skilled in the relevant art to understand the claimed subject matter, the various embodiments, and the various modifications that are suited to the particular uses contemplated.
[00118] Although the Detailed Description describes certain embodiments and the best mode contemplated, the technology can be practiced in many ways no matter how detailed the Detailed Description appears. Embodiments can vary considerably in their implementation details, while still being encompassed by the specification. Particular terminology used when describing certain features or aspects of various embodiments should not be taken to imply that the terminology is being redefined herein to be restricted to any specific characteristics, features, or aspects of the technology with which that terminology is associated. In general, the terms used in the following claims should not be construed to limit the technology to the specific embodiments disclosed in the specification, unless those terms are explicitly defined herein. Accordingly, the actual scope of the technology encompasses not only the disclosed embodiments, but also all equivalent ways of practicing or implementing the embodiments.
[00119] The language used in the specification has been principally selected for readability and instructional purposes. It may not have been selected to delineate or circumscribe the subject matter. It is therefore intended that the scope of the technology be limited not by this Detailed Description, but rather by any claims that issue on an application based hereon. Accordingly, the disclosure of various embodiments is intended to be illustrative, but not limiting, of the scope of the technology as set forth in the following claims.
Claims
1 . A method performed by a processor of a computing device, the method comprising: acquiring a signal that includes values generated by an inertial measurement unit over an interval of time; applying a classification model to the signal to obtain a plurality of outputs, each of which is representative of a classification of a different one of a plurality of segments of the signal; identifying a plurality of walking bouts by joining consecutive ones of the plurality of segments that are classified as ambulatory; producing a plurality of measurements by deriving a measurement for each of the plurality of walking bouts; determining an optimal scope for aggregation by - defining a plurality of candidate scopes corresponding to different durations, for each of the plurality of candidate scopes, calculating, based on the plurality of measurements, an aggregated measurement within that candidate scope, and determining a reliability metric based on a comparison of the aggregated measurement to aggregated measurements calculated for adjacent candidate scopes, plotting the plurality of reliability metrics against the plurality of candidate scopes in a plot, and identifying the optimal scope based on an analysis of the plot; aggregating the values included in the signal via a mathematical operator, such that a plurality of aggregated values are produced for a
plurality of windows defined across the signal in accordance with the optimal scope; generating a tabular data structure in which the plurality of aggregated values are paired together; and estimating test-retest reliability of the signal based on an analysis of the tabular data structure.
2. The method of claim 1 , wherein the inertial measurement unit is included in the computing device.
3. The method of claim 1 , wherein the classification indicates whether each of the plurality of segments is ambulatory or non-ambulatory.
4. The method of claim 1 , wherein the plurality of windows do not overlap with one another.
5. The method of claim 1 , wherein each of the plurality of windows overlaps at least one other of the plurality of windows.
6. The method of claim 1 , further comprising: performing a preprocessing operation in which:
(i) the values included in the signal are resampled at a predetermined resolution,
(ii) bias is removed from the resampled values, and
(iii) noise Is removed from the resampled values by applying a low- pass filter or a high-pass filter thereto.
7. A non-transitory medium with instructions stored therein that, when executed by a processor of a computing device, cause the computing device to perform operations comprising:
obtaining a plurality of measurements that are derived via analysis of a signal that is generated over an interval of time, wherein each of the plurality of measurements is associated with a corresponding one of a plurality of segments of the signal; defining a plurality of candidate scopes for aggregation that correspond to different durations; for each of the plurality of candidate scopes, calculating, based on the plurality of measurements, an aggregated measurement within that candidate scope; and computing an intraclass correlation coefficient based on a comparison of the aggregated measurement to aggregated measurement calculated for adjacent candidate scopes; and identifying an optimal scope from among the plurality of candidate scopes based on an analysis of the plurality of intraclass correlation coefficients computed for the plurality of candidate scopes.
8. The non-transitory medium of claim 7, wherein the operations further comprise: aggregating, via a mathematical operator, data included in the signal, such that a plurality of aggregated values are produced for a plurality of windows defined across the signal in accordance with the optimal scope; generating a tabular data structure in which the plurality of aggregated values are paired together; and estimating test-retest reliability of the signal based on an analysis of the tabular data structure.
9. The non-transitory medium of claim 8, wherein said establishing comprises:
for each of the plurality of intraclass correlation coefficients, computing a point estimate and/or a confidence interval, and determining the test-retest reliability based on the point estimates computed for the plurality of intraclass correlation coefficients, the confidence intervals computed for the plurality of intraclass correlation coefficients, or a combination thereof.
10. The non-transitory medium of claim 8, wherein the mathematical operator is a max operator, a min operator, or a 95th-percentile operator.
1 1 . The non-transitory medium of claim 8, further comprising: determining, based on the test-retest reliability, reliability of a digital biomarker that is computable using the signal.
12. The non-transitory medium of claim 7, wherein the signal is representative of (i) discrete measurements generated by a sensor and/or (ii) inputs provided by an individual.
13. The non-transitory medium of claim 7, wherein the signal includes discrete measurements generated by a sensor, and wherein the operations further comprise: applying a classification model to the signal to obtain a plurality of outputs, each of which is representative of a classification of a different portion of the signal; creating the plurality of segments by joining consecutive portions that have the same classification; and producing the plurality of measurements by deriving a measurement for each of the plurality of segments.
14. The non-transitory medium of claim 7, wherein the signal includes inputs provided by one or more individuals, and wherein the operations further comprise: coding the inputs into a numerical format.
15. The non-transitory medium of claim 14, wherein the inputs are representative of responses provided by individuals in response to being queried, and wherein the operations further comprise: performing at least one of: filtering the coded data to remove outlier responses, and binning the coded data such that the responses are assigned to categories corresponding to characteristics of the individuals.
16. The non-transitory medium of claim 15, wherein the categories correspond to different age ranges.
17. The non-transitory medium of claim 15, wherein the categories correspond to different geographical locations.
18. The non-transitory medium of claim 15, wherein the categories correspond to different disease classifications.
19. A method performed by a processor of a computing device, the method comprising: defining a plurality of candidate scopes for aggregation of data that is included in a signal generated over an interval of time, wherein each of the plurality of candidate scopes is associated with a different duration for segmenting the interval of time into windows;
for each of the plurality of candidate scopes, producing an aggregated measurement within that candidate scope; and determining a reliability metric based on a comparison of the aggregated measurement to aggregated measurements calculated for adjacent candidate scopes; and identifying an optimal scope for aggregation from among the plurality of candidate scopes based on an analysis of the plurality of reliability metrics determined for the plurality of candidate scopes.
20. The method of claim 19, further comprising: aggregating, via a mathematical operator, the data such that a plurality of aggregated values are produced for a plurality of windows defined across the signal in accordance with the optimal scope; generating a tabular data structure in which the plurality of aggregated values are paired together; and estimating test-retest reliability of the signal based on an analysis of the tabular data structure.
21 . The method of claim 20, wherein the plurality of windows are defined such that no overlapping occurs, and wherein the pairing of the plurality of aggregated values is governed by a parameter that specifies a maximum duration over which pairings are sought.
22. The method of claim 20, wherein the plurality of windows are defined such that overlapping occurs, and
wherein said generating includes populating pairings of the plurality of aggregated values and shifted pairings of the plurality of aggregated values in the tabular data structure.
23. The method of claim 20, wherein the plurality of windows are defined such that no overlapping occurs, wherein the pairing of the plurality of aggregated values is governed by a parameter that specifies a gap over which pairings are sought, and wherein said generating includes populating, for each of the plurality of aggregated values, all possible pairings within the gap in the tabular data structure.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263476112P | 2022-12-19 | 2022-12-19 | |
| PCT/US2023/083969 WO2024137323A1 (en) | 2022-12-19 | 2023-12-14 | Establishing optimal aggregation of data in signals generated in free-living scenarios |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4637548A1 true EP4637548A1 (en) | 2025-10-29 |
Family
ID=91589857
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23908182.1A Pending EP4637548A1 (en) | 2022-12-19 | 2023-12-14 | Establishing optimal aggregation of data in signals generated in free-living scenarios |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4637548A1 (en) |
| JP (1) | JP2025541420A (en) |
| WO (1) | WO2024137323A1 (en) |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9053222B2 (en) * | 2002-05-17 | 2015-06-09 | Lawrence A. Lynn | Patient safety processor |
| EP4030993B1 (en) * | 2019-09-17 | 2026-04-15 | F. Hoffmann-La Roche AG | Improvements in personalized healthcare for patients with movement disorders |
-
2023
- 2023-12-14 EP EP23908182.1A patent/EP4637548A1/en active Pending
- 2023-12-14 WO PCT/US2023/083969 patent/WO2024137323A1/en not_active Ceased
- 2023-12-14 JP JP2025535918A patent/JP2025541420A/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024137323A1 (en) | 2024-06-27 |
| JP2025541420A (en) | 2025-12-18 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Rodriguez-León et al. | Mobile and wearable technology for the monitoring of diabetes-related parameters: Systematic review | |
| US20230082019A1 (en) | Systems and methods for monitoring brain health status | |
| Low | Harnessing consumer smartphone and wearable sensors for clinical cancer research | |
| US9955869B2 (en) | System and method for supporting health management services | |
| US20190117143A1 (en) | Methods and Apparatus for Assessing Depression | |
| US20200075167A1 (en) | Dynamic activity recommendation system | |
| CN115769302A (en) | Epidemic disease monitoring system | |
| JP2018524137A (en) | Method and system for assessing psychological state | |
| US12315604B2 (en) | Recurring remote monitoring with real-time exchange to analyze health data and generate action plans | |
| US11816750B2 (en) | System and method for enhanced curation of health applications | |
| WO2020005822A1 (en) | Activity tracking and classification for diabetes management system, apparatus, and method | |
| Shukur et al. | Diabetes at a glance: assessing AI strategies for early diabetes detection and intervention | |
| Riadhusin et al. | Ubiquitous health monitoring using bio-wearable devices | |
| Elango et al. | Super artificial intelligence medical healthcare services and smart wearable system based on IoT for remote health monitoring | |
| O'Brien et al. | Automate, illuminate, predict: a universal framework for integrating wearable sensors in healthcare | |
| Mila et al. | MASC: wearable design for infectious disease detection through machine learning | |
| US20190279752A1 (en) | Generation of adherence-improvement programs | |
| Shen et al. | Conformal prediction quantifies wearable cuffless blood pressure with certainty | |
| WO2025093921A1 (en) | Systems and methods for providing personalized health risk assessments and recommendations | |
| WO2024137323A1 (en) | Establishing optimal aggregation of data in signals generated in free-living scenarios | |
| US20240006067A1 (en) | System by which patients receiving treatment and at risk for iatrogenic cytokine release syndrome are safely monitored | |
| US20240074661A1 (en) | Wearable monitor and application | |
| WO2025128637A1 (en) | Systems and methods for identifying disease progression biomarkers using non-ambulatory patient data | |
| Nawaz et al. | A novel methodology for patient prescreening using wireless body area networks (WBANs) | |
| JP2024519249A (en) | System for determining similarity of a sequence of glucose values - Patents.com |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250703 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |