WO2024257656A1 - データ処理装置及びプログラム - Google Patents

データ処理装置及びプログラム Download PDF

Info

Publication number
WO2024257656A1
WO2024257656A1 PCT/JP2024/020404 JP2024020404W WO2024257656A1 WO 2024257656 A1 WO2024257656 A1 WO 2024257656A1 JP 2024020404 W JP2024020404 W JP 2024020404W WO 2024257656 A1 WO2024257656 A1 WO 2024257656A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
video data
still image
information
generation unit
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2024/020404
Other languages
English (en)
French (fr)
Inventor
裕子 石若
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
SoftBank Corp
Original Assignee
SoftBank Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by SoftBank Corp filed Critical SoftBank Corp
Publication of WO2024257656A1 publication Critical patent/WO2024257656A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • AHUMAN NECESSITIES
    • A01AGRICULTURE; FORESTRY; ANIMAL HUSBANDRY; HUNTING; TRAPPING; FISHING
    • A01KANIMAL HUSBANDRY; AVICULTURE; APICULTURE; PISCICULTURE; FISHING; REARING OR BREEDING ANIMALS, NOT OTHERWISE PROVIDED FOR; NEW BREEDS OF ANIMALS
    • A01K61/00Culture of aquatic animals
    • A01K61/90Sorting, grading, counting or marking live aquatic animals, e.g. sex determination
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features

Definitions

  • the present invention relates to a data processing device and a program.
  • Patent Document 1 describes a method for automatically generating training data that reproduces the various swimming movements of fish using musical scores that correspond to existing music pieces.
  • a data processing device may include a video data acquisition unit that acquires video data of an object.
  • the data processing device may include a data generation unit that converts information on the time-series movement of the object in the video data into still image elements, and generates still image data including the elements of the converted still image.
  • the data generation unit may generate the single still image data consisting of the multiple parts by performing at least one of generating one part of the multiple parts constituting the single still image data from one frame of the multiple frames of the video data and generating one part of the multiple parts constituting the single still image data from multiple frames of the video data.
  • the data generation unit may generate the multiple parts by converting information on three-dimensional positions of the multiple objects included in the multiple frames of the video data into pixel value information for the multiple frames of the video data.
  • the video data acquisition unit may acquire multiple video data captured in parallel by multiple cameras that capture the object from multiple different directions, and the data generation unit may generate the multiple parts by converting information on three-dimensional positions of the multiple objects included in the multiple frames of the video data into pixel value information for each of the multiple video data.
  • the data generation unit may generate the still image data that represents information about the time-series movement of the object in the video data by at least one of pixel values, feature vectors, and edges.
  • the data generation unit may convert information on changes in the positions of the objects in time series in the video data into the elements of the still image, and generate the still image data including the elements of the converted still image.
  • the data generation unit may generate the still image data expressing information on changes in the positions of the objects in time series in the video data by pixel values.
  • the data generation unit may convert information on three-dimensional positions where objects exist for each frame into pixel value information.
  • the data generation unit may generate the still image data expressing information on changes in the positions of the objects in time series in the video data by feature vectors.
  • the data generation unit may convert information on three-dimensional positions where objects exist for each frame into feature vector information.
  • the data generation unit may generate the still image data expressing information on changes in the positions of the objects in time series in the video data by edges.
  • the data generation unit may convert information on three-dimensional positions where objects exist for each frame into edge information.
  • the data generation unit may convert information on time-series shape changes of the object in the video data into the elements of the still image, and generate the still image data including the elements of the converted still image.
  • the data generation unit may generate the still image data expressing the information on time-series shape changes of the object in the video data by pixel values.
  • the data generation unit may convert information on three-dimensional positions where multiple structural points of the object exist for each frame into pixel value information.
  • the data generation unit may generate the still image data expressing the information on time-series shape changes of the object in the video data by feature vectors.
  • the data generation unit may convert information on three-dimensional positions where multiple structural points of the object exist for each frame into feature vector information.
  • the data generation unit may generate the still image data expressing the information on time-series shape changes of the object in the video data by edges.
  • the data generation unit may convert information on three-dimensional positions where multiple structural points of the object exist for each frame into edge information.
  • the data generation unit may convert information about the time-series movement of the object in the video data into text elements, and generate text data including the converted text elements.
  • the data generation unit may convert information on the time-series movement of the object in the video data into genome elements, and generate genome data including the converted genome elements.
  • Any of the data processing devices may further include a classification result acquisition unit that inputs the still image data generated by the data generation unit into a still image classification network trained to classify still image data and acquires the classification result output from the still image classification network, and an alert output unit that outputs an alert when the classification result acquired by the classification result acquisition unit satisfies a predetermined condition.
  • Any of the data processing devices may further include a conversion data acquisition unit that acquires still image data including still image elements converted from information on the time-series movement of an object in the video data, and a video data generation unit that generates the video data from the still image data.
  • a data processing device may include a still image data acquisition unit that acquires still image data including still image elements converted from information on the time-series movement of an object in video data.
  • the data processing device may include a video data generation unit that generates the video data from the still image data.
  • a data processing device may include a video data acquisition unit that acquires video data of an object.
  • the data processing device may include a data generation unit that converts information on the time series movement of the object in the video data into text elements and generates text data including the converted text elements.
  • the data generation unit may generate the text data that represents the information on the time series movement of the object in the video data by at least one of characters, words, fonts, and ASCII codes.
  • the data generation unit may convert information on changes in the positions of multiple objects in the video data in time series into text elements and generate the text data including the converted text elements.
  • the data generation unit may generate the text data that represents the information on changes in the positions of multiple objects in the video data in time series by characters.
  • the data generation unit may convert information on the three-dimensional position where an object exists for each frame into character information.
  • the data generation unit may generate the text data that represents the information on changes in the positions of multiple objects in the video data in time series by words.
  • the data generating unit may convert information on a three-dimensional position where an object exists for each frame into information on words.
  • the data generating unit may generate the text data expressing information on time-series positional changes of a plurality of objects in the video data by ASCII code.
  • the data generating unit may convert information on a three-dimensional position where an object exists for each frame into information on ASCII code.
  • the data generating unit may generate the text data expressing information on time-series positional changes of a plurality of objects in the video data by at least one of characters and words and a font.
  • the data generating unit may convert information on a three-dimensional position where an object exists for each frame into at least one of characters and words and a font.
  • the data generating unit may convert information on time-series shape changes of an object in the video data into text elements, and generate the text data including the converted text elements.
  • the data generating unit may generate the text data expressing information on time-series shape changes of an object in the video data by characters.
  • the data generation unit may generate the text data expressing information on time-series shape changes of an object in the video data by ASCII code.
  • the data generation unit may generate the text data expressing information on time-series shape changes of an object in the video data by at least one of characters and words and a font.
  • the data generation unit may generate the text data from a plurality of video data in which an object is captured from a plurality of different directions.
  • the data generation unit may convert information on time-series movement of an object in the plurality of video data into text elements and generate the text data including the converted text elements.
  • the data processing device may include a classification result acquisition unit that inputs the text data generated by the data generation unit to a still image classification network and acquires a classification result output from the still image classification network.
  • the data processing device may include an alert output unit that outputs an alert according to the classification result by the classification result acquisition unit.
  • a data processing device may include a video data acquisition unit that acquires video data of an object.
  • the data processing device may include a data generation unit that converts information on the time series movement of the object in the video data into genome elements and generates genome data including the converted genome elements.
  • the data generation unit may generate the genome data representing information on time series positional changes of multiple objects in the video data by base sequences.
  • the data generation unit may convert information on three-dimensional positions where an object exists for each frame into base sequence information.
  • the data generation unit may convert information on time series shape changes of an object in the video data into genome elements and generate the genome data including the converted genome elements.
  • the data generation unit may generate the genome data representing information on time series shape changes of an object in the video data by base sequences.
  • the data generation unit may convert information on three-dimensional positions where multiple structural points of an object exist for each frame into base sequence information.
  • the data generation unit may generate the genome data from a plurality of video data obtained by capturing images of an object from a plurality of different directions.
  • the data generation unit may convert information on the time series movement of an object in the plurality of video data into genome elements, and generate the genome data including the converted genome elements.
  • the data generation unit may generate the genome data representing information on the time series movement of an object in the plurality of video data by a base sequence.
  • the data processing device may include a classification result acquisition unit that inputs the genome data generated by the data generation unit to a genome classification network and acquires a classification result output from the genome classification network.
  • the data processing device may include an alert output unit that outputs an alert according to the classification result by the classification result acquisition unit.
  • a program for causing a computer to function as the data processing device.
  • FIG. 1 illustrates a schematic diagram of an example of a data processing device 100.
  • 2 is an explanatory diagram for explaining a conversion process for converting moving image data 200 into still image data 300.
  • FIG. 1 is an explanatory diagram for explaining a conversion process for converting a plurality of pieces of video data 200 into still image data 300.
  • FIG. 1 is an explanatory diagram for explaining a conversion process for converting a plurality of pieces of video data 200 into still image data 300.
  • FIG. 4 is an explanatory diagram for explaining a conversion process for converting video data 200 into text data 400.
  • FIG. 4 is an explanatory diagram for explaining a conversion process for converting video data 200 into text data 400.
  • FIG. 4 is an explanatory diagram for explaining a conversion process for converting video data 200 into text data 400.
  • FIG. 1 is an explanatory diagram for explaining a conversion process for converting video data 200 into genome data 500.
  • FIG. 1 illustrates an example of a functional configuration of a data processing device 100.
  • An example of the hardware configuration of a computer 1200 functioning as the data processing device 100 is shown in schematic form.
  • the data processing device 100 converts the video data into data with a lower dimension. For example, the data processing device 100 converts the video data into two-dimensional data.
  • the data processing device 100 converts video data into still image data.
  • the data processing device 100 may convert information contained in the video data into still image elements, thereby converting the video data into still image data.
  • the data processing device 100 performs object detection on the video data and converts information on the time-series movement of the object into still image elements, thereby converting the video data into still image data.
  • the data processing device 100 may convert the video data into still image data according to conversion information including a conversion rule.
  • the data processing device 100 may realize object monitoring by analyzing the converted still image data. This can reduce the processing load.
  • neural networks that analyze still image data have been studied very extensively, and many neural networks already exist, making it possible to reuse such neural networks.
  • the data processing device 100 converts video data into text data.
  • the data processing device 100 may convert information contained in the video data into text elements, thereby converting the video data into text data.
  • the data processing device 100 performs object detection on the video data, and converts information on the time-series movement of the object into text elements, thereby converting the video data into text data.
  • the data processing device 100 may convert the video data into text data according to conversion information including a conversion rule.
  • the data processing device 100 may realize object monitoring by analyzing the converted text data. This can reduce the processing load.
  • neural networks that analyze text data have been studied very extensively, and many neural networks already exist, making it possible to reuse such neural networks.
  • the data processing device 100 may convert the video data into data in which specific characters are listed or data in which numbers are listed.
  • the data processing device 100 converts the video data into genome data, for example.
  • the genome data may be data composed of a base sequence.
  • the data processing device 100 may convert the video data into genome data by converting information contained in the video data into genome elements. For example, the data processing device 100 performs object detection on the video data and converts information on the time-series movement of the object into genome elements, thereby converting the video data into genome data.
  • the data processing device 100 may convert the video data into genome data according to conversion information including a conversion rule.
  • the data processing device 100 may realize object monitoring by analyzing the converted genome data. This can reduce the processing load.
  • neural networks that analyze genome data have been studied very widely, and many neural networks already exist, making it possible to reuse such neural networks.
  • FIG. 1 shows a schematic diagram of an example of a data processing device 100.
  • the data processing device 100 may acquire video data of an object, and convert the acquired video data into data of a lower dimension than the video data (the converted data may be referred to as converted data).
  • Examples of data of a lower dimension than video data include still image data, text data, and genome data.
  • the object may be anything that changes over time.
  • the object may be something whose position changes over time, something whose shape changes over time, or something whose position and shape change over time.
  • the object may be something whose position changes over time and not whose shape changes.
  • the object may be something whose position changes over time and not whose position changes over time.
  • the object may be the subject of monitoring.
  • the object is a fish and the subject of monitoring is a school of fish, but the present invention is not limited to this.
  • the data processing device 100 may acquire video data captured by the camera 102.
  • the camera 102 may be built into the data processing device 100.
  • the camera 102 may be externally attached to the data processing device 100.
  • the camera 102 may be connected to the data processing device 100 via any network.
  • the data processing device 100 may receive video data from the communication terminal 30.
  • the communication terminal 30 transmits video data captured by, for example, a camera 32 to the data processing device 100.
  • the camera 32 may be built into the communication terminal 30.
  • the camera 32 may be externally attached to the communication terminal 30.
  • the camera 32 may be connected to the communication terminal 30 via any network.
  • the communication terminal 30 may be a PC (Personal Computer), a tablet terminal, a smartphone, or the like.
  • the data processing device 100 and the communication terminal 30 may communicate via a network 20.
  • the network 20 may include the Internet.
  • the network 20 may include a LAN (Local Area Network).
  • the network 20 may include a mobile communication network.
  • the mobile communication network may conform to any of the following communication methods: 5G (5th Generation) communication method, LTE (Long Term Evolution) communication method, 3G (3rd Generation) communication method, and 6G (6th Generation) communication method or later.
  • the data processing device 100 may pre-store a classification network that has been trained to classify the converted data.
  • the data processing device 100 may input the converted data to the classification network and obtain the classification result output from the classification network.
  • the data processing device 100 may output an alert when the classification result satisfies a predetermined condition.
  • the data processing device 100 acquires video data capturing a school of fish in a fish farm, converts information on the time series movements of multiple fish in the video data into still image elements, and generates still image data including the converted still image elements.
  • the data processing device 100 converts, for example, information on the time series changes in position of multiple fish in the video data into still image elements, and generates still image data including the converted still image elements.
  • the data processing device 100 may generate still image data for any time unit, such as seconds, minutes, hours, days, weeks, months, years, etc.
  • the data processing device 100 may generate a still image classification network that classifies the still image data by performing machine learning using the large amount of still image data generated.
  • the data processing device 100 may classify the still image data by inputting the newly generated still image data into the generated still image classification network.
  • the data processing device 100 may use an existing still image classification network. For example, the data processing device 100 acquires and stores in advance a still image classification network generated by another device. The data processing device 100 then inputs still image data generated subsequently into the still image classification network, thereby classifying the still image data.
  • the data processing device 100 may further include a function for generating video data from the converted data.
  • the data processing device 100 may generate video data from the converted data that it has generated.
  • the data processing device 100 may convert the converted data into video data by using the conversion information that was used when converting the video data into the converted data.
  • the data processing device 100 may generate video data from converted data generated by another device.
  • the data processing device 100 may receive conversion information used when converting video data into converted data from the other device that generated the converted data, and may use the conversion information to convert the converted data received from the other device into video data.
  • the data processing device 100 may have a function to generate video data from converted data, without having a function to generate converted data from video data.
  • FIG. 2 is an explanatory diagram for explaining the conversion process for converting video data 200 into still image data 300.
  • the camera 102 converts video data 200 capturing an image of a school of fish in a fish farm 40 into still image data 300.
  • the data processing device 100 generates one portion 302 constituting the still image data 300 from one frame of the video data 200.
  • the data processing device 100 may generate still image data 300 consisting of multiple portions 302 by generating multiple portions 302, each of which corresponds to one of the multiple frames of the video data 200.
  • the data processing device 100 may generate one portion 302 from multiple frames of the video data 200.
  • the data processing device 100 may generate still image data 300 from multiple video data 200 captured in parallel by multiple cameras.
  • Figure 3 shows examples of frame 211 of video data 210 captured from above, frame 221 of video data 220 captured from below, frame 231 of video data 230 captured from side A, frame 241 of video data 240 captured from side B, frame 251 of video data 250 captured from side C, and frame 261 of video data 260 captured from side D at a certain timing.
  • the data processing device 100 for example, generates part 311 from frame 211, generates part 321 from frame 221, generates part 331 from frame 231, generates part 341 from frame 241, generates part 351 from frame 251, and generates part 361 from frame 261.
  • the data processing device 100 connects these as shown in FIG. 4.
  • the data processing device 100 similarly generates and connects still image parts for the next frame. By repeating this process, the data processing device 100 generates still image data 300.
  • the data processing device 100 may generate one portion from multiple consecutive frames, rather than generating one portion from one frame, as described above. For example, the data processing device 100 generates one portion 311 from 10 consecutive frames of the video data 210 captured from above, generates one portion 321 from 10 consecutive frames of the video data 220 captured from below, generates one portion 331 from 10 consecutive frames of the video data 230 captured from side A, generates one portion 341 from 10 consecutive frames of the video data 240 captured from side B, generates one portion 351 from 10 consecutive frames of the video data 250 captured from side C, and generates one portion 361 from 10 consecutive frames of the video data 260 captured from side D.
  • 10 frames is just an example, and the number of frames may be other numbers.
  • FIG. 5 is an explanatory diagram for explaining the conversion process for converting video data 200 into text data 400.
  • video data 200 capturing an image of a school of fish in a fish farm 40 by a camera 102 is converted into text data 400.
  • the data processing device 100 acquires video data 200 from the camera 102.
  • the data processing device 100 converts information on the time-series movement of the school of fish in the video data 200 into text elements, and generates text data 400 that includes the converted text elements.
  • the data processing device 100 generates one portion 402 constituting the text data 400 from one frame of the video data 200.
  • the data processing device 100 may generate text data 400 consisting of multiple portions 402 by generating multiple portions 402 each corresponding to one of the multiple frames of the video data 200. Note that the data processing device 100 may generate one portion 402 from multiple frames of the video data 200.
  • the data processing device 100 may generate the text data 400 for any period of time. For example, when generating the text data 400 for each day, the data processing device 100 generates the text data 400 corresponding to one day by generating a plurality of portions 402 from the frames of the video data 200 for one day.
  • the data processing device 100 may generate text data 400 from multiple video data 200 captured in parallel by multiple cameras.
  • six video data 200 capturing images of a school of fish in a fish farm 40 from six directions, as shown in Figure 3, are converted into text data 400.
  • FIG. 7 is an explanatory diagram for explaining the conversion process for converting video data 200 into genome data 500.
  • video data 200 capturing an image of a school of fish in a fish farm 40 by a camera 102 is converted into genome data 500.
  • the data processing device 100 acquires video data 200 from the camera 102.
  • the data processing device 100 converts information on the time-series movement of the school of fish in the video data 200 into genome elements, and generates genome data 500 including the converted genome elements.
  • the data processing device 100 generates one portion constituting the genome data 500 from one frame of the video data 200.
  • the data processing device 100 may generate genome data 500 consisting of multiple portions by generating multiple portions each corresponding to one of the multiple frames of the video data 200.
  • the data processing device 100 may generate one portion 402 from multiple frames of the video data 200.
  • the data processing device 100 may generate genome data 500 for any period of time. For example, when generating genome data 500 for each day, the data processing device 100 generates genome data 500 corresponding to one day by generating multiple parts from frames of one day's worth of video data 200.
  • the data processing device 100 may generate genome data 500 from multiple video data 200 captured in parallel by multiple cameras, similar to the still image data 300 and text data 400.
  • FIG. 8 shows an example of a schematic functional configuration of the data processing device 100.
  • the data processing device 100 includes a conversion information storage unit 112, a model storage unit 114, a video data acquisition unit 116, a data generation unit 118, a classification result acquisition unit 120, a model generation unit 122, an output unit 130, a conversion data acquisition unit 142, and a video data generation unit 144.
  • the conversion information storage unit 112 stores conversion information including conversion rules for converting the video data 200.
  • the conversion information storage unit 112 may store conversion information that has been registered in advance.
  • the conversion information storage unit 112 stores conversion information including conversion rules for converting the video data 200 into still image data 300.
  • the conversion information storage unit 112 stores conversion information including conversion rules for converting the video data 200 into text data 400.
  • the conversion information storage unit 112 stores conversion information including conversion rules for converting the video data 200 into genome data 500.
  • the model storage unit 114 stores a classification network trained to classify converted data.
  • the model storage unit 114 may store a classification network registered in advance.
  • the model storage unit 114 stores, for example, a classification network generated by another device.
  • the conversion information storage unit 112 stores a still image classification network.
  • the conversion information storage unit 112 stores a text classification network.
  • the conversion information storage unit 112 stores a genome classification network.
  • the video data 200 may be accompanied by the detection results of an object.
  • the video data 200 may be accompanied by information on the three-dimensional position of an object.
  • the information on the three-dimensional position of an object may be identified by a method using the output of a distance measuring sensor arranged separately from the camera 102 or the camera 32, a method using the image analysis results of the camera 102 or the camera 32, or the like. Note that the method is not limited to these, and any method may be used as long as it is possible to identify the three-dimensional position of an object.
  • the data generation unit 118 may convert information on time-series positional changes of multiple objects in the video data 200 into still image elements, and generate still image data 300 including the converted still image elements.
  • the data generating unit 118 generates still image data 300 that expresses, for example, information on the change in the positions of multiple objects in the video data 200 over time using pixel values.
  • the data generating unit 118 converts, for example, information on the three-dimensional position of an object for each frame into pixel value information.
  • the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of the multiple objects included in the frame as pixel values of three consecutive pixels.
  • the pixel values of the three consecutive pixels indicate the X coordinate, Y coordinate, and Z coordinate of the object.
  • the data generating unit 118 can generate still image data 300 corresponding to the video data 200 by similarly assigning pixel values for other frames of the video data 200.
  • the method of converting the information on the X coordinate, Y coordinate, and Z coordinate of each of the multiple objects into pixel values is not limited to this.
  • the information on the X coordinate, Y coordinate, and Z coordinate of each of the multiple objects may be compressed and converted into pixel values using existing compression technology.
  • Frame divisions may be expressed by arrangement, or one or more pixel values may be used to indicate frame divisions. Other methods may also be used to indicate frame divisions.
  • Grayscale may be used as pixel values. This results in the generation of grayscale still image data 300 that represents information on the change in the positions of multiple objects over time in the video data 200. If multiple objects in the video data 200 are moving in a similar manner, the still image data 300 may have the appearance of a similar pattern being repeated. Furthermore, if multiple objects in the video data 200 are moving rapidly for a certain period of time, the still image data 300 may have the appearance of a part that becomes darker, lighter, or changes more drastically, thereby indicating that the part is different from the other parts.
  • Color may be used as the pixel value. This allows information about changes in the positions of multiple objects over time to be expressed more efficiently than when grayscale is used.
  • color still image data 300 is generated that represents information about changes in the positions of multiple objects over time in the video data 200. If multiple objects in the video data 200 are moving in a similar manner, the still image data 300 may have the same color tone. Also, if multiple objects in the video data 200 are moving rapidly for a certain period of time, the still image data 300 may have an appearance that indicates that the part is different from the other parts by becoming darker or lighter in color or by changing color more drastically.
  • the data generation unit 118 generates still image data 300 that expresses, for example, information on the change in the positions of multiple objects over time in the video data 200 using feature vectors.
  • the data generation unit 118 converts, for example, information on the three-dimensional positions of objects for each frame into feature vector information.
  • the data generation unit 118 may use vectors that quantify position coordinates, velocity change, and direction as a method of expressing feature vectors, or may use a method of learning the features themselves.
  • feature vectors may be generated after taking statistics on the positions of multiple objects over time. Statistics may also be taken from the feature vectors, and these may then be used as feature vectors in a higher layer.
  • the data generation unit 118 generates still image data 300 that uses edges to represent, for example, information on the change in the positions of multiple objects over time in the video data 200.
  • the data generation unit 118 converts, for example, information on the three-dimensional positions of objects for each frame into edge information.
  • the data generation unit 118 may realize the conversion between information on the change in the positions of multiple objects over time and edge information using, for example, a method of binarizing and differentiating, the Canny method, or a method of learning edges or silhouettes using a DNN (Deep Neural Network).
  • DNN Deep Neural Network
  • the data generator 118 may convert information about time-series shape changes of objects in the video data 200 into still image elements, and generate still image data 300 that includes the converted still image elements.
  • the data generation unit 118 generates still image data 300 that expresses, for example, information on the time series shape changes of an object in the video data 200 using pixel values.
  • the data generation unit 118 converts, for example, information on the three-dimensional positions of multiple structural points of an object for each frame into pixel value information. For example, if the object is a living thing or a robot, the structural points of the object may be the positions of joints. If the object does not have joints, the structural points of the object may be each point of the object that affects the movement of the object.
  • the data generation unit 118 assigns the X-coordinate, Y-coordinate, and Z-coordinate of each of the multiple structural points of an object included in the frame as pixel values of three consecutive pixels.
  • the pixel values of three consecutive pixels indicate the X-coordinate, Y-coordinate, and Z-coordinate of the object.
  • the data generation unit 118 can generate still image data 300 corresponding to the video data 200 by similarly assigning pixel values for other frames of the video data 200.
  • the method of converting the information of the X-coordinate, Y-coordinate, and Z-coordinate of each of the multiple objects into pixel values is not limited to this.
  • the information of the X-coordinate, Y-coordinate, and Z-coordinate of each of the multiple objects may be compressed and converted into pixel values using existing compression technology.
  • the frame division may be expressed by the arrangement, or one or more pixel values indicating the frame division may be used. Methods other than these may be used for the frame division.
  • Grayscale may be used as pixel values. This results in the generation of grayscale still image data 300 that represents information on the time series changes in shape of an object in the video data 200. If an object in the video data 200 is moving in a similar manner, the still image data 300 may have the appearance of a similar pattern being repeated. Also, if an object in the video data 200 is moving violently for a certain period of time, the still image data 300 may have the appearance of that part becoming darker, lighter, or changing more dramatically, thereby indicating that that part is different from the other parts.
  • Color may be used as the pixel value. This allows information about changes in the shape of an object over time to be expressed more efficiently than when grayscale is used.
  • color still image data 300 is generated that represents information about changes in the shape of an object over time in the video data 200. If objects in the video data 200 are moving in a similar manner, the still image data 300 may have the same color tone. Also, if an object in the video data 200 is moving violently for a certain period of time, the color of that part may become darker or lighter, or the color tone may change drastically, thereby giving the appearance that the part is different from the other parts.
  • the data generator 118 generates still image data 300 that expresses, for example, information on the time series shape changes of an object in the video data 200 using a feature vector.
  • the data generator 118 converts, for example, information on the three-dimensional position where an object exists for each frame into feature vector information.
  • the data generator 118 generates still image data 300 that uses edges to represent, for example, information on the time-series shape changes of an object in the video data 200.
  • the data generator 118 converts, for example, information on the three-dimensional position where an object exists for each frame into edge information.
  • the data generation unit 118 generates still image data 300 from multiple video data 200 in which an object is captured from multiple different directions.
  • the data generation unit 118 may convert information on the time-series movement of an object in the multiple video data 200 into still image elements, and generate still image data 300 including the converted still image elements.
  • the data generation unit 118 may generate still image data 300 that represents information on the time-series movement of an object in the multiple video data 200 by at least one of pixel values, feature vectors, and edges.
  • the data generation unit 118 may convert information on time-series positional changes of multiple objects in multiple video data 200 into still image elements, and generate still image data 300 including the converted still image elements.
  • the data generating unit 118 generates still image data 300 that expresses, for example, information on the change in the positions of multiple objects in multiple video data 200 over time using pixel values. For each of the multiple video data 200, the data generating unit 118 converts, for example, information on the three-dimensional position of an object for each frame into pixel value information. As a specific example, for one frame of the video data 200, the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of the multiple objects included in the frame as pixel values of three consecutive pixels. As a result, for the portion of the still image data 300 corresponding to that frame, the pixel values of three consecutive pixels indicate the X coordinate, Y coordinate, and Z coordinate of the object. The data generating unit 118 can generate still image data 300 corresponding to the video data 200 by similarly assigning pixel values for other frames of the video data 200.
  • the data generation unit 118 generates still image data 300 that expresses, for example, information on the change in the positions of multiple objects over time in multiple video data 200 using feature vectors. For example, the data generation unit 118 converts information on the three-dimensional position where an object exists for each frame of multiple video data 200 into feature vector information.
  • the data generation unit 118 generates still image data 300 that uses edges to represent, for example, information on the change in the positions of multiple objects over time in multiple video data 200.
  • the data generation unit 118 converts, for example, information on the three-dimensional position of an object for each frame of multiple video data 200 into edge information.
  • the data generator 118 may convert information on time-series shape changes of objects in multiple video data 200 into still image elements, and generate still image data 300 including the converted still image elements.
  • the data generating unit 118 generates still image data 300 that expresses, for example, information on the time series shape changes of an object in multiple video data 200 using pixel values. For example, the data generating unit 118 converts, for each frame of the multiple video data 200, information on the three-dimensional positions where multiple structural points of the object exist into pixel value information. As a specific example, for one frame of the video data 200, the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of the multiple structural points of the object contained in the frame as pixel values of three consecutive pixels. As a result, for the portion of the still image data 300 corresponding to that frame, the pixel values of three consecutive pixels indicate the X coordinate, Y coordinate, and Z coordinate of the object. The data generating unit 118 can generate still image data 300 corresponding to the video data 200 by similarly assigning pixel values for other frames of the video data 200.
  • the data generation unit 118 generates still image data 300 that expresses, for example, information on the time series shape changes of an object in multiple video data 200 using feature vectors. For example, the data generation unit 118 converts information on the three-dimensional position where an object exists for each frame of multiple video data 200 into feature vector information.
  • the data generation unit 118 may convert information on changes in the positions of multiple objects over time in the video data 200 into text elements, and generate text data 400 including the converted text elements.
  • the data generating unit 118 generates text data 400 that expresses, for example, information on the change in the positions of multiple objects in the video data 200 over time using characters.
  • the data generating unit 118 converts, for example, information on the three-dimensional position of an object for each frame into character information.
  • the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of the multiple objects included in the frame as three consecutive characters.
  • three consecutive characters indicate the X coordinate, Y coordinate, and Z coordinate of the object.
  • the data generating unit 118 can generate text data 400 corresponding to the video data 200 by similarly assigning the characters for other frames of the video data 200.
  • the method of converting the information on the X coordinate, Y coordinate, and Z coordinate of each of the multiple objects into characters is not limited to this.
  • the information on the X coordinate, Y coordinate, and Z coordinate of each of the multiple objects may be compressed and converted into characters using existing compression technology.
  • Frame divisions may be expressed by placement, or one or more characters may be used to indicate the frame division. Other methods for dividing frames may also be used.
  • Japanese characters may be used as the characters. This generates Japanese text data 400 that represents information on changes in the time series positions of multiple objects in the video data 200. If multiple objects in the video data 200 are moving in a similar manner, the text data 400 may have similar content, with similar characters being repeated. Furthermore, if multiple objects in the video data 200 are moving rapidly for a certain period of time, the text data 400 may use different characters in that portion than in the other portions. Characters from a country other than Japan may be used as the characters. Furthermore, characters from multiple countries may be mixed as characters.
  • the data generation unit 118 generates text data 400 that expresses, by words, information on the time-series changes in the positions of multiple objects in the video data 200.
  • the data generation unit 118 converts, for example, information on the three-dimensional positions where objects exist for each frame into word information.
  • the data generation unit 118 generates text data 400 that expresses, for example, information on the time-series changes in the positions of multiple objects in the video data 200 using ASCII code.
  • the data generation unit 118 converts, for example, information on the three-dimensional positions where objects exist for each frame into ASCII code information.
  • the data generation unit 118 generates text data 400 that expresses, for example, information on changes in the positions of multiple objects over time in the video data 200 using at least one of characters and words, and a font.
  • the data generation unit 118 converts, for example, information on the three-dimensional positions where objects exist for each frame into at least one of characters and words, and font information.
  • the data generation unit 118 may convert information about time-series shape changes of objects in the video data 200 into text elements, and generate text data 400 including the converted text elements.
  • the data generation unit 118 generates text data 400 that uses characters to represent, for example, information on the time-series shape changes of an object in the video data 200.
  • the data generation unit 118 converts, for example, information on the three-dimensional positions of multiple structural points of an object for each frame into character information.
  • the data generation unit 118 assigns the X-coordinate, Y-coordinate, and Z-coordinate of each of the multiple structural points of an object included in the frame as three consecutive characters.
  • the three consecutive characters indicate the X-coordinate, Y-coordinate, and Z-coordinate of the object.
  • the data generation unit 118 can generate text data 400 corresponding to the video data 200 by similarly assigning characters for other frames of the video data 200.
  • the method of converting the information on the X-coordinate, Y-coordinate, and Z-coordinate of each of the multiple structural points of an object into characters is not limited to this.
  • the information on the X-coordinate, Y-coordinate, and Z-coordinate of each of the multiple structural points of an object may be compressed and converted into characters using existing compression technology.
  • the frame division may be expressed by arrangement, or one or more characters may be used to indicate the frame division. Methods other than these may be used for the frame division.
  • the data generation unit 118 generates text data 400 that expresses, by words, information on the time-series shape changes of an object in the video data 200, for example.
  • the data generation unit 118 converts, for example, information on the three-dimensional position where an object exists for each frame into word information.
  • the data generation unit 118 generates text data 400 that expresses, for example, information on the time-series shape changes of an object in the video data 200 using ASCII code.
  • the data generation unit 118 converts, for example, information on the three-dimensional position where an object exists for each frame into ASCII code information.
  • the data generation unit 118 generates text data 400 that expresses, for example, information on the time-series shape changes of an object in the video data 200 using at least one of characters and words, and a font.
  • the data generation unit 118 converts, for example, information on the three-dimensional position where an object exists for each frame into at least one of characters and words, and font information.
  • the data generation unit 118 generates text data 400 from multiple video data 200 in which an object is captured from multiple different directions.
  • the data generation unit 118 may convert information on the time-series movement of an object in the multiple video data 200 into text elements, and generate text data 400 including the converted text elements.
  • the data generation unit 118 may generate text data 400 that represents information on the time-series movement of an object in the multiple video data 200 by at least one of characters, words, fonts, and ASCII codes.
  • the data generation unit 118 may convert information on changes in the time series positions of multiple objects in multiple video data 200 into text elements, and generate text data 400 including the converted text elements.
  • the data generating unit 118 generates text data 400 that expresses, for example, information on the change in the positions of multiple objects in multiple video data 200 over time using characters. For each of the multiple video data 200, the data generating unit 118 converts, for example, information on the three-dimensional position of an object for each frame into character information. As a specific example, for one frame of the video data 200, the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of the multiple objects included in the frame as three consecutive characters. As a result, for the portion of the still image data 300 corresponding to that frame, the three consecutive characters indicate the X coordinate, Y coordinate, and Z coordinate of the object. The data generating unit 118 can generate text data 400 corresponding to the video data 200 by similarly assigning characters for other frames of the video data 200.
  • the data generation unit 118 generates text data 400 that expresses, by words, information on the change in the positions of multiple objects over time in multiple video data 200. For example, the data generation unit 118 converts information on the three-dimensional positions where objects exist for each frame of multiple video data 200 into word information.
  • the data generation unit 118 generates text data 400 that expresses, for example, information on the change in the positions of multiple objects over time in multiple video data 200 using ASCII code.
  • the data generation unit 118 converts, for example, information on the three-dimensional position where an object exists for each frame of each of the multiple video data 200 into ASCII code information.
  • the data generator 118 may convert information on time-series shape changes of objects in multiple video data 200 into text elements, and generate text data 400 including the converted text elements.
  • the data generating unit 118 generates text data 400 that expresses, for example, information on the time-series shape changes of an object in multiple video data 200 using characters. For example, the data generating unit 118 converts, for each frame of the multiple video data 200, information on the three-dimensional positions where multiple structural points of the object exist into character information. As a specific example, for one frame of the video data 200, the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of the multiple structural points of the object contained in the frame as three consecutive characters. As a result, for the portion of the text data 400 that corresponds to that frame, the three consecutive characters indicate the X coordinate, Y coordinate, and Z coordinate of the object. The data generating unit 118 can generate text data 400 corresponding to the video data 200 by similarly assigning characters for other frames of the video data 200.
  • the data generation unit 118 generates text data 400 that expresses, for example, information on time-series shape changes of objects in multiple video data 200 using at least one of characters and words and a font. For example, the data generation unit 118 converts information on the three-dimensional position where an object exists for each frame of multiple video data 200 into at least one of characters and words and font information.
  • the data generation unit 118 generates genome data 500 from the video data 200.
  • the data generation unit 118 may convert information on the time series movement of an object in the video data 200 into genome elements, and generate genome data 500 including the converted genome elements.
  • the data generation unit 118 may generate genome data 500 that represents information on the time series movement of an object in the video data 200 by a base sequence.
  • the data generating unit 118 generates genome data 500 that expresses information on the time-series positional changes of multiple objects in the video data 200 using base sequences. For example, the data generating unit 118 converts information on the three-dimensional position of an object for each frame into base sequence information. As a specific example, the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of multiple objects included in one frame of the video data 200 as a base sequence. As a result, for a portion of the genome data 500 corresponding to the frame, consecutive base sequences indicate the X coordinate, Y coordinate, and Z coordinate of the object.
  • the data generating unit 118 can generate genome data 500 corresponding to the video data 200 by similarly assigning the base sequences for other frames of the video data 200.
  • the method of converting the information on the X coordinate, Y coordinate, and Z coordinate of each of multiple objects into a base sequence is not limited to this.
  • the information on the X coordinate, Y coordinate, and Z coordinate of each of multiple objects may be compressed and converted into a base sequence using existing compression technology.
  • Frame divisions may be expressed by arrangement, or a base sequence indicating the frame division may be used. Methods other than these may also be used to indicate frame divisions.
  • the data generation unit 118 may convert information on time-series shape changes of objects in the video data 200 into genome elements, and generate genome data 500 including the converted genome elements.
  • the data generating unit 118 generates genome data 500 that represents, for example, information on the time-series shape change of an object in the video data 200 by a base sequence.
  • the data generating unit 118 converts, for example, information on the three-dimensional positions of the object's structural points for each frame into base sequence information.
  • the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of the structural points of the object included in one frame of the video data 200 as a base sequence.
  • the base sequence indicates the X coordinate, Y coordinate, and Z coordinate of the object for the part of the genome data 500 that corresponds to the frame.
  • the data generating unit 118 can generate genome data 500 corresponding to the video data 200 by similarly assigning the information as base sequences for other frames of the video data 200.
  • the method of converting the information on the X coordinate, Y coordinate, and Z coordinate of each of the structural points of the object into a base sequence is not limited to this.
  • the information on the X coordinate, Y coordinate, and Z coordinate of each of the structural points of the object may be compressed and converted into a base sequence using existing compression technology.
  • Frame divisions may be expressed by arrangement, or a base sequence indicating the frame division may be used. Methods other than these may also be used to indicate frame divisions.
  • the data generation unit 118 generates genome data 500 from multiple video data 200 in which an object is captured from multiple different directions.
  • the data generation unit 118 may convert information on the time series movement of the object in the multiple video data 200 into genome elements, and generate genome data 500 including the converted genome elements.
  • the data generation unit 118 may generate text data 400 that represents information on the time series movement of the object in the multiple video data 200 by a base sequence.
  • the data generation unit 118 may convert information on time-series positional changes of multiple objects in multiple video data 200 into genome elements, and generate genome data 500 including the converted genome elements.
  • the data generating unit 118 generates genome data 500 that expresses, for example, information on the change in the positions of a plurality of objects in a plurality of video data 200 over time by using base sequences. For example, the data generating unit 118 converts, for each of the plurality of video data 200, information on the three-dimensional position where an object exists for each frame into information on base sequences. As a specific example, for one frame of the video data 200, the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of the plurality of objects included in the frame as base sequences. As a result, for a portion of the still image data 300 corresponding to the frame, the base sequences indicate the X coordinate, Y coordinate, and Z coordinate of the object. The data generating unit 118 can generate text data 400 corresponding to the video data 200 by similarly assigning the base sequences to other frames of the video data 200. ⁇ Genome data: multiple videos: shape changes>
  • the data generation unit 118 may convert information on time-series shape changes of objects in multiple video data 200 into genome elements, and generate genome data 500 including the converted genome elements.
  • the data generating unit 118 generates genome data 500 that expresses, for example, information on the time-series shape changes of an object in multiple video data 200 using a base sequence. For example, the data generating unit 118 converts, for each frame of multiple video data 200, information on the three-dimensional positions of multiple structural points of the object into base sequence information. As a specific example, for one frame of the video data 200, the data generating unit 118 assigns the X coordinate, Y coordinate, and Z coordinate of each of the multiple structural points of the object contained in the frame as a base sequence. As a result, for the portion of the genome data 500 corresponding to that frame, three consecutive characters indicate the X coordinate, Y coordinate, and Z coordinate of the object. The data generating unit 118 can generate genome data 500 corresponding to the video data 200 by similarly assigning base sequences for other frames of the video data 200.
  • the classification result acquisition unit 120 inputs the converted data generated by the data generation unit 118 to the classification network stored in the model storage unit 114 and acquires the classification result output from the classification network.
  • the classification result acquisition unit 120 inputs the still image data 300 generated by the classification result acquisition unit 120 to the still image classification network and acquires the classification result output from the still image classification network.
  • the classification result acquisition unit 120 inputs the text data 400 generated by the classification result acquisition unit 120 to the text classification network and acquires the classification result output from the text classification network.
  • the classification result acquisition unit 120 inputs the genome data 500 generated by the classification result acquisition unit 120 to the genome classification network and acquires the classification result output from the genome classification network.
  • the model generation unit 122 generates a classification network that classifies the converted data by executing machine learning using the converted data generated by the data generation unit 118.
  • the model generation unit 122 may store the generated classification network in the model storage unit 114.
  • the model generation unit 122 generates a still image classification network that classifies the still image data by executing machine learning using the still image data 300 generated by the data generation unit 118.
  • the model generation unit 122 generates a text classification network that classifies the text data by executing machine learning using the text data 400 generated by the data generation unit 118.
  • the model generation unit 122 generates a genome classification network that classifies the genome data by executing machine learning using the genome data 500 generated by the data generation unit 118.
  • the output unit 130 performs various output controls.
  • the output unit 130 may have an alert output unit 132, a display output unit 134, and an audio output unit 136.
  • the alert output unit 132 outputs an alert according to the classification result obtained by the classification result acquisition unit 120.
  • the alert output unit 132 may display and output alert information on a display provided in the data processing device 100, output an alert as sound from a speaker provided in the data processing device 100, display and output alert information on a display provided in the communication terminal 30, or output an alert as sound from a speaker provided in the communication terminal 30.
  • the alert output unit 132 outputs an alert when the classification results acquired by the classification result acquisition unit 120 satisfy a predetermined condition. For example, in a situation where the video data acquisition unit 116 is continuously acquiring video data 200, the data generation unit 118 is continuously generating converted data, and the classification result acquisition unit 120 is continuously acquiring classification results, the alert output unit 132 outputs an alert when the degree of difference between successive classification results exceeds a predetermined threshold. Also, for example, after multiple pieces of converted data generated by the data generation unit 118 are accumulated, the classification result acquisition unit 120 acquires classification results collectively for the multiple converted data, and the alert output unit 132 outputs an alert when the classification results include a classification result that indicates an abnormal value.
  • the display output unit 134 may display and output the converted data generated by the data generation unit 118.
  • the display output unit 134 displays and outputs the converted data generated by the data generation unit 118 on a display provided in the data processing device 100.
  • the display output unit 134 displays and outputs the converted data generated by the data generation unit 118 on a display provided in the communication terminal 30.
  • the display output unit 134 displays and outputs the still image data 300 generated by the data generation unit 118.
  • the display output unit 134 may continuously display and output the still image data 300 continuously generated by the data generation unit 118.
  • the display output unit 134 displays and outputs the text data 400 generated by the data generation unit 118.
  • the display output unit 134 may continuously display and output the text data 400 continuously generated by the data generation unit 118.
  • the display output unit 134 displays and outputs the genome data 500 generated by the data generation unit 118.
  • the display output unit 134 may continuously display and output the genome data 500 continuously generated by the data generation unit 118. This makes it possible for a viewer viewing the display to become aware that some abnormality may have occurred in the object.
  • the display output unit 134 may display and output the classification results acquired by the classification result acquisition unit 120.
  • the display output unit 134 displays and outputs the classification results acquired by the classification result acquisition unit 120 on a display provided in the data processing device 100.
  • the display output unit 134 displays and outputs the classification results acquired by the classification result acquisition unit 120 on a display provided in the communication terminal 30.
  • the audio output unit 136 may output the converted data generated by the data generation unit 118 as audio.
  • the display output unit 134 may output, for example, the text data 400 generated by the data generation unit 118 as audio.
  • the audio output unit 136 may output the text data 400 generated by the data generation unit 118 as audio to a speaker provided in the data processing device 100.
  • the audio output unit 136 may output the text data 400 generated by the data generation unit 118 as audio to a speaker provided in the communication terminal 30. This makes it possible for a person who hears the audio to become aware that some abnormality may have occurred in the object.
  • the conversion data acquisition unit 142 acquires conversion data including elements converted from information on the time series movement of an object in the video data.
  • the conversion data acquisition unit 142 acquires, for example, still image data including still image elements converted from information on the time series movement of an object in the video data.
  • the conversion data acquisition unit 142 acquires, for example, text data including text elements converted from information on the time series movement of an object in the video data.
  • the conversion data acquisition unit 142 acquires, for example, genome data including genome elements converted from information on the time series movement of an object in the video data.
  • the video data generation unit 144 generates video data from the converted data acquired by the converted data acquisition unit 142.
  • the video data generation unit 144 When the converted data acquisition unit 142 acquires the converted data generated by the data generation unit 118, the video data generation unit 144 generates video data from the converted data using the conversion information stored in the conversion information storage unit 112 and used to generate the converted data.
  • the video data generation unit 144 generates video data 200 from, for example, still image data 300 generated by the data generation unit 118.
  • the video data generation unit 144 generates video data 200 from, for example, text data 400 generated by the data generation unit 118.
  • the video data generation unit 144 generates video data 200 from, for example, genome data 500 generated by the data generation unit 118.
  • the converted data acquisition unit 142 may also acquire conversion information used by the other device when generating the converted data.
  • the video data generation unit 144 may use the conversion information to generate video data from the converted data. For example, the video data generation unit 144 generates video data from still image data generated by another device. For example, the video data generation unit 144 generates video data from text data generated by another device. For example, the video data generation unit 144 generates video data from genome data generated by another device.
  • the display output unit 134 may display and output the video data generated by the video data generation unit 144.
  • the display output unit 134 displays and outputs the video data generated by the video data generation unit 144 on a display provided in the data processing device 100.
  • the display output unit 134 displays and outputs the video data generated by the video data generation unit 144 on a display provided in the communication terminal 30.
  • the data processing device 100 includes all of the conversion information storage unit 112, the model storage unit 114, the video data acquisition unit 116, the data generation unit 118, the classification result acquisition unit 120, the model generation unit 122, the output unit 130, the conversion data acquisition unit 142, and the video data generation unit 144.
  • the data processing device 100 may have a function of generating conversion data from video data, but may not have a function of generating video data from conversion data. In this case, the data processing device 100 may not have the conversion data acquisition unit 142 and the video data generation unit 144. Also, the data processing device 100 may have a function of generating video data from conversion data, but may not have a function of generating conversion data from video data. In this case, the data processing device 100 may not have the model storage unit 114, the video data acquisition unit 116, the data generation unit 118, the classification result acquisition unit 120, and the model generation unit 122.
  • FIG. 9 shows a schematic diagram of an example of a hardware configuration of a computer 1200 functioning as a data processing device 100.
  • a program installed on the computer 1200 can cause the computer 1200 to function as one or more "parts" of an apparatus according to the present embodiment, or to execute operations or one or more "parts” associated with an apparatus according to the present embodiment, and/or to execute a process or steps of a process according to the present embodiment.
  • Such a program can be executed by the CPU 1212 to cause the computer 1200 to execute specific operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.
  • the computer 1200 includes a CPU 1212, a RAM 1214, and a graphics controller 1216, which are interconnected by a host controller 1210.
  • the computer 1200 also includes input/output units such as a communication interface 1222, a storage device 1224, a DVD drive, and an IC card drive, which are connected to the host controller 1210 via an input/output controller 1220.
  • the DVD drive may be a DVD-ROM drive, a DVD-RAM drive, or the like.
  • the storage device 1224 may be a hard disk drive, a solid state drive, or the like.
  • the computer 1200 also includes a ROM 1230 and a legacy input/output unit such as a keyboard, which are connected to the input/output controller 1220 via an input/output chip 1240.
  • the CPU 1212 operates according to the programs stored in the ROM 1230 and the RAM 1214, thereby controlling each unit.
  • the graphics controller 1216 acquires image data generated by the CPU 1212 into a frame buffer or the like provided in the RAM 1214 or into itself, and causes the image data to be displayed on the display device 1218.
  • the communication interface 1222 communicates with other electronic devices via a network.
  • the storage device 1224 stores programs and data used by the CPU 1212 in the computer 1200.
  • the DVD drive reads programs or data from a DVD-ROM or the like and provides them to the storage device 1224.
  • the IC card drive reads programs and data from an IC card and/or writes programs and data to an IC card.
  • ROM 1230 stores therein a boot program or the like to be executed by computer 1200 upon activation, and/or a program that depends on the hardware of computer 1200.
  • I/O chip 1240 may also connect various I/O units to I/O controller 1220 via USB ports, parallel ports, serial ports, keyboard ports, mouse ports, etc.
  • the programs are provided by a computer-readable storage medium such as a DVD-ROM or an IC card.
  • the programs are read from the computer-readable storage medium, installed in storage device 1224, RAM 1214, or ROM 1230, which are also examples of computer-readable storage media, and executed by CPU 1212.
  • the information processing described in these programs is read by computer 1200, and brings about cooperation between the programs and the various types of hardware resources described above.
  • An apparatus or method may be constructed by realizing the operation or processing of information according to the use of computer 1200.
  • the CPU 1212 may also cause all or a necessary portion of a file or database stored in an external recording medium such as the storage device 1224, a DVD drive (DVD-ROM), an IC card, etc. to be read into the RAM 1214, and perform various types of processing on the data on the RAM 1214. The CPU 1212 may then write back the processed data to the external recording medium.
  • an external recording medium such as the storage device 1224, a DVD drive (DVD-ROM), an IC card, etc.
  • CPU 1212 may perform various types of processing on data read from RAM 1214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search/replacement, etc., as described throughout this disclosure and specified by the instruction sequence of the program, and write back the results to RAM 1214.
  • CPU 1212 may also search for information in a file, database, etc. in the recording medium.
  • CPU 1212 may search for an entry whose attribute value of the first attribute matches a specified condition from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.
  • the above-described programs or software modules may be stored in a computer-readable storage medium on the computer 1200 or in the vicinity of the computer 1200.
  • a recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can be used as a computer-readable storage medium, thereby providing the programs to the computer 1200 via the network.
  • the blocks in the flowcharts and block diagrams in this embodiment may represent stages of a process in which an operation is performed or "parts" of a device responsible for performing the operation. Particular stages and “parts" may be implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable storage medium, and/or a processor provided with computer-readable instructions stored on a computer-readable storage medium.
  • the dedicated circuitry may include digital and/or analog hardware circuitry and may include integrated circuits (ICs) and/or discrete circuits.
  • the programmable circuitry may include reconfigurable hardware circuitry including AND, OR, XOR, NAND, NOR, and other logical operations, flip-flops, registers, and memory elements, such as, for example, field programmable gate arrays (FPGAs) and programmable logic arrays (PLAs).
  • FPGAs field programmable gate arrays
  • PDAs programmable logic arrays
  • a computer-readable storage medium may include any tangible device capable of storing instructions that are executed by a suitable device, such that a computer-readable storage medium having instructions stored thereon comprises an article of manufacture that includes instructions that can be executed to create means for performing the operations specified in the flowchart or block diagram.
  • Examples of computer-readable storage media may include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, and the like.
  • Computer-readable storage media may include floppy disks, diskettes, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), electrically erasable programmable read-only memories (EEPROMs), static random access memories (SRAMs), compact disk read-only memories (CD-ROMs), digital versatile disks (DVDs), Blu-ray disks, memory sticks, integrated circuit cards, and the like.
  • RAMs random access memories
  • ROMs read-only memories
  • EPROMs or flash memories erasable programmable read-only memories
  • EEPROMs electrically erasable programmable read-only memories
  • SRAMs static random access memories
  • CD-ROMs compact disk read-only memories
  • DVDs digital versatile disks
  • Blu-ray disks memory sticks, integrated circuit cards, and the like.
  • the computer readable instructions may include either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk (registered trademark), JAVA (registered trademark), C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages.
  • ISA instruction set architecture
  • machine instructions machine-dependent instructions
  • microcode firmware instructions
  • state setting data or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk (registered trademark), JAVA (registered trademark), C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages.
  • the computer-readable instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, or to a programmable circuit, either locally or over a local area network (LAN), a wide area network (WAN) such as the Internet, so that the processor of the general-purpose computer, special-purpose computer, or other programmable data processing apparatus, or to a programmable circuit, executes the computer-readable instructions to generate means for performing the operations specified in the flowcharts or block diagrams.
  • processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Environmental Sciences (AREA)
  • Zoology (AREA)
  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Biodiversity & Conservation Biology (AREA)
  • Animal Husbandry (AREA)
  • Marine Sciences & Fisheries (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Theoretical Computer Science (AREA)
  • Image Analysis (AREA)
  • Farming Of Fish And Shellfish (AREA)

Abstract

オブジェクトを撮像した動画データを取得する動画データ取得部と、前記動画データにおける前記オブジェクトの時系列の動きの情報を静止画の要素に変換して、変換した前記静止画の前記要素を含む静止画データを生成するデータ生成部とを備える、データ処理装置を提供する。前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の動きの情報を、画素値、特徴ベクトル、及びエッジの少なくともいずれかによって表す前記静止画データを生成してよい。

Description

データ処理装置及びプログラム
 本発明は、データ処理装置及びプログラムに関する。
 特許文献1には、魚の多様な泳動を既存の楽曲に対応する楽譜を用いて再現したトレーニングデータを自動で生成する手法について記載されている。
 [先行技術文献]
 [特許文献]
 [特許文献1]特許第7152535号
一般的開示
 本発明の一実施態様によれば、データ処理装置が提供される。前記データ処理装置は、オブジェクトを撮像した動画データを取得する動画データ取得部を備えてよい。前記データ処理装置は、前記動画データにおける前記オブジェクトの時系列の動きの情報を静止画の要素に変換して、変換した前記静止画の前記要素を含む静止画データを生成するデータ生成部を備えてよい。
 前記データ処理装置において、前記データ生成部は、前記動画データの複数のフレームのうちの1つのフレームから、一つの静止画データを構成する複数の部分のうちの1つの部分を生成すること、及び、前記動画データの複数のフレームのうちの複数のフレームから、前記一つの静止画データを構成する前記複数の部分のうちの1つの部分を生成することの少なくともいずれかを実行することによって、前記複数の部分からなる前記一つの静止画データを生成してよい。前記いずれかのデータ処理装置において、前記データ生成部は、前記動画データの前記複数のフレームについて、前記フレームに含まれる複数の前記オブジェクトが存在する3次元位置の情報を、画素値の情報に変換することによって、前記複数の部分を生成してよい。前記いずれかのデータ処理装置において、前記動画データ取得部は、前記オブジェクトを複数の異なる方向から撮像する複数のカメラによって並行して撮像された複数の前記動画データを取得してよく、前記データ生成部は、前記複数の動画データ毎に、フレーム毎の、複数のオブジェクトが存在する3次元位置の情報を、画素値の情報に変換することによって、前記複数の部分を生成してよい。前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の動きの情報を、画素値、特徴ベクトル、及びエッジの少なくともいずれかによって表す前記静止画データを生成してよい。
 前記いずれかのデータ処理装置において、前記データ生成部は、前記動画データにおける複数の前記オブジェクトの時系列の位置の変化の情報を前記静止画の前記要素に変換して、変換した前記静止画の前記要素を含む前記静止画データを生成してよい。前記データ生成部は、前記動画データにおける複数のオブジェクトの時系列の位置の変化の情報を、画素値によって表す前記静止画データを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトが存在する3次元位置の情報を、画素値の情報に変換してよい。前記データ生成部は、前記動画データにおける複数のオブジェクトの時系列の位置の変化の情報を、特徴ベクトルによって表す前記静止画データを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトが存在する3次元位置の情報を、特徴ベクトルの情報に変換してよい。前記データ生成部は、前記動画データにおける複数のオブジェクトの時系列の位置の変化の情報を、エッジによって表す前記静止画データを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトが存在する3次元位置の情報を、エッジの情報に変換してよい。
 前記いずれかのデータ処理装置において、前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の形状変化の情報を前記静止画の前記要素に変換して、変換した前記静止画の前記要素を含む前記静止画データを生成してよい。前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の形状変化の情報を、画素値によって表す前記静止画データを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、画素値の情報に変換してよい。前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の形状変化の情報を、特徴ベクトルによって表す前記静止画データを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、特徴ベクトルの情報に変換してよい。前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の形状変化の情報を、エッジによって表す前記静止画データを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、エッジの情報に変換してよい。
 前記いずれかのデータ処理装置において、前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の動きの情報をテキストの要素に変換して、変換した前記テキストの前記要素を含むテキストデータを生成してよい。
 前記いずれかのデータ処理装置において、前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の動きの情報をゲノムの要素に変換して、変換した前記ゲノムの前記要素を含むゲノムデータを生成してよい。
 前記いずれかのデータ処理装置は、前記データ生成部によって生成された前記静止画データを、静止画データを分類するように学習された静止画分類ネットワークに入力して、前記静止画分類ネットワークから出力された分類結果を取得する分類結果取得部と、前記分類結果取得部が取得した前記分類結果が予め定められた条件を満たす場合に、アラートを出力するアラート出力部とを更に備えてよい。
 前記いずれかのデータ処理装置は、動画データにおけるオブジェクトの時系列の動きの情報から変換された静止画の要素を含む静止画データを取得する変換データ取得部と、前記静止画データから前記動画データを生成する動画データ生成部とを更に備えてよい。
 本発明の一実施態様によれば、データ処理装置が提供される。前記データ処理装置は、動画データにおけるオブジェクトの時系列の動きの情報から変換された静止画の要素を含む静止画データを取得する静止画データ取得部を備えてよい。前記データ処理装置は、前記静止画データから前記動画データを生成する動画データ生成部を備えてよい。
 本発明の一実施態様によれば、データ処理装置が提供される。前記データ処理装置は、オブジェクトを撮像した動画データを取得する動画データ取得部を備えてよい。前記データ処理装置は、前記動画データにおける前記オブジェクトの時系列の動きの情報をテキストの要素に変換して、変換した前記テキストの前記要素を含むテキストデータを生成するデータ生成部を備えてよい。前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の動きの情報を、文字、単語、フォント、及びアスキーコードの少なくともいずれかによって表す前記テキストデータを生成してよい。前記データ生成部は、前記動画データにおける複数のオブジェクトの時系列の位置の変化の情報をテキストの要素に変換して、変換したテキストの要素を含む前記テキストデータを生成してよい。前記データ生成部は、前記動画データにおける複数のオブジェクトの時系列の位置の変化の情報を、文字によって表す前記テキストデータを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトが存在する3次元位置の情報を、文字の情報に変換してよい。前記データ生成部は、前記動画データにおける複数のオブジェクトの時系列の位置の変化の情報を、単語によって表す前記テキストデータを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトが存在する3次元位置の情報を、単語の情報に変換してよい。前記データ生成部は、前記動画データにおける複数のオブジェクトの時系列の位置の変化の情報を、アスキーコードによって表す前記テキストデータを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトが存在する3次元位置の情報を、アスキーコードの情報に変換してよい。前記データ生成部は、前記動画データにおける複数のオブジェクトの時系列の位置の変化の情報を、文字及び単語の少なくともいずれかと、フォントとによって表す前記テキストデータを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトが存在する3次元位置の情報を、文字及び単語の少なくともいずれかと、フォントの情報に変換してよい。前記データ生成部は、前記動画データにおけるオブジェクトの時系列の形状変化の情報をテキストの要素に変換して、変換したテキストの要素を含む前記テキストデータを生成してよい。前記データ生成部は、前記動画データにおけるオブジェクトの時系列の形状変化の情報を、文字によって表す前記テキストデータを生成してよい。前記データ生成部は、前記動画データにおけるオブジェクトの時系列の形状変化の情報を、アスキーコードによって表す前記テキストデータを生成してよい。前記データ生成部は、前記動画データにおけるオブジェクトの時系列の形状変化の情報を、文字及び単語の少なくともいずれかと、フォントとによって表す前記テキストデータを生成してよい。前記データ生成部は、オブジェクトを複数の異なる方向から撮像した複数の動画データから、前記テキストデータを生成してよい。前記データ生成部は、前記複数の動画データにおけるオブジェクトの時系列の動きの情報をテキストの要素に変換して、変換したテキストの要素を含む前記テキストデータを生成してよい。前記データ処理装置は、前記データ生成部によって生成された前記テキストデータを、静止画分類ネットワークに入力して、前記静止画分類ネットワークから出力された分類結果を取得する分類結果取得部を備えてよい。前記データ処理装置は、前記分類結果取得部による分類結果に応じたアラートを出力するアラート出力部を備えてよい。
 本発明の一実施態様によれば、データ処理装置が提供される。前記データ処理装置は、オブジェクトを撮像した動画データを取得する動画データ取得部を備えてよい。前記データ処理装置は、前記動画データにおける前記オブジェクトの時系列の動きの情報をゲノムの要素に変換して、変換した前記ゲノムの前記要素を含むゲノムデータを生成するデータ生成部を備えてよい。前記データ生成部は、前記動画データにおける複数のオブジェクトの時系列の位置の変化の情報を、塩基配列によって表す前記ゲノムデータを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトが存在する3次元位置の情報を、塩基配列の情報に変換してよい。前記データ生成部は、前記動画データにおけるオブジェクトの時系列の形状変化の情報をゲノムの要素に変換して、変換したゲノムの要素を含む前記ゲノムデータを生成してよい。前記データ生成部は、前記動画データにおけるオブジェクトの時系列の形状変化の情報を、塩基配列によって表す前記ゲノムデータを生成してよい。前記データ生成部は、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、塩基配列の情報に変換してよい。前記データ生成部は、オブジェクトを複数の異なる方向から撮像した複数の動画データから、前記ゲノムデータを生成してよい。前記データ生成部は、前記複数の動画データにおけるオブジェクトの時系列の動きの情報をゲノムの要素に変換して、変換したゲノムの要素を含む前記ゲノムデータを生成してよい。前記データ生成部は、前記複数の動画データにおけるオブジェクトの時系列の動きの情報を、塩基配列によって表す前記ゲノムデータを生成してよい。前記データ処理装置は、前記データ生成部によって生成された前記ゲノムデータを、ゲノム分類ネットワークに入力して、前記ゲノム分類ネットワークから出力された分類結果を取得する分類結果取得部を備えてよい。前記データ処理装置は、前記分類結果取得部による分類結果に応じたアラートを出力するアラート出力部を備えてよい。
 本発明の一実施態様によれば、コンピュータを、前記データ処理装置として機能させるためのプログラムが提供される。
 なお、上記の発明の概要は、本発明の必要な特徴の全てを列挙したものではない。また、これらの特徴群のサブコンビネーションもまた、発明となりうる。
データ処理装置100の一例を概略的に示す。 動画データ200を静止画データ300に変換する変換処理について説明するための説明図である。 複数の動画データ200を静止画データ300に変換する変換処理について説明するための説明図である。 複数の動画データ200を静止画データ300に変換する変換処理について説明するための説明図である。 動画データ200をテキストデータ400に変換する変換処理について説明するための説明図である。 動画データ200をテキストデータ400に変換する変換処理について説明するための説明図である。 動画データ200をゲノムデータ500に変換する変換処理について説明するための説明図である。 データ処理装置100の機能構成の一例を概略的に示す。 データ処理装置100として機能するコンピュータ1200のハードウェア構成の一例を概略的に示す。
 以下、発明の実施の形態を通じて本発明を説明するが、以下の実施形態は請求の範囲にかかる発明を限定するものではない。また、実施形態の中で説明されている特徴の組み合わせの全てが発明の解決手段に必須であるとは限らない。
 魚の養殖において、魚の状態を把握することは、無駄な餌、大量死、病気等を防ぐために、非常に重要である。しかし、養殖場を四六時中監視することは困難であり、養殖場を撮影した動画を自動的に分析して分析結果を提供するモニタリングシステムが望まれる。このようなモニタリングシステムは、魚に限らず、任意のオブジェクトに対して必要とされている。このようなモニタリングシステムは、3次元空間と時間とを有する4次元の情報を含む動画データを分析する必要があるが、動画データをそのまま分析することは、処理負荷が非常に高い。それに対して、本実施形態に係るデータ処理装置100においては、動画データを、より低い次元のデータに変換する。例えば、データ処理装置100は、動画データを、2次元データに変換する。
 具体例として、データ処理装置100は、動画データを、静止画データに変換する。データ処理装置100は、動画データに含まれる情報を、静止画の要素に変換することによって、動画データを静止画データに変換してよい。例えば、データ処理装置100は、動画データに対するオブジェクト検出を実行し、オブジェクトの時系列の動きの情報を静止画の要素に変換することによって、動画データを静止画データに変換する。データ処理装置100は、変換ルールを含む変換情報に従って、動画データを静止画データに変換してよい。データ処理装置100は、変換した静止画データを分析することによってオブジェクトのモニタリングを実現してよい。これにより、処理負荷を低減することができる。また、静止画データを分析するニューラルネットワークは非常に幅広く研究されていて、多くのニューラルネットワークがすでに存在しており、そのようなニューラルネットワークを流用することを可能とすることができる。
 具体例として、データ処理装置100は、動画データを、テキストデータに変換する。データ処理装置100は、動画データに含まれる情報を、テキストの要素に変換することによって、動画データをテキストデータに変換してよい。例えば、データ処理装置100は、動画データに対するオブジェクト検出を実行し、オブジェクトの時系列の動きの情報をテキストの要素に変換することによって、動画データをテキストデータに変換する。データ処理装置100は、変換ルールを含む変換情報に従って、動画データをテキストデータに変換してよい。データ処理装置100は、変換したテキストデータを分析することによってオブジェクトのモニタリングを実現してよい。これにより、処理負荷を低減することができる。また、テキストデータを分析するニューラルネットワークは非常に幅広く研究されていて、多くのニューラルネットワークがすでに存在しており、そのようなニューラルネットワークを流用することを可能とすることができる。
 データ処理装置100は、動画データを、特定の文字を羅列したデータや、数字を羅列したデータに変換してもよい。データ処理装置100は、例えば、動画データを、ゲノムデータに変換する。ゲノムデータとは、塩基配列によって構成されるデータであってよい。データ処理装置100は、動画データに含まれる情報を、ゲノムの要素に変換することによって、動画データをゲノムデータに変換してよい。例えば、データ処理装置100は、動画データに対するオブジェクト検出を実行し、オブジェクトの時系列の動きの情報をゲノムの要素に変換することによって、動画データをゲノムデータに変換する。データ処理装置100は、変換ルールを含む変換情報に従って、動画データをゲノムデータに変換してよい。データ処理装置100は、変換したゲノムデータを分析することによってオブジェクトのモニタリングを実現してよい。これにより、処理負荷を低減することができる。また、ゲノムデータを分析するニューラルネットワークは非常に幅広く研究されていて、多くのニューラルネットワークがすでに存在しており、そのようなニューラルネットワークを流用することを可能とすることができる。
 図1は、データ処理装置100の一例を概略的に示す。データ処理装置100は、オブジェクトを撮像した動画データを取得して、取得した動画データを、当該動画データよりも次元の低いデータに変換してよい(変換後のデータを変換データと記載する場合がある。)。動画データよりも次元の低いデータの例として、静止画データ、テキストデータ、及びゲノムデータ等が挙げられる。
 オブジェクトは、時系列に変化する任意のものであってよい。オブジェクトは、時系列に位置が変化するものであってよく、時系列に形状が変化するものであってよく、時系列に位置及び形状が変化するものであってよい。オブジェクトは、時系列に形状は変化せずに位置のみが変化するものであってもよい。オブジェクトは、時系列に位置は変化せずに形状のみが変化するものであってもよい。
 オブジェクトは、モニタリングの対象となるものであってよい。本実施形態では、オブジェクトが魚であり、モニタリング対象が魚群であるケースを主に例に挙げて説明するが、これに限られない。
 データ処理装置100は、カメラ102によって撮像された動画データを取得してよい。カメラ102は、データ処理装置100に内蔵されていてよい。カメラ102は、データ処理装置100に外付けされてもよい。カメラ102は、任意のネットワークを介してデータ処理装置100に接続されてもよい。
 データ処理装置100は、通信端末30から動画データを受信してもよい。通信端末30は、例えば、カメラ32によって撮像された動画データをデータ処理装置100に送信する。カメラ32は、通信端末30に内蔵されていてよい。カメラ32は、通信端末30に外付けされてもよい。カメラ32は、任意のネットワークを介して通信端末30に接続されてもよい。
 通信端末30は、PC(Personal Computer)、タブレット端末、及びスマートフォン等であってよい。データ処理装置100と通信端末30とは、ネットワーク20を介して通信してよい。ネットワーク20は、インターネットを含んでよい。ネットワーク20は、LAN(Local Area Network)を含んでよい。ネットワーク20は、移動体通信ネットワークを含んでよい。移動体通信ネットワークは、5G(5th Generation)通信方式、LTE(Long Term Evolution)通信方式、3G(3rd Generation)通信方式、及び6G(6th Generation)通信方式以降の通信方式のいずれに準拠していてもよい。
 データ処理装置100は、変換データを分類するように学習された分類ネットワークを予め記憶しておいてよい。データ処理装置100は、分類ネットワークに変換データを入力して、分類ネットワークから出力された分類結果を取得してよい。データ処理装置100は、分類結果が予め定められた条件を満たす場合に、アラートを出力してよい。
 具体例として、データ処理装置100は、養殖場の魚群を撮像した動画データを取得して、動画データにおける複数の魚の時系列の動きの情報を静止画の要素に変換して、変換した静止画の要素を含む静止画データを生成する。データ処理装置100は、例えば、動画データにおける複数の魚の時系列の位置の変化の情報を静止画の要素に変換して、変換した静止画の要素を含む静止画データを生成する。データ処理装置100は、秒単位、分単位、時間単位、日単位、週単位、月単位、及び年単位等のように、任意の時間単位毎に静止画データを生成してよい。
 例えば、データ処理装置100は、1時間単位で静止画データを生成する。この場合、1つの静止画データには、1時間の間の複数の魚の時系列の位置の変化の特性が含まれることになる。例えば、5時間分の動画データから、5個の静止画データが生成されることになる。5時間の間、複数の魚に特に異常がなく、正常に泳いでいた場合には、5つの静止画データは、互いに類似した画像となる。それに対して、複数の魚が、4時間は正常に泳いでいたが、最後の1時間に、それまでとは異なる激しい動きをした場合、最初の4時間に対応する4つの静止画データは互いに類似した画像となるが、最後の1時間に対応する1つの静止画データは、4つの静止画データとは異なる特徴を示す画像となり得る。すなわち、静止画データを分類することによって、魚群の状態の変化をモニタリングすることができる。
 データ処理装置100は、生成した多数の静止画データを用いた機械学習を実行することによって、静止画データを分類する静止画分類ネットワークを生成してよい。データ処理装置100は、生成した静止画分類ネットワークに、新たに生成した静止画データを入力することによって、当該静止画データを分類してよい。
 データ処理装置100は、既存の静止画分類ネットワークを用いてもよい。例えば、データ処理装置100は、他の装置によって生成された静止画分類ネットワークを予め取得して記憶しておく。そして、データ処理装置100は、以降に生成した静止画データを当該静止画分類ネットワークに入力することによって、当該静止画データを分類する。
 データ処理装置100は、変換データから動画データを生成する機能を更に備えてもよい。データ処理装置100は、自身が生成した変換データから、動画データを生成してよい。データ処理装置100は、動画データを変換データに変換する際に用いた変換情報を利用して、変換データを動画データに変換してよい。
 データ処理装置100は、他の装置が生成した変換データから、動画データを生成してもよい。データ処理装置100は、変換データを生成した他の装置から、動画データを変換データに変換する際に用いた変換情報を受信しておき、当該変換情報を利用して、他の装置から受信した変換データを動画データに変換してよい。
 データ処理装置100は、動画データから変換データを生成する機能は有さずに、変換データから動画データを生成する機能を有してもよい。
 図2は、動画データ200を静止画データ300に変換する変換処理について説明するための説明図である。ここでは、カメラ102が、養殖場40の魚群を撮像した動画データ200を、静止画データ300に変換するケースについて説明する。
 データ処理装置100は、カメラ102から動画データ200を取得する。データ処理装置100は、動画データ200における魚群の時系列の動きの情報を静止画の要素に変換して、変換した静止画の要素を含む静止画データ300を生成する。
 例えば、データ処理装置100は、動画データ200の1つのフレームから、静止画データ300を構成する1つの部分302を生成する。データ処理装置100は、それぞれが動画データ200の複数のフレームのそれぞれに対応する複数の部分302を生成することによって、複数の部分302からなる静止画データ300を生成してよい。なお、データ処理装置100は、動画データ200の複数のフレームから、1つの部分302を生成してもよい。
 データ処理装置100は、任意の期間毎に静止画データ300を生成してよい。例えば、1日毎に静止画データ300を生成する場合、データ処理装置100は、1日分の動画データ200のフレームから、複数の部分302を生成することによって、1日に対応する静止画データ300を生成する。
 データ処理装置100は、複数のカメラによって並行して撮像された複数の動画データ200から、静止画データ300を生成してもよい。
 図3及び図4は、複数の動画データ200を静止画データ300に変換する変換処理について説明するための説明図である。ここでは、6方向から養殖場40の魚群を撮像した6つの動画データ200を静止画データ300に変換するケースについて説明する。図3、図4に示す例においては、上、下、及び横の4方向(横A、横B、横C、横D)から養殖場40の魚群が撮像されている。
 図3では、あるタイミングにおける、上から撮像した動画データ210のフレーム211、下から撮像した動画データ220のフレーム221、横Aから撮像した動画データ230のフレーム231、横Bから撮像した動画データ240のフレーム241、横Cから撮像した動画データ250のフレーム251、横Dから撮像した動画データ260のフレーム261を例示している。
 データ処理装置100は、例えば、フレーム211から部分311を生成し、フレーム221から部分321を生成し、フレーム231から部分331を生成し、フレーム241から部分341を生成し、フレーム251から部分351を生成し、フレーム261から部分361を生成する。データ処理装置100は、図4に示すように、これらを連結する。データ処理装置100は、次のフレームについても、同様に静止画の部分を生成して、連結する。これを繰り返すことによって、データ処理装置100は、静止画データ300を生成する。
 なお、データ処理装置100は、上述したように、1つのフレームから1つの部分を生成するのではなく、複数の連続するフレームから1つの部分を生成してもよい。例えば、データ処理装置100は、上から撮像した動画データ210の連続する10フレームから1つの部分311を生成し、下から撮像した動画データ220の連続する10フレームから1つの部分321を生成し、横Aから撮像した動画データ230の連続する10フレームから1つの部分331を生成し、横Bから撮像した動画データ240の連続する10フレームから1つの部分341を生成し、横Cから撮像した動画データ250の連続する10フレームから1つの部分351を生成し、横Dから撮像した動画データ260の連続する10フレームから1つの部分361を生成する。10フレームというのは一例であり、フレームの数は他の数であってもよい。
 図5は、動画データ200をテキストデータ400に変換する変換処理について説明するための説明図である。ここでは、カメラ102が、養殖場40の魚群を撮像した動画データ200を、テキストデータ400に変換するケースについて説明する。
 データ処理装置100は、カメラ102から動画データ200を取得する。データ処理装置100は、動画データ200における魚群の時系列の動きの情報をテキストの要素に変換して、変換したテキストの要素を含むテキストデータ400を生成する。
 例えば、データ処理装置100は、動画データ200の1つのフレームから、テキストデータ400を構成する1つの部分402を生成する。データ処理装置100は、それぞれが動画データ200の複数のフレームのそれぞれに対応する複数の部分402を生成することによって、複数の部分402からなるテキストデータ400を生成してよい。なお、データ処理装置100は、動画データ200の複数のフレームから、1つの部分402を生成してもよい。
 データ処理装置100は、任意の期間毎にテキストデータ400を生成してよい。例えば、1日毎にテキストデータ400を生成する場合、データ処理装置100は、1日分の動画データ200のフレームから、複数の部分402を生成することによって、1日に対応するテキストデータ400を生成する。
 データ処理装置100は、複数のカメラによって並行して撮像された複数の動画データ200から、テキストデータ400を生成してよい。ここでは、図3に示すように、6方向から養殖場40の魚群を撮像した6つの動画データ200をテキストデータ400に変換するケースについて説明する。
 データ処理装置100は、例えば、フレーム211から部分411を生成し、フレーム221から部分421を生成し、フレーム231から部分431を生成し、フレーム241から部分441を生成し、フレーム251から部分451を生成し、フレーム261から部分461を生成する。データ処理装置100は、図6に示すように、これらを連結する。データ処理装置100は、次のフレームについても、同様にテキストの部分を生成して、連結する。これを繰り返すことによって、データ処理装置100は、テキストデータ400を生成する。
 なお、データ処理装置100は、上述したように、1つのフレームから1つの部分を生成するのではなく、複数の連続するフレームから1つの部分を生成してもよい。例えば、データ処理装置100は、上から撮像した動画データ210の連続する10フレームから1つの部分411を生成し、下から撮像した動画データ220の連続する10フレームから1つの部分421を生成し、横Aから撮像した動画データ230の連続する10フレームから1つの部分431を生成し、横Bから撮像した動画データ240の連続する10フレームから1つの部分441を生成し、横Cから撮像した動画データ250の連続する10フレームから1つの部分451を生成し、横Dから撮像した動画データ260の連続する10フレームから1つの部分461を生成する。10フレームというのは一例であり、フレームの数は他の数であってもよい。
 図7は、動画データ200をゲノムデータ500に変換する変換処理について説明するための説明図である。ここでは、カメラ102が、養殖場40の魚群を撮像した動画データ200を、ゲノムデータ500に変換するケースについて説明する。
 データ処理装置100は、カメラ102から動画データ200を取得する。データ処理装置100は、動画データ200における魚群の時系列の動きの情報をゲノムの要素に変換して、変換したゲノムの要素を含むゲノムデータ500を生成する。
 例えば、データ処理装置100は、動画データ200の1つのフレームから、ゲノムデータ500を構成する1つの部分を生成する。データ処理装置100は、それぞれが動画データ200の複数のフレームのそれぞれに対応する複数の部分を生成することによって、複数の部分からなるゲノムデータ500を生成してよい。なお、データ処理装置100は、動画データ200の複数のフレームから、1つの部分402を生成してもよい。
 データ処理装置100は、任意の期間毎にゲノムデータ500を生成してよい。例えば、1日毎にゲノムデータ500を生成する場合、データ処理装置100は、1日分の動画データ200のフレームから、複数の部分を生成することによって、1日に対応するゲノムデータ500を生成する。
 データ処理装置100は、静止画データ300及びテキストデータ400と同様に、複数のカメラによって並行して撮像された複数の動画データ200から、ゲノムデータ500を生成してよい。
 図8は、データ処理装置100の機能構成の一例を概略的に示す。データ処理装置100は、変換情報記憶部112、モデル記憶部114、動画データ取得部116、データ生成部118、分類結果取得部120、モデル生成部122、出力部130、変換データ取得部142、及び動画データ生成部144を備える。
 変換情報記憶部112は、動画データ200を変換するための変換ルールを含む変換情報を記憶する。変換情報記憶部112は、予め登録された変換情報を記憶してよい。例えば、変換情報記憶部112は、動画データ200を静止画データ300に変換するための変換ルールを含む変換情報を記憶する。例えば、変換情報記憶部112は、動画データ200をテキストデータ400に変換するための変換ルールを含む変換情報を記憶する。例えば、変換情報記憶部112は、動画データ200をゲノムデータ500に変換するための変換ルールを含む変換情報を記憶する。
 モデル記憶部114は、変換データを分類するように学習された分類ネットワークを記憶する。モデル記憶部114は、予め登録された分類ネットワークを記憶してよい。モデル記憶部114は、例えば、他の装置によって生成された分類ネットワークを記憶する。例えば、変換情報記憶部112は、静止画分類ネットワークを記憶する。例えば、変換情報記憶部112は、テキスト分類ネットワークを記憶する。例えば、変換情報記憶部112は、ゲノム分類ネットワークを記憶する。
 動画データ取得部116は、オブジェクトを撮像した動画データ200を取得する。動画データ取得部116は、変換対象の動画データ200を取得する。動画データ取得部116は、カメラ102から動画データ200を受信してよい。動画データ取得部116は、通信端末30から動画データ200を受信してよい。動画データ取得部116は、他の装置から動画データ200を受信してもよい。動画データ取得部116は、可搬型の記憶デバイス等から、動画データ200を取得してもよい。
 動画データ200には、オブジェクトの検出結果が付帯されてもよい。動画データ200には、オブジェクトの3次元位置の情報が付帯されてもよい。オブジェクトの3次元位置の情報は、カメラ102やカメラ32とは別途配置された測距センサによる出力を用いる方法や、カメラ102やカメラ32の画像解析結果を用いる方法等によって特定され得る。なお、これらに限られず、オブジェクトの3次元位置を特定することができれば、どのような手法が用いられてもよい。
 データ生成部118は、動画データ取得部116が取得した動画データ200を変換して、変換データを生成する。データ生成部118は、変換情報記憶部112に記憶されている変換情報を用いて、動画データ200から変換データを生成してよい。
 例えば、データ生成部118は、動画データ200から、静止画データ300を生成する。データ生成部118は、動画データ200におけるオブジェクトの時系列の動きの情報を静止画の要素に変換して、変換した静止画の要素を含む静止画データ300を生成してよい。データ生成部118は、動画データ200におけるオブジェクトの時系列の動きの情報を、画素値、特徴ベクトル、及びエッジの少なくともいずれかによって表す静止画データ300を生成してよい。このように、静止画の要素の例として、画素値、特徴ベクトル、及びエッジが挙げられるが、これらに限られない。
 データ生成部118は、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を静止画の要素に変換して、変換した静止画の要素を含む静止画データ300を生成してよい。
 データ生成部118は、例えば、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、画素値によって表す静止画データ300を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、画素値の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれる複数のオブジェクトのそれぞれのX座標、Y座標、Z座標を、連続する3つのピクセルの画素値として割り当てていく。これにより、静止画データ300における当該フレームに対応する部分について、連続する3つのピクセルの画素値が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に画素値として割り当てていくことによって、動画データ200に対応する静止画データ300を生成し得る。複数のオブジェクトのそれぞれのX座標、Y座標、Z座標の情報を画素値に変換する方法は、これに限られない。既存の圧縮技術を用いて、複数のオブジェクトのそれぞれのX座標、Y座標、Z座標の情報を圧縮して、画素値に変換してもよい。フレームの区切りについては、配置によって表現してよく、フレームの区切りを示す1又は複数の画素値を用いてもよい。フレームの区切りについて、これら以外の方法を用いてもよい。
 画素値として、グレースケールが用いられてよい。これにより、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を表す、グレースケールの静止画データ300が生成されることになる。当該静止画データ300は、動画データ200において、複数のオブジェクトが同じような動きをしている場合には、同じような模様が繰り返される見た目となり得る。また、当該静止画データ300は、動画データ200において、複数のオブジェクトが一時的に激しい動きをしている場合には、その部分が濃くなったり、薄くなったり、変化が激しくなったりすることによって、その部分が他の部分と異なることを示す見た目となり得る。
 画素値として、カラーが用いられてもよい。これにより、グレースケールを用いる場合と比較して、複数のオブジェクトの時系列の位置の変化の情報を効率的に表現することができる。カラーを用いた場合、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を表す、カラーの静止画データ300が生成されることになる。当該静止画データ300は、動画データ200において、複数のオブジェクトが同じような動きをしている場合には、同じような色合いとなり得る。また、当該静止画データ300は、動画データ200において、複数のオブジェクトが一時的に激しい動きをしている場合には、その部分の色彩が濃くなったり、薄くなったり、色合いの変化が激しくなったりすることによって、その部分が他の部分と異なることを示す見た目となり得る。
 データ生成部118は、例えば、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、特徴ベクトルによって表す静止画データ300を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、特徴ベクトルの情報に変換する。データ生成部118は、特徴ベクトルを表現する方法として、位置座標、速度変化量、及び方向を数値化したベクトル等を用いてもよく、特徴そのものを学習する方法を用いてもよい。魚群等のように、対象となるオブジェクトの数が多い場合には、複数のオブジェクトの時系列の位置について、統計をとった後に、特徴ベクトルを生成してもよい。また、特徴ベクトルから統計をとって、それらをまた上位層の特徴ベクトルとしてもよい。
 データ生成部118は、例えば、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、エッジによって表す静止画データ300を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、エッジの情報に変換する。データ生成部118は、例えば、2値化して微分をとる方法、キャニー法、DNN(Deep Neural Network)によってエッジの学習やシルエットの学習をする方法等を用いて、複数のオブジェクトの時系列の位置の変化の情報と、エッジの情報との間の変換を実現してよい。
 データ生成部118は、動画データ200におけるオブジェクトの時系列の形状変化の情報を静止画の要素に変換して、変換した静止画の要素を含む静止画データ300を生成してもよい。
 データ生成部118は、例えば、動画データ200におけるオブジェクトの時系列の形状変化の情報を、画素値によって表す静止画データ300を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、画素値の情報に変換する。例えば、オブジェクトが生物やロボット等である場合、オブジェクトの構造点とは、関節の位置であってよい。オブジェクトが関節を有さないものである場合、オブジェクトの構造点は、オブジェクトの動きに影響を与えるオブジェクトの各ポイントであってよい。
 具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれるオブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標を、連続する3つのピクセルの画素値として割り当てていく。これにより、静止画データ300における当該フレームに対応する部分について、連続する3つのピクセルの画素値が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に画素値として割り当てていくことによって、動画データ200に対応する静止画データ300を生成し得る。複数のオブジェクトのそれぞれのX座標、Y座標、Z座標の情報を画素値に変換する方法は、これに限られない。既存の圧縮技術を用いて、複数のオブジェクトのそれぞれのX座標、Y座標、Z座標の情報を圧縮して、画素値に変換してもよい。フレームの区切りについては、配置によって表現してよく、フレームの区切りを示す1又は複数の画素値を用いてもよい。フレームの区切りについて、これら以外の方法を用いてもよい。
 画素値として、グレースケールが用いられてよい。これにより、動画データ200におけるオブジェクトの時系列の形状変化の情報を表す、グレースケールの静止画データ300が生成されることになる。当該静止画データ300は、動画データ200において、オブジェクトが同じような動きをしている場合には、同じような模様が繰り返される見た目となり得る。また、当該静止画データ300は、動画データ200において、オブジェクトが一時的に激しい動きをしている場合には、その部分が濃くなったり、薄くなったり、変化が激しくなったりすることによって、その部分が他の部分と異なることを示す見た目となり得る。
 画素値として、カラーが用いられてもよい。これにより、グレースケールを用いる場合と比較して、オブジェクトの時系列の形状変化の情報を効率的に表現することができる。カラーを用いた場合、動画データ200におけるオブジェクトの時系列の形状変化の情報を表す、カラーの静止画データ300が生成されることになる。当該静止画データ300は、動画データ200において、オブジェクトが同じような動きをしている場合には、同じような色合いとなり得る。また、当該静止画データ300は、動画データ200において、オブジェクトが一時的に激しい動きをしている場合には、その部分の色彩が濃くなったり、薄くなったり、色合いの変化が激しくなったりすることによって、その部分が他の部分と異なることを示す見た目となり得る。
 データ生成部118は、例えば、動画データ200におけるオブジェクトの時系列の形状変化の情報を、特徴ベクトルによって表す静止画データ300を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、特徴ベクトルの情報に変換する。
 データ生成部118は、例えば、動画データ200におけるオブジェクトの時系列の形状変化の情報を、エッジによって表す静止画データ300を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、エッジの情報に変換する。
 例えば、データ生成部118は、オブジェクトを複数の異なる方向から撮像した複数の動画データ200から、静止画データ300を生成する。ここでは、1つの動画データ200から静止画データ300を生成する場合と異なる点について主に説明する。データ生成部118は、複数の動画データ200におけるオブジェクトの時系列の動きの情報を静止画の要素に変換して、変換した静止画の要素を含む静止画データ300を生成してよい。データ生成部118は、複数の動画データ200におけるオブジェクトの時系列の動きの情報を、画素値、特徴ベクトル、及びエッジの少なくともいずれかによって表す静止画データ300を生成してよい。
 データ生成部118は、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報を静止画の要素に変換して、変換した静止画の要素を含む静止画データ300を生成してよい。
 データ生成部118は、例えば、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、画素値によって表す静止画データ300を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、画素値の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれる複数のオブジェクトのそれぞれのX座標、Y座標、Z座標を、連続する3つのピクセルの画素値として割り当てていく。これにより、静止画データ300における当該フレームに対応する部分について、連続する3つのピクセルの画素値が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に画素値として割り当てていくことによって、動画データ200に対応する静止画データ300を生成し得る。
 データ生成部118は、例えば、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、特徴ベクトルによって表す静止画データ300を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、特徴ベクトルの情報に変換する。
 データ生成部118は、例えば、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、エッジによって表す静止画データ300を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、エッジの情報に変換する。
 データ生成部118は、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報を静止画の要素に変換して、変換した静止画の要素を含む静止画データ300を生成してもよい。
 データ生成部118は、例えば、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報を、画素値によって表す静止画データ300を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、画素値の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれるオブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標を、連続する3つのピクセルの画素値として割り当てていく。これにより、静止画データ300における当該フレームに対応する部分について、連続する3つのピクセルの画素値が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に画素値として割り当てていくことによって、動画データ200に対応する静止画データ300を生成し得る。
 データ生成部118は、例えば、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報を、特徴ベクトルによって表す静止画データ300を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、特徴ベクトルの情報に変換する。
 データ生成部118は、例えば、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報を、エッジによって表す静止画データ300を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、エッジの情報に変換する。
 例えば、データ生成部118は、動画データ200から、テキストデータ400を生成する。データ生成部118は、動画データ200におけるオブジェクトの時系列の動きの情報をテキストの要素に変換して、変換したテキストの要素を含むテキストデータ400を生成してよい。データ生成部118は、動画データ200におけるオブジェクトの時系列の動きの情報を、文字、単語、フォント、及びアスキーコードの少なくともいずれかによって表すテキストデータ400を生成してよい。このように、テキストの要素の例として、文字、単語、フォント、及びアスキーコードが挙げられるが、これらに限られない。
 データ生成部118は、動画データ200における複数のオブジェクトの時系列の位置の変化の情報をテキストの要素に変換して、変換したテキストの要素を含むテキストデータ400を生成してよい。
 データ生成部118は、例えば、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、文字によって表すテキストデータ400を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、文字の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれる複数のオブジェクトのそれぞれのX座標、Y座標、Z座標を、連続する3つの文字として割り当てていく。これにより、静止画データ300における当該フレームに対応する部分について、連続する3つの文字が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に文字として割り当てていくことによって、動画データ200に対応するテキストデータ400を生成し得る。複数のオブジェクトのそれぞれのX座標、Y座標、Z座標の情報を文字に変換する方法は、これに限られない。既存の圧縮技術を用いて、複数のオブジェクトのそれぞれのX座標、Y座標、Z座標の情報を圧縮して、文字に変換してもよい。フレームの区切りについては、配置によって表現してよく、フレームの区切りを示す1又は複数の文字を用いてもよい。フレームの区切りについて、これら以外の方法を用いてもよい。
 文字として、日本語の文字が用いられてよい。これにより、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を表す、日本語のテキストデータ400が生成されることになる。当該テキストデータ400は、動画データ200において、複数のオブジェクトが同じような動きをしている場合には、同じような文字が繰り返される、同じような内容となり得る。また、当該テキストデータ400は、動画データ200において、複数のオブジェクトが一時的に激しい動きをしている場合には、その部分だけ、他の部分と用いられる文字がことなることになり得る。文字として、日本以外の国の文字が用いられても良い。また、文字として、複数の国の文字が混在してもよい。
 データ生成部118は、例えば、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、単語によって表すテキストデータ400を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、単語の情報に変換する。
 データ生成部118は、例えば、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、アスキーコードによって表すテキストデータ400を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、アスキーコードの情報に変換する。
 データ生成部118は、例えば、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、文字及び単語の少なくともいずれかと、フォントとによって表すテキストデータ400を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、文字及び単語の少なくともいずれかと、フォントの情報に変換する。
 データ生成部118は、動画データ200におけるオブジェクトの時系列の形状変化の情報をテキストの要素に変換して、変換したテキストの要素を含むテキストデータ400を生成してもよい。
 データ生成部118は、例えば、動画データ200におけるオブジェクトの時系列の形状変化の情報を、文字によって表すテキストデータ400を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、文字の情報に変換する。
 具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれるオブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標を、連続する3つの文字として割り当てていく。これにより、テキストデータ400における当該フレームに対応する部分について、連続する3つの文字が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に文字として割り当てていくことによって、動画データ200に対応するテキストデータ400を生成し得る。オブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標の情報を文字に変換する方法は、これに限られない。既存の圧縮技術を用いて、オブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標の情報を圧縮して、文字に変換してもよい。フレームの区切りについては、配置によって表現してよく、フレームの区切りを示す1又は複数の文字を用いてもよい。フレームの区切りについて、これら以外の方法を用いてもよい。
 文字として、日本語の文字が用いられてよい。これにより、動画データ200におけるオブジェクトの時系列の形状変化の情報を表す、日本語のテキストデータ400が生成されることになる。当該テキストデータ400は、動画データ200において、オブジェクトが同じような動きをしている場合には、同じような文字が繰り返される、同じような内容となり得る。また、当該テキストデータ400は、動画データ200において、オブジェクトが一時的に激しい動きをしている場合には、その部分だけ、他の部分と用いられる文字が異なることになり得る。文字として、日本以外の国の文字が用いられても良い。また、文字として、複数の国の文字が混在してもよい。
 データ生成部118は、例えば、動画データ200におけるオブジェクトの時系列の形状変化の情報を、単語によって表すテキストデータ400を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、単語の情報に変換する。
 データ生成部118は、例えば、動画データ200におけるオブジェクトの時系列の形状変化の情報を、アスキーコードによって表すテキストデータ400を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、アスキーコードの情報に変換する。
 データ生成部118は、例えば、動画データ200におけるオブジェクトの時系列の形状変化の情報を、文字及び単語の少なくともいずれかと、フォントとによって表すテキストデータ400を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、文字及び単語の少なくともいずれかと、フォントの情報に変換する。
 例えば、データ生成部118は、オブジェクトを複数の異なる方向から撮像した複数の動画データ200から、テキストデータ400を生成する。ここでは、1つの動画データ200からテキストデータ400を生成する場合と異なる点について主に説明する。データ生成部118は、複数の動画データ200におけるオブジェクトの時系列の動きの情報をテキストの要素に変換して、変換したテキストの要素を含むテキストデータ400を生成してよい。データ生成部118は、複数の動画データ200におけるオブジェクトの時系列の動きの情報を、文字、単語、フォント、及びアスキーコードの少なくともいずれかによって表すテキストデータ400を生成してよい。
 データ生成部118は、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報をテキストの要素に変換して、変換したテキストの要素を含むテキストデータ400を生成してよい。
 データ生成部118は、例えば、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、文字によって表すテキストデータ400を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、文字の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれる複数のオブジェクトのそれぞれのX座標、Y座標、Z座標を、連続する3つの文字として割り当てていく。これにより、静止画データ300における当該フレームに対応する部分について、連続する3つの文字が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に文字として割り当てていくことによって、動画データ200に対応するテキストデータ400を生成し得る。
 データ生成部118は、例えば、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、単語によって表すテキストデータ400を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、単語の情報に変換する。
 データ生成部118は、例えば、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、アスキーコードによって表すテキストデータ400を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、アスキーコードの情報に変換する。
 データ生成部118は、例えば、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、文字及び単語の少なくともいずれかと、フォントとによって表すテキストデータ400を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、文字及び単語の少なくともいずれかと、フォントの情報に変換する。
 データ生成部118は、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報をテキストの要素に変換して、変換したテキストの要素を含むテキストデータ400を生成してもよい。
 データ生成部118は、例えば、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報を、文字によって表すテキストデータ400を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、文字の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれるオブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標を、連続する3つの文字として割り当てていく。これにより、テキストデータ400における当該フレームに対応する部分について、連続する3つの文字が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に文字として割り当てていくことによって、動画データ200に対応するテキストデータ400を生成し得る。
 データ生成部118は、例えば、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報を、単語によって表すテキストデータ400を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、単語の情報に変換する。
 データ生成部118は、例えば、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報を、アスキーコードによって表すテキストデータ400を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、アスキーコードの情報に変換する。
 データ生成部118は、例えば、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報を、文字及び単語の少なくともいずれかと、フォントとによって表すテキストデータ400を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、文字及び単語の少なくともいずれかと、フォントの情報に変換する。
 例えば、データ生成部118は、動画データ200から、ゲノムデータ500を生成する。データ生成部118は、動画データ200におけるオブジェクトの時系列の動きの情報をゲノムの要素に変換して、変換したゲノムの要素を含むゲノムデータ500を生成してよい。データ生成部118は、動画データ200におけるオブジェクトの時系列の動きの情報を、塩基配列によって表すゲノムデータ500を生成してよい。
 データ生成部118は、動画データ200における複数のオブジェクトの時系列の位置の変化の情報をゲノムの要素に変換して、変換したゲノムの要素を含むゲノムデータ500を生成してよい。
 データ生成部118は、例えば、動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、塩基配列によって表すゲノムデータ500を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトが存在する3次元位置の情報を、塩基配列の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれる複数のオブジェクトのそれぞれのX座標、Y座標、Z座標を、塩基配列として割り当てていく。これにより、ゲノムデータ500における当該フレームに対応する部分について、連続する塩基配列が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に塩基配列として割り当てていくことによって、動画データ200に対応するゲノムデータ500を生成し得る。複数のオブジェクトのそれぞれのX座標、Y座標、Z座標の情報を塩基配列に変換する方法は、これに限られない。既存の圧縮技術を用いて、複数のオブジェクトのそれぞれのX座標、Y座標、Z座標の情報を圧縮して、塩基配列に変換してもよい。フレームの区切りについては、配置によって表現してよく、フレームの区切りを示す塩基配列を用いてもよい。フレームの区切りについて、これら以外の方法を用いてもよい。
 データ生成部118は、動画データ200におけるオブジェクトの時系列の形状変化の情報をゲノムの要素に変換して、変換したゲノムの要素を含むゲノムデータ500を生成してもよい。
 データ生成部118は、例えば、動画データ200におけるオブジェクトの時系列の形状変化の情報を、塩基配列によって表すゲノムデータ500を生成する。データ生成部118は、例えば、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、塩基配列の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれるオブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標を、塩基配列として割り当てていく。これにより、ゲノムデータ500における当該フレームに対応する部分について、塩基配列が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に塩基配列として割り当てていくことによって、動画データ200に対応するゲノムデータ500を生成し得る。オブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標の情報を塩基配列に変換する方法は、これに限られない。既存の圧縮技術を用いて、オブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標の情報を圧縮して、塩基配列に変換してもよい。フレームの区切りについては、配置によって表現してよく、フレームの区切りを示す塩基配列を用いてもよい。フレームの区切りについて、これら以外の方法を用いてもよい。
 例えば、データ生成部118は、オブジェクトを複数の異なる方向から撮像した複数の動画データ200から、ゲノムデータ500を生成する。ここでは、1つの動画データ200からゲノムデータ500を生成する場合と異なる点について主に説明する。データ生成部118は、複数の動画データ200におけるオブジェクトの時系列の動きの情報をゲノムの要素に変換して、変換したゲノムの要素を含むゲノムデータ500を生成してよい。データ生成部118は、複数の動画データ200におけるオブジェクトの時系列の動きの情報を、塩基配列によって表すテキストデータ400を生成してよい。
 データ生成部118は、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報をゲノムの要素に変換して、変換したゲノムの要素を含むゲノムデータ500を生成してよい。
 データ生成部118は、例えば、複数の動画データ200における複数のオブジェクトの時系列の位置の変化の情報を、塩基配列によって表すゲノムデータ500を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトが存在する3次元位置の情報を、塩基配列の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれる複数のオブジェクトのそれぞれのX座標、Y座標、Z座標を、塩基配列として割り当てていく。これにより、静止画データ300における当該フレームに対応する部分について、塩基配列が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に塩基配列として割り当てていくことによって、動画データ200に対応するテキストデータ400を生成し得る。
<ゲノムデータ:複数動画:形状変化>
 データ生成部118は、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報をゲノムの要素に変換して、変換したゲノムの要素を含むゲノムデータ500を生成してもよい。
 データ生成部118は、例えば、複数の動画データ200におけるオブジェクトの時系列の形状変化の情報を、塩基配列によって表すゲノムデータ500を生成する。データ生成部118は、例えば、複数の動画データ200毎に、フレーム毎の、オブジェクトの複数の構造点が存在する3次元位置の情報を、塩基配列の情報に変換する。具体例として、データ生成部118は、動画データ200の1つのフレームについて、フレームに含まれるオブジェクトの複数の構造点のそれぞれのX座標、Y座標、Z座標を、塩基配列として割り当てていく。これにより、ゲノムデータ500における当該フレームに対応する部分について、連続する3つの文字が、オブジェクトのX座標、Y座標、Z座標を示すことになる。データ生成部118は、動画データ200の他のフレームについて、同様に塩基配列として割り当てていくことによって、動画データ200に対応するゲノムデータ500を生成し得る。
 分類結果取得部120は、データ生成部118によって生成された変換データを、モデル記憶部114に記憶されている分類ネットワークに入力して、分類ネットワークから出力された分類結果を取得する。分類結果取得部120は、例えば、分類結果取得部120によって生成された静止画データ300を、静止画分類ネットワークに入力して、静止画分類ネットワークから出力された分類結果を取得する。分類結果取得部120は、例えば、分類結果取得部120によって生成されたテキストデータ400を、テキスト分類ネットワークに入力して、テキスト分類ネットワークから出力された分類結果を取得する。分類結果取得部120は、例えば、分類結果取得部120によって生成されたゲノムデータ500を、ゲノム分類ネットワークに入力して、ゲノム分類ネットワークから出力された分類結果を取得する。
 モデル生成部122は、データ生成部118によって生成された変換データを用いた機械学習を実行することによって、変換データを分類する分類ネットワークを生成する。モデル生成部122は、生成した分類ネットワークを、モデル記憶部114に記憶させてよい。モデル生成部122は、例えば、データ生成部118によって生成された静止画データ300を用いた機械学習を実行することによって、静止画データを分類する静止画分類ネットワークを生成する。モデル生成部122は、例えば、データ生成部118によって生成されたテキストデータ400を用いた機械学習を実行することによって、テキストデータを分類するテキスト分類ネットワークを生成する。モデル生成部122は、例えば、データ生成部118によって生成されたゲノムデータ500を用いた機械学習を実行することによって、ゲノムデータを分類するゲノム分類ネットワークを生成する。
 出力部130は、各種出力制御を実行する。出力部130は、アラート出力部132、表示出力部134、及び音声出力部136を有してよい。
 アラート出力部132は、分類結果取得部120による分類結果に応じたアラートを出力する。アラート出力部132は、データ処理装置100が備えるディスプレイにアラート情報を表示出力させたり、データ処理装置100が備えるスピーカにアラートを音声出力させたり、通信端末30が備えるディスプレイにアラート情報を表示出力させたり、通信端末30が備えるスピーカにアラートを音声出力させたりしてよい。
 例えば、アラート出力部132は、分類結果取得部120が取得した分類結果が予め定められた条件を満たす場合に、アラートを出力する。例えば、動画データ取得部116が継続的に動画データ200を取得しており、データ生成部118が継続的に変換データを生成しており、分類結果取得部120が継続的に分類結果を取得している状況において、アラート出力部132は、連続する分類結果の相違度が予め定められた閾値を超えたことに応じて、アラートを出力する。また、例えば、データ生成部118によって生成された変換データが複数蓄積された後、分類結果取得部120が、複数の変換データについて、まとめて分類結果を取得した場合において、アラート出力部132は、分類結果の中に、異常値を示す分類結果が含まれていた場合に、アラートを出力する。
 表示出力部134は、データ生成部118によって生成された変換データを表示出力させてよい。表示出力部134は、例えば、データ生成部118によって生成された変換データをデータ処理装置100が備えるディスプレイに表示出力させる。表示出力部134は、例えば、データ生成部118によって生成された変換データを通信端末30が備えるディスプレイに表示出力させる。
 例えば、表示出力部134は、データ生成部118によって生成された静止画データ300を表示出力させる。表示出力部134は、データ生成部118によって継続的に生成された静止画データ300を、継続的に表示出力させてもよい。例えば、表示出力部134は、データ生成部118によって生成されたテキストデータ400を表示出力させる。表示出力部134は、データ生成部118によって継続的に生成されたテキストデータ400を、継続的に表示出力させてもよい。例えば、表示出力部134は、データ生成部118によって生成されたゲノムデータ500を表示出力させる。表示出力部134は、データ生成部118によって継続的に生成されたゲノムデータ500を、継続的に表示出力させてもよい。これらによって、表示を閲覧した閲覧者に、オブジェクトに何らかの異常が発生した可能性があることに気づかせることができる。
 表示出力部134は、分類結果取得部120によって取得された分類結果を表示出力させてよい。表示出力部134は、例えば、分類結果取得部120によって取得された分類結果をデータ処理装置100が備えるディスプレイに表示出力させる。表示出力部134は、例えば、分類結果取得部120によって取得された分類結果を通信端末30が備えるディスプレイに表示出力させる。
 音声出力部136は、データ生成部118によって生成された変換データを音声出力させてよい。表示出力部134は、例えば、データ生成部118によって生成されたテキストデータ400を音声出力させる。音声出力部136は、データ生成部118によって生成されたテキストデータ400をデータ処理装置100が備えるスピーカに音声出力させてよい。音声出力部136は、データ生成部118によって生成されたテキストデータ400を通信端末30が備えるスピーカに音声出力させてよい。これらによって、音声を聞いた者に、オブジェクトに何らかの異常が発生した可能性があることに気づかせることができる。
 変換データ取得部142は、動画データにおけるオブジェクトの時系列の動きの情報から変換された要素を含む変換データを取得する。変換データ取得部142は、例えば、動画データにおけるオブジェクトの時系列の動きの情報から変換された静止画の要素を含む静止画データを取得する。変換データ取得部142は、例えば、動画データにおけるオブジェクトの時系列の動きの情報から変換されたテキストの要素を含むテキストデータを取得する。変換データ取得部142は、例えば、動画データにおけるオブジェクトの時系列の動きの情報から変換されたゲノムの要素を含むゲノムデータを取得する。
 動画データ生成部144は、変換データ取得部142が取得した変換データから動画データを生成する。動画データ生成部144は、データ生成部118によって生成された変換データを変換データ取得部142が取得した場合、変換情報記憶部112に記憶されている、当該変換データの生成に用いられた変換情報を用いて、変換データから動画データを生成する。動画データ生成部144は、例えば、データ生成部118によって生成された静止画データ300から動画データ200を生成する。動画データ生成部144は、例えば、データ生成部118によって生成されたテキストデータ400から動画データ200を生成する。動画データ生成部144は、例えば、データ生成部118によって生成されたゲノムデータ500から動画データ200を生成する。
 他の装置によって生成された変換データを取得した場合、変換データ取得部142が、他の装置が当該変換データを生成する際に用いた変換情報を合わせて取得してよい。動画データ生成部144は、当該変換情報を用いて、当該変換データから動画データを生成してよい。動画データ生成部144は、例えば、他の装置によって生成された静止画データから動画データを生成する。動画データ生成部144は、例えば、他の装置によって生成されたテキストデータから動画データを生成する。動画データ生成部144は、例えば、他の装置によって生成されたゲノムデータから動画データを生成する。
 表示出力部134は、動画データ生成部144によって生成された動画データを表示出力させてよい。例えば、表示出力部134は、動画データ生成部144によって生成された動画データをデータ処理装置100が備えるディスプレイに表示出力させる。例えば、表示出力部134は、動画データ生成部144によって生成された動画データを通信端末30が備えるディスプレイに表示出力させる。
 なお、データ処理装置100が変換情報記憶部112、モデル記憶部114、動画データ取得部116、データ生成部118、分類結果取得部120、モデル生成部122、出力部130、変換データ取得部142、及び動画データ生成部144のすべてを備えることは必須とは限らない。上述した通り、データ処理装置100は、動画データから変換データを生成する機能を有して、変換データから動画データを生成する機能を有さなくてもよい。この場合、データ処理装置100は、変換データ取得部142及び動画データ生成部144を備えなくてよい。また、データ処理装置100は、変換データから動画データを生成する機能を有して、動画データから変換データを生成する機能を有さなくてもよい。この場合、データ処理装置100は、モデル記憶部114、動画データ取得部116、データ生成部118、分類結果取得部120、及びモデル生成部122を備えなくてよい。
 図9は、データ処理装置100として機能するコンピュータ1200のハードウェア構成の一例を概略的に示す。コンピュータ1200にインストールされたプログラムは、コンピュータ1200を、本実施形態に係る装置の1又は複数の「部」として機能させ、又はコンピュータ1200に、本実施形態に係る装置に関連付けられるオペレーション又は当該1又は複数の「部」を実行させることができ、及び/又はコンピュータ1200に、本実施形態に係るプロセス又は当該プロセスの段階を実行させることができる。そのようなプログラムは、コンピュータ1200に、本明細書に記載のフローチャート及びブロック図のブロックのうちのいくつか又はすべてに関連付けられた特定のオペレーションを実行させるべく、CPU1212によって実行されてよい。
 本実施形態によるコンピュータ1200は、CPU1212、RAM1214、及びグラフィックコントローラ1216を含み、それらはホストコントローラ1210によって相互に接続されている。コンピュータ1200はまた、通信インタフェース1222、記憶装置1224、DVDドライブ、及びICカードドライブのような入出力ユニットを含み、それらは入出力コントローラ1220を介してホストコントローラ1210に接続されている。DVDドライブは、DVD-ROMドライブ及びDVD-RAMドライブ等であってよい。記憶装置1224は、ハードディスクドライブ及びソリッドステートドライブ等であってよい。コンピュータ1200はまた、ROM1230及びキーボードのようなレガシの入出力ユニットを含み、それらは入出力チップ1240を介して入出力コントローラ1220に接続されている。
 CPU1212は、ROM1230及びRAM1214内に格納されたプログラムに従い動作し、それにより各ユニットを制御する。グラフィックコントローラ1216は、RAM1214内に提供されるフレームバッファ等又はそれ自体の中に、CPU1212によって生成されるイメージデータを取得し、イメージデータがディスプレイデバイス1218上に表示されるようにする。
 通信インタフェース1222は、ネットワークを介して他の電子デバイスと通信する。記憶装置1224は、コンピュータ1200内のCPU1212によって使用されるプログラム及びデータを格納する。DVDドライブは、プログラム又はデータをDVD-ROM等から読み取り、記憶装置1224に提供する。ICカードドライブは、プログラム及びデータをICカードから読み取り、及び/又はプログラム及びデータをICカードに書き込む。
 ROM1230はその中に、アクティブ化時にコンピュータ1200によって実行されるブートプログラム等、及び/又はコンピュータ1200のハードウェアに依存するプログラムを格納する。入出力チップ1240はまた、様々な入出力ユニットをUSBポート、パラレルポート、シリアルポート、キーボードポート、マウスポート等を介して、入出力コントローラ1220に接続してよい。
 プログラムは、DVD-ROM又はICカードのようなコンピュータ可読記憶媒体によって提供される。プログラムは、コンピュータ可読記憶媒体から読み取られ、コンピュータ可読記憶媒体の例でもある記憶装置1224、RAM1214、又はROM1230にインストールされ、CPU1212によって実行される。これらのプログラム内に記述される情報処理は、コンピュータ1200に読み取られ、プログラムと、上記様々なタイプのハードウェアリソースとの間の連携をもたらす。装置又は方法が、コンピュータ1200の使用に従い情報のオペレーション又は処理を実現することによって構成されてよい。
 例えば、通信がコンピュータ1200及び外部デバイス間で実行される場合、CPU1212は、RAM1214にロードされた通信プログラムを実行し、通信プログラムに記述された処理に基づいて、通信インタフェース1222に対し、通信処理を命令してよい。通信インタフェース1222は、CPU1212の制御の下、RAM1214、記憶装置1224、DVD-ROM、又はICカードのような記録媒体内に提供される送信バッファ領域に格納された送信データを読み取り、読み取られた送信データをネットワークに送信し、又はネットワークから受信した受信データを記録媒体上に提供される受信バッファ領域等に書き込む。
 また、CPU1212は、記憶装置1224、DVDドライブ(DVD-ROM)、ICカード等のような外部記録媒体に格納されたファイル又はデータベースの全部又は必要な部分がRAM1214に読み取られるようにし、RAM1214上のデータに対し様々なタイプの処理を実行してよい。CPU1212は次に、処理されたデータを外部記録媒体にライトバックしてよい。
 様々なタイプのプログラム、データ、テーブル、及びデータベースのような様々なタイプの情報が記録媒体に格納され、情報処理を受けてよい。CPU1212は、RAM1214から読み取られたデータに対し、本開示の随所に記載され、プログラムの命令シーケンスによって指定される様々なタイプのオペレーション、情報処理、条件判断、条件分岐、無条件分岐、情報の検索/置換等を含む、様々なタイプの処理を実行してよく、結果をRAM1214に対しライトバックする。また、CPU1212は、記録媒体内のファイル、データベース等における情報を検索してよい。例えば、各々が第2の属性の属性値に関連付けられた第1の属性の属性値を有する複数のエントリが記録媒体内に格納される場合、CPU1212は、当該複数のエントリの中から、第1の属性の属性値が指定されている条件に一致するエントリを検索し、当該エントリ内に格納された第2の属性の属性値を読み取り、それにより予め定められた条件を満たす第1の属性に関連付けられた第2の属性の属性値を取得してよい。
 上で説明したプログラム又はソフトウエアモジュールは、コンピュータ1200上又はコンピュータ1200近傍のコンピュータ可読記憶媒体に格納されてよい。また、専用通信ネットワーク又はインターネットに接続されたサーバシステム内に提供されるハードディスク又はRAMのような記録媒体が、コンピュータ可読記憶媒体として使用可能であり、それによりプログラムを、ネットワークを介してコンピュータ1200に提供する。
 本実施形態におけるフローチャート及びブロック図におけるブロックは、オペレーションが実行されるプロセスの段階又はオペレーションを実行する役割を持つ装置の「部」を表わしてよい。特定の段階及び「部」が、専用回路、コンピュータ可読記憶媒体上に格納されるコンピュータ可読命令と共に供給されるプログラマブル回路、及び/又はコンピュータ可読記憶媒体上に格納されるコンピュータ可読命令と共に供給されるプロセッサによって実装されてよい。専用回路は、デジタル及び/又はアナログハードウェア回路を含んでよく、集積回路(IC)及び/又はディスクリート回路を含んでよい。プログラマブル回路は、例えば、フィールドプログラマブルゲートアレイ(FPGA)、及びプログラマブルロジックアレイ(PLA)等のような、論理積、論理和、排他的論理和、否定論理積、否定論理和、及び他の論理演算、フリップフロップ、レジスタ、並びにメモリエレメントを含む、再構成可能なハードウェア回路を含んでよい。
 コンピュータ可読記憶媒体は、適切なデバイスによって実行される命令を格納可能な任意の有形なデバイスを含んでよく、その結果、そこに格納される命令を有するコンピュータ可読記憶媒体は、フローチャート又はブロック図で指定されたオペレーションを実行するための手段を作成すべく実行され得る命令を含む、製品を備えることになる。コンピュータ可読記憶媒体の例としては、電子記憶媒体、磁気記憶媒体、光記憶媒体、電磁記憶媒体、半導体記憶媒体等が含まれてよい。コンピュータ可読記憶媒体のより具体的な例としては、フロッピー(登録商標)ディスク、ディスケット、ハードディスク、ランダムアクセスメモリ(RAM)、リードオンリメモリ(ROM)、消去可能プログラマブルリードオンリメモリ(EPROM又はフラッシュメモリ)、電気的消去可能プログラマブルリードオンリメモリ(EEPROM)、静的ランダムアクセスメモリ(SRAM)、コンパクトディスクリードオンリメモリ(CD-ROM)、デジタル多用途ディスク(DVD)、ブルーレイ(登録商標)ディスク、メモリスティック、集積回路カード等が含まれてよい。
 コンピュータ可読命令は、アセンブラ命令、命令セットアーキテクチャ(ISA)命令、マシン命令、マシン依存命令、マイクロコード、ファームウェア命令、状態設定データ、又はSmalltalk(登録商標)、JAVA(登録商標)、C++等のようなオブジェクト指向プログラミング言語、及び「C」プログラミング言語又は同様のプログラミング言語のような従来の手続型プログラミング言語を含む、1又は複数のプログラミング言語の任意の組み合わせで記述されたソースコード又はオブジェクトコードのいずれかを含んでよい。
 コンピュータ可読命令は、汎用コンピュータ、特殊目的のコンピュータ、若しくは他のプログラム可能なデータ処理装置のプロセッサ、又はプログラマブル回路が、フローチャート又はブロック図で指定されたオペレーションを実行するための手段を生成するために当該コンピュータ可読命令を実行すべく、ローカルに又はローカルエリアネットワーク(LAN)、インターネット等のようなワイドエリアネットワーク(WAN)を介して、汎用コンピュータ、特殊目的のコンピュータ、若しくは他のプログラム可能なデータ処理装置のプロセッサ、又はプログラマブル回路に提供されてよい。プロセッサの例としては、コンピュータプロセッサ、処理ユニット、マイクロプロセッサ、デジタル信号プロセッサ、コントローラ、マイクロコントローラ等を含む。
 以上、本発明を実施の形態を用いて説明したが、本発明の技術的範囲は上記実施の形態に記載の範囲には限定されない。上記実施の形態に、多様な変更又は改良を加えることが可能であることが当業者に明らかである。その様な変更又は改良を加えた形態も本発明の技術的範囲に含まれ得ることが、請求の範囲の記載から明らかである。
 請求の範囲、明細書、及び図面中において示した装置、システム、プログラム、及び方法における動作、手順、ステップ、及び段階などの各処理の実行順序は、特段「より前に」、「先立って」などと明示しておらず、また、前の処理の出力を後の処理で用いるのでない限り、任意の順序で実現しうることに留意すべきである。請求の範囲、明細書、及び図面中の動作フローに関して、便宜上「まず、」、「次に、」などを用いて説明したとしても、この順で実施することが必須であることを意味するものではない。
20 ネットワーク、30 通信端末、32 カメラ、40 養殖場、100 データ処理装置、102 カメラ、112 変換情報記憶部、114 モデル記憶部、116 動画データ取得部、118 データ生成部、120 分類結果取得部、122 モデル生成部、130 出力部、132 アラート出力部、134 表示出力部、136 音声出力部、142 変換データ取得部、144 動画データ生成部、200 動画データ、210、220、230、240、250、260 動画データ、211、221、231、241、251、261 フレーム、300 静止画データ、302、311、321、331、341、351、361 部分、400 テキストデータ、402、411、421、431、441、451、461 部分、500 ゲノムデータ、1200 コンピュータ、1210 ホストコントローラ、1212 CPU、1214 RAM、1216 グラフィックコントローラ、1218 ディスプレイデバイス、1220 入出力コントローラ、1222 通信インタフェース、1224 記憶装置、1230 ROM、1240 入出力チップ

Claims (12)

  1.  オブジェクトを撮像した動画データを取得する動画データ取得部と、
     前記動画データにおける前記オブジェクトの時系列の動きの情報を静止画の要素に変換して、変換した前記静止画の前記要素を含む静止画データを生成するデータ生成部と
     を備える、データ処理装置。
  2.  前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の動きの情報を、画素値、特徴ベクトル、及びエッジの少なくともいずれかによって表す前記静止画データを生成する、請求項1に記載のデータ処理装置。
  3.  前記データ生成部は、前記動画データにおける複数の前記オブジェクトの時系列の位置の変化の情報を前記静止画の前記要素に変換して、変換した前記静止画の前記要素を含む前記静止画データを生成する、請求項1に記載のデータ処理装置。
  4.  前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の形状変化の情報を前記静止画の前記要素に変換して、変換した前記静止画の前記要素を含む前記静止画データを生成する、請求項1に記載のデータ処理装置。
  5.  前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の動きの情報をテキストの要素に変換して、変換した前記テキストの前記要素を含むテキストデータを生成する、請求項1に記載のデータ処理装置。
  6.  前記データ生成部は、前記動画データにおける前記オブジェクトの時系列の動きの情報をゲノムの要素に変換して、変換した前記ゲノムの前記要素を含むゲノムデータを生成する、請求項1に記載のデータ処理装置。
  7.  前記データ生成部によって生成された前記静止画データを、静止画データを分類するように学習された静止画分類ネットワークに入力して、前記静止画分類ネットワークから出力された分類結果を取得する分類結果取得部と、
     前記分類結果取得部が取得した前記分類結果が予め定められた条件を満たす場合に、アラートを出力するアラート出力部と
     を更に備える、請求項1から6のいずれか一項に記載のデータ処理装置。
  8.  動画データにおけるオブジェクトの時系列の動きの情報から変換された静止画の要素を含む静止画データを取得する変換データ取得部と、
     前記静止画データから前記動画データを生成する動画データ生成部と
     を更に備える、請求項1から6のいずれか一項に記載のデータ処理装置。
  9.  動画データにおけるオブジェクトの時系列の動きの情報から変換された静止画の要素を含む静止画データを取得する静止画データ取得部と、
     前記静止画データから前記動画データを生成する動画データ生成部と
     を備えるデータ処理装置。
  10.  オブジェクトを撮像した動画データを取得する動画データ取得部と、
     前記動画データにおける前記オブジェクトの時系列の動きの情報をテキストの要素に変換して、変換した前記テキストの前記要素を含むテキストデータを生成するデータ生成部と
     を備える、データ処理装置。
  11.  オブジェクトを撮像した動画データを取得する動画データ取得部と、
     前記動画データにおける前記オブジェクトの時系列の動きの情報をゲノムの要素に変換して、変換した前記ゲノムの前記要素を含むゲノムデータを生成するデータ生成部と
     を備える、データ処理装置。
  12.  コンピュータを、請求項1から6、9、10、11のいずれか一項に記載のデータ処理装置として機能させるためのプログラム。
PCT/JP2024/020404 2023-06-13 2024-06-04 データ処理装置及びプログラム Ceased WO2024257656A1 (ja)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2023096904A JP7519506B1 (ja) 2023-06-13 2023-06-13 データ処理装置及びプログラム
JP2023-096904 2023-06-13

Publications (1)

Publication Number Publication Date
WO2024257656A1 true WO2024257656A1 (ja) 2024-12-19

Family

ID=91895796

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2024/020404 Ceased WO2024257656A1 (ja) 2023-06-13 2024-06-04 データ処理装置及びプログラム

Country Status (2)

Country Link
JP (1) JP7519506B1 (ja)
WO (1) WO2024257656A1 (ja)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06309462A (ja) * 1993-02-26 1994-11-04 Fujitsu Ltd 動画像処理装置
JP2015184798A (ja) * 2014-03-20 2015-10-22 ソニー株式会社 情報処理装置、情報処理方法及びコンピュータプログラム
CN115497156A (zh) * 2021-06-01 2022-12-20 阿里巴巴新加坡控股有限公司 动作识别方法和装置、电子设备及计算机可读存储介质

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2006276948A (ja) 2005-03-28 2006-10-12 Seiko Epson Corp 画像処理装置、画像処理方法、画像処理プログラムおよび画像処理プログラムを記憶した記録媒体
JP6340675B1 (ja) 2017-03-01 2018-06-13 株式会社Jストリーム オブジェクト抽出装置、オブジェクト認識システム及びメタデータ作成システム

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06309462A (ja) * 1993-02-26 1994-11-04 Fujitsu Ltd 動画像処理装置
JP2015184798A (ja) * 2014-03-20 2015-10-22 ソニー株式会社 情報処理装置、情報処理方法及びコンピュータプログラム
CN115497156A (zh) * 2021-06-01 2022-12-20 阿里巴巴新加坡控股有限公司 动作识别方法和装置、电子设备及计算机可读存储介质

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
KEI TERAHARA , MASAYUKI KASHIMA , KIMINORI SATO , MUTSUMI WATANABE : "Research on automatic lost article detection by analyzing human's walking locus", IEICE TECHNICAL REPORT, IEICE, JP, vol. 109, no. 471- PRMU2009-257, HIP2009-142, 8 March 2010 (2010-03-08), JP, pages 139 - 144, XP009559352, ISSN: 0913-5685 *

Also Published As

Publication number Publication date
JP7519506B1 (ja) 2024-07-19
JP2024178625A (ja) 2024-12-25

Similar Documents

Publication Publication Date Title
Brownlee Deep learning for computer vision: image classification, object detection, and face recognition in python
CN110610510B (zh) 目标跟踪方法、装置、电子设备及存储介质
CN111783620B (zh) 表情识别方法、装置、设备及存储介质
US10776662B2 (en) Weakly-supervised spatial context networks to recognize features within an image
CN111370020A (zh) 一种将语音转换成唇形的方法、系统、装置和存储介质
WO2018176186A1 (en) Semantic image segmentation using gated dense pyramid blocks
CN108154236A (zh) 用于评估群组层面的认知状态的技术
US20220207665A1 (en) Intelligence-based editing and curating of images
Appiah et al. A single-chip FPGA implementation of real-time adaptive background model
US20250182438A1 (en) Detection of moment of perception
KR102788804B1 (ko) 전자 장치 및 그 제어 방법
Quinn et al. British sign language recognition in the wild based on multi-class SVM
Alamri et al. Intelligent real-life key-pixel image detection system for early Arabic sign language learners
Tvoroshenko et al. Object identification method based on image keypoint descriptors
US12020510B2 (en) Person authentication apparatus, control method, and non-transitory storage medium
JP7519506B1 (ja) データ処理装置及びプログラム
CN114842295B (zh) 绝缘子故障检测模型的获得方法、装置及电子设备
CN115690615A (zh) 一种面向视频流的深度学习目标识别方法及系统
JP7518124B2 (ja) 画像処理装置、プログラム、及び画像処理方法
WO2024231989A1 (ja) 情報解析システム、情報解析方法及びプログラム
CN116212393B (zh) 一种关卡元素布局生成方法、装置及关卡编辑器
Anita et al. Early Fire Detection using SVM based Machine Learning Algorithm
Paduraru A state-aware, hierarchical deep learning framework for automated visual glitch detection in games
RU2774624C1 (ru) Способ и система определения синтетических изменений лиц в видео
JP7545433B2 (ja) 画像処理装置、プログラム、画像処理システム、及び画像処理方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24823275

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE