WO2026022985A1 - 撮像装置、撮像支援方法、およびプログラム - Google Patents

撮像装置、撮像支援方法、およびプログラム

Info

Publication number
WO2026022985A1
WO2026022985A1 PCT/JP2024/026516 JP2024026516W WO2026022985A1 WO 2026022985 A1 WO2026022985 A1 WO 2026022985A1 JP 2024026516 W JP2024026516 W JP 2024026516W WO 2026022985 A1 WO2026022985 A1 WO 2026022985A1
Authority
WO
WIPO (PCT)
Prior art keywords
information
imaging
imaging device
unit
image
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/JP2024/026516
Other languages
English (en)
French (fr)
Inventor
博仁 栗山
秀幸 桑島
拓也 清水
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Maxell Ltd
Original Assignee
Maxell Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Maxell Ltd filed Critical Maxell Ltd
Priority to PCT/JP2024/026516 priority Critical patent/WO2026022985A1/ja
Publication of WO2026022985A1 publication Critical patent/WO2026022985A1/ja
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N23/00Cameras or camera modules comprising electronic image sensors; Control thereof
    • H04N23/60Control of cameras or camera modules

Definitions

  • the present invention relates to an imaging device, an imaging support method, and a program.
  • Patent Document 1 does not sufficiently consider configurations for providing users with response output technology using artificial intelligence in a more suitable manner.
  • the object of the present invention is to provide a more suitable response output technology.
  • the present application includes multiple means for solving the above problem, but one example would be an imaging device comprising an imaging unit, a display unit, and a control unit, where the control unit is configured to perform the following processes: sending to a generating artificial intelligence instruction information requesting imaging support information for supporting a user in imaging using the imaging unit; receiving from the generating artificial intelligence response information including the imaging support information; and displaying on the display unit the imaging support information included in the received response information or information based on the imaging support information.
  • an imaging support method may be configured such that a control unit in an imaging device executes the following processes: sending instruction information to a generating artificial intelligence requesting imaging support information for supporting a user in imaging using the imaging unit; receiving response information including the imaging support information from the generating artificial intelligence; and displaying the imaging support information included in the received response information or information based on the imaging support information on a display unit.
  • a program may be configured to cause a processor constituting a control unit in an imaging device to execute the following processes: sending instruction information to a generating artificial intelligence requesting imaging support information to assist a user in capturing an image using the imaging unit; receiving response information including the imaging support information from the generating artificial intelligence; and displaying the imaging support information included in the received response information or information based on the imaging support information on a display unit.
  • the present invention provides a more suitable response output technology.
  • Other issues, configurations, and advantages will be made clear in the description of the embodiments below.
  • FIG. 1 is a diagram illustrating an example of an artificial intelligence response output device and system according to an embodiment of the present invention.
  • 1 is a diagram illustrating an example of an artificial intelligence response output device according to an embodiment of the present invention.
  • 1 is a diagram showing an example of the operation of an AI response output device and system according to an embodiment of the present invention;
  • FIG. 1 is a diagram illustrating an example of a system including an imaging device and a generation artificial intelligence server according to an embodiment of the present invention.
  • 1 is a diagram illustrating an example of a configuration of an imaging device according to an embodiment of the present invention.
  • FIG. 2 is a diagram showing an example of a display screen of an imaging device according to an embodiment of the present invention.
  • 10A and 10B are diagrams illustrating an example of a change in composition before and after receiving advice from a generation artificial intelligence according to an embodiment of the present invention.
  • 10A and 10B are diagrams showing an example of how an unnecessary object is extracted from a through image obtained by an imaging device according to an embodiment of the present invention.
  • 10A and 10B are diagrams showing an example in which an original image and a temporarily processed image are displayed in an imaging device according to an embodiment of the present invention.
  • the AI response output device may be called a display device. If the AI response output device has an audio output function, it may be called an audio output device.
  • the AI response output device may simply be called an information processing device.
  • a system including an AI response output device and a large-scale language model server that holds a large-scale language model may be called an AI response output system.
  • the AI response output device provides a user with a response service based on a large-scale language model, which is an AI, and is of assistance to the user, the AI response output device or the display output of the AI response output device can serve as an artificial intelligence (AI) assistant for the user.
  • AI artificial intelligence
  • the AI response output device may be called an AI assistant device or an AI assistant display device.
  • a system including an AI response output device and a large-scale language model server that holds a large-scale language model may be called an AI assistant system or an AI assistant display system.
  • the AI response output device serves as an interface between the user and the AI, and may therefore be called an AI interface device.
  • a system including an AI response output device and a large-scale language model server that holds a large-scale language model may be called an AI interface system.
  • Example 1 As a first embodiment of the present invention, an AI response output device and system for outputting a response from a large-scale language model AI will be described.
  • FIG. 1A an example of an AI response output device 10010 of the present invention will be described. Furthermore, in the case where the AI response output device 10010 is linked to a large-scale language model server 19001 via communication or the like, an example of a system in which the AI response output device 10010 includes the large-scale language model server 19001 and/or a multimodal large-scale language model server 20001 will be described.
  • the AI response output device 10010 has a display unit 10011.
  • the display unit 10011 may be a flat display, a screen that projects an image from the back, or a floating image that forms an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight.
  • the display unit 10011 may also be a plasma display.
  • the display unit 10011 may also be an organic EL display in which the pixels emit light themselves.
  • the display unit 10011 may also be provided with a touch operation input sensor and configured as a touch panel.
  • the audio output unit 1140 provided in the AI response output device 10010 is composed of a speaker.
  • the AI response output device 10010 also has a microphone 1139, which can pick up the user's voice. By receiving audio input from the microphone 1139 or by receiving user operation input via the operation input unit described below, the AI response output device 10010 can obtain user input that serves as the basis for instruction sentences (prompts) for the large-scale language model, which is the AI.
  • the AI response output device 10010 may be equipped with a local large-scale language model.
  • the response of the large-scale language model may be output as a display output from the display unit 10011 and/or as an audio output from the audio output unit 1140.
  • the AI response output device 10010 may not have a local large-scale language model, but may communicate with an external large-scale language model server 19001, and output the response received from the large-scale language model server 19001 as a display output on the display unit 10011 and/or as an audio output on the audio output unit 1140.
  • the artificial intelligence response output device 10010 may also have a local large-scale language model and may be further configured to communicate with an external large-scale language model server 19001 having a large-scale language model or an external large-scale language model server 20001 having a multimodal large-scale language model.
  • the response of the local large-scale language model and a response received from the large-scale language model server 19001 or the multimodal large-scale language model server 20001 may be switched, and either one may be output as a display output from the display unit 10011 and/or an audio output from the audio output unit 1140.
  • a response generated based on both the response of the local large-scale language model and the response received from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 may be output as a display output from the display unit 10011 and/or an audio output from the audio output unit 1140.
  • the configuration when the AI response output device 10010 communicates and cooperates with an external large-scale language model server 19001 or large-scale language model server 20001 is as follows.
  • the AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000 via the communication unit 1132.
  • the communication between the communication unit 1132 and the communication device 19011 is shown as wireless, but wired communication is also acceptable.
  • the communication path from the communication unit 1132 to the communication device 19011 may include wired and wireless portions, or may go via a router or repeater.
  • the communication path from the communication unit 1132 to the Internet 19000 may include wired and wireless portions, or may go via a router or repeater.
  • the AI response output device 10010 can communicate with the large-scale language model server 19001 via the communication device 19011 and the Internet 19000. Furthermore, the AI response output device 10010 can communicate with the large-scale language model server 19001 or the large-scale language model server 20001, and a second server 19002 different from these servers, via a communication device 19011 and the Internet 19000. A configuration including the AI response output device 10010 and the large-scale language model server 19001 or the large-scale language model server 20001 may be considered as a single system.
  • large-scale language model can be considered to refer to the local large-scale language model provided by the AI response output device 10010, the large-scale language model provided by the large-scale language model server 19001, and the multimodal large-scale language model provided by the large-scale language model server 20001.
  • FIG. 1A shows an example in which the display unit 10011 displays elements in two display areas: a prompt display area 10051 in which the user inputs prompts to the large-scale language model, which is an artificial intelligence; and an artificial intelligence response display area 10061 in which responses from the large-scale language model are displayed.
  • the prompt display area 10051 displays an icon 10052 indicating the user, text 10053 such as natural language or software code as a component of the prompt, an image 10054 as a component of the prompt, and a video 10055 as a component of the prompt.
  • text 10053 such as natural language or software code
  • the artificial intelligence response display area 10061 displays an icon 10062 indicating the artificial intelligence or artificial intelligence assistant, text 10063 such as natural language or software code as a component of the response from the artificial intelligence, an image 10064 as a component of the response from the artificial intelligence, and a video 10065 as a component of the response from the artificial intelligence.
  • the display example of the display unit 10011 of the AI response output device 10010 shown in Figure 1A is merely an example. Depending on the implementation example in which the AI response output device 10010 is used, a display different from the example shown in Figure 1A may be displayed.
  • LLMs Large-scale language models
  • GPT-1 Global System for Mobile Communications
  • GPT-2 Global System for Mobile Communications
  • GPT-3 Generalized Synchronization Protocol
  • InstructGPT InstructGPT
  • ChatGPT ChatGPT
  • GPT-1 Global System for Mobile Communications
  • GPT-2 GPT-3
  • InstructGPT InstructGPT
  • ChatGPT ChatGPT
  • These technologies can be used in this embodiment as well.
  • these large-scale language models are artificial intelligence models generated through large-scale pre-training of the natural language contained in the numerous documents and texts that exist in the human world. The number of parameters in artificial intelligence models exceeds one billion.
  • An example of a base model is a model called a Transformer. Reference 1, for example, is published as an example of training for these models.
  • large-scale language models are capable of natural language translation, natural language text proofreading, and natural language text summarization.
  • Advanced models are capable of natural language question-answering (also known as dialogue or conversation), natural language suggestion generation, and programming code generation.
  • these AI models have a very large number of parameters, training requires vast amounts of data and computational resources. Therefore, training this level of AI for specific applications is extremely resource-inefficient. Therefore, large-scale pre-training is performed to generate models as foundation models that can be applied to various applications.
  • the large-scale language model server 19001 shown in Figure 1A may be equipped with such a large-scale language model and configured to be accessible on various terminals via an API (Application Programming Interface).
  • API Application Programming Interface
  • the AI response output device 10010 shown in Figure 1A may be equipped with a local large-scale language model and configured to be used by the AI response output device 10010 itself.
  • the learning of any large-scale language model itself can be generated by separate large-scale pre-learning, and the generated large-scale language model can be replicated and provided on the large-scale language model server 19001 or the AI response output device 10010.
  • replicating the large-scale language model that is the base model generated by large-scale pre-learning and using it on individual servers and terminals allows the resources used for learning to be shared, resulting in good resource efficiency.
  • a large-scale language model may be configured to perform additional learning such as transfer learning on individual servers or devices depending on the application and purpose.
  • large-scale language models can be pre-trained on natural languages and perform input/output processing targeting natural languages.
  • multimodal large-scale language model AI capable of processing not only natural language text information but also types of information other than natural language text information can also be applied to embodiments of the present invention.
  • Figure 1A shows a large-scale language model server 20001 having a multimodal large-scale language model.
  • Specific examples of multimodal large-scale language model AI include GPT-4 (see Reference 2) and Gato (see Reference 3). These technologies can also be used in this embodiment.
  • these multimodal large-scale language models are AI models generated through large-scale pre-training on the natural language contained in the numerous documents and texts present in the human world, as well as types of information other than natural language text information (e.g., images, videos, audio, etc.). In addition to this, there are also models that have undergone reinforcement learning based on human feedback.
  • types of information other than natural language text information such as images, videos, and audio, may be referred to as non-natural language information sources.
  • an artificial intelligence response output device 10010 that accepts input from a user to an artificial intelligence such as a large-scale language model, and outputs a response from the artificial intelligence such as a large-scale language model in response to the user input.
  • the AI response output device 10010 includes a display unit 10011, a control unit 1110, a memory 1109, a non-volatile memory 1108, an external power input interface 1111, an operation input unit 1107, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, a posture sensor 1113, a communication unit 1132, an audio output unit 1140, a microphone 1139, a video signal input unit 1131, an audio signal input unit 1133, an imaging unit 1180, and the like.
  • the AI response output device 10010 may have a large screen, such as a monitor or television.
  • the display unit 10011 may be a flat display, a screen that projects an image from the back, or a device that displays a floating image by forming an optical image in the air. If the display unit 10011 is a flat display, it may be a liquid crystal display having a liquid crystal panel and a backlight. The display unit 10011 may also be a plasma display. The display unit 10011 may also be an organic EL display in which the pixels are self-luminous. If the display unit 10011 is a panel, it may be called a display panel. The display unit 10011 may be provided with a touch operation input sensor and configured to accept touch operation input by the finger of the user 230. In this case, the display unit 10011 may be configured as a touch panel. By the user's operation input via the touch panel, the artificial intelligence response output device 10010 can obtain the user input that serves as the basis for instructions (prompts) for the large-scale language model, which is the artificial intelligence.
  • the communication unit 1132 may be configured with a Wi-Fi communication interface, a Bluetooth (registered trademark) communication interface, or a mobile communication interface such as 4G or 5G. Using these communication methods, the communication unit 1132 of the AI response output device 10010 can communicate with a communication device 19011 connected to the Internet 19000.
  • the communication path between the communication unit 1132 and the communication device 19011 may include wired and wireless portions, or may go via a router or repeater. In the wired case, the communication unit 1132 may have an Ethernet connection interface as hardware and communicate using a LAN communication method. This allows the AI response output device 10010 to communicate with various servers connected to the Internet 19000.
  • the AI response output device 10010 is equipped with a control unit 1110 such as a CPU and memory 1109, and the control unit 1110 controls the display unit 10011, communication unit 1132, etc.
  • a control unit 1110 such as a CPU and memory 1109
  • the control unit 1110 controls the display unit 10011, communication unit 1132, etc.
  • the power supply 1106 converts AC current input from the outside via the external power supply input interface 1111 into DC current and supplies the necessary DC current to each part of the AI response output device 10010.
  • the secondary battery 1112 stores the power supplied from the power supply 1106. Furthermore, the secondary battery 1112 supplies power to each part requiring power via the external power supply input interface 1111 when power is not being supplied from the outside.
  • the operation input unit 1107 is, for example, an operation button, a signal receiving unit or an infrared light receiving unit of a remote controller, and inputs signals for operations other than the user's touch operation on the touch operation input sensor of the display unit 10011. Separate from a user touching the touch operation input sensor of the display unit 10011, the operation input unit 1107 may also be used by, for example, an administrator to operate the AI response output device 10010. By the user's operation input via the operation input unit 1107, the AI response output device 10010 can obtain user input that serves as the basis for instructions (prompts) for the large-scale language model, which is the AI. A modified configuration is also possible in which the touch operation input sensor of the display unit 10011 is also included as part of the operation input unit 1107.
  • Video signal input unit 1131 connects to an external video output device and inputs video data.
  • the video signal input unit 1131 can be configured with a variety of digital video input interfaces. For example, it can be configured with a video input interface conforming to the HDMI (registered trademark) (High-Definition Multimedia Interface) standard, a video input interface conforming to the DVI (Digital Visual Interface) standard, or a video input interface conforming to the DisplayPort standard. Alternatively, an analog video input interface such as analog RGB or composite video may be provided.
  • the video signal input unit 1131 can also be any of various USB interfaces.
  • Audio signal input unit 1133 connects to an external audio output device and inputs audio data.
  • Audio signal input unit 1133 may be configured as an HDMI-standard audio input interface, an optical digital terminal interface, or a coaxial digital terminal interface. Audio signal input unit 1133 may also be various USB interfaces. In the case of an HDMI-standard interface, video signal input unit 1131 and audio signal input unit 1133 may be configured as an interface in which the terminal and cable are integrated.
  • the audio output unit 1140 is capable of outputting audio based on audio data input to the audio signal input unit 1133.
  • the audio output unit 1140 is also capable of outputting audio based on audio data stored in the storage unit 1170.
  • the audio output unit 1140 may be configured as a speaker.
  • the audio output unit 1140 may also output built-in operation sounds or error warning sounds.
  • the audio output unit 1140 may be configured to output an audio signal as a digital signal to an external device, such as the Audio Return Channel function defined in the HDMI standard.
  • the audio output unit 1140 may be configured to output an audio signal as an analog signal to an external device such as headphones.
  • Microphone 1039 is a microphone that picks up sounds around the AI response output device 10010, converts them into signals, and generates audio signals.
  • the microphone may also be configured to record a person's voice, such as a user's voice, and the control unit 1110, described below, performs voice recognition processing on the generated audio signal to obtain text information from the audio signal.
  • AI response output device 10010 can obtain user input that will serve as the basis for instructions (prompts) for the large-scale language model, which is the AI.
  • the imaging unit 1180 is a camera with an image sensor.
  • the camera may be provided on the front side of the display unit 10011 of the AI response output device 10010, or on the back side of the display unit 10011. Both a front camera and a back camera may be provided. In this embodiment, the imaging unit 1180 will be described as having both a front camera and a back camera.
  • the storage unit 1170 is a storage device that records various types of information, such as video data, image data, and audio data.
  • the storage unit 1170 may be configured as a magnetic recording medium recording device, such as a hard disk drive (HDD), or a semiconductor element memory, such as a solid state drive (SSD).
  • HDD hard disk drive
  • SSD solid state drive
  • various types of information, such as video data, image data, and audio data may be recorded in the storage unit 1170 before product shipment.
  • the storage unit 1170 may also record various types of information, such as video data, image data, and audio data, acquired from external devices, external servers, etc. via the communication unit 1132.
  • the video data, image data, etc. recorded in the storage unit 1170 is output to the display unit 10011.
  • the video data, image data, etc. recorded in the storage unit 1170 may also be output to external devices, external servers, etc. via the communication unit 1132.
  • the video control unit 1160 performs various controls related to the video signal input to the display unit 10011.
  • the video control unit 1160 may be referred to as a video processing circuit, and may be configured with hardware such as an ASIC, FPGA, or video processor.
  • the video control unit 1160 may also be referred to as a video processing unit or image processing unit.
  • the video control unit 1160 controls video switching, such as which video signal to input to the display unit 10011, between the video signal to be stored in the memory 1109 and the video signal (video data) input to the video signal input unit 1131.
  • the video control unit 1160 may also control image processing of the video signal input from the video signal input unit 1131 and the video signal to be stored in the memory 1109.
  • image processing examples include scaling, which enlarges, reduces, or transforms an image; brightness adjustment, which changes the brightness; contrast adjustment, which changes the contrast curve of an image; and Retinex processing, which decomposes an image into its light components and changes the weighting of each component.
  • the attitude sensor 1113 is a sensor configured as a gravity sensor or an acceleration sensor, or a combination of these, and is capable of detecting the attitude of the AI response output device 10010. Based on the attitude detection results of the attitude sensor 1113, the control unit 1110 may control the operation of each connected unit.
  • the non-volatile memory 1108 stores various data used by the AI response output device 10010.
  • Data stored in the non-volatile memory 1108 includes, for example, data for various operations to be displayed on the display unit 10011 of the AI response output device 10010, display icons, object data for user operation, and layout information.
  • the memory 1109 stores video data to be displayed on the display unit 10011, data for controlling the device, and the like.
  • the control unit 1110 may read various software from the storage unit 1170 and expand and store it in the memory 1109.
  • the local LLM processing unit 10028 has a memory capable of holding a large-scale language model (LLM), and can execute inference of the large-scale language model based on the control of the control unit 1110.
  • the hardware may be configured with a so-called GPU (Graphics Processing Unit) or the like.
  • the local LLM processing unit 10028 may not only execute inference, but also learn. Note that the local LLM processing unit 10028 is not necessarily required in cases where it is not necessary to execute inference of a large-scale language model in the local environment of the artificial intelligence response output device 10010.
  • the control unit 1110 controls the operation of each connected unit.
  • the control unit 1110 may also work in conjunction with a program stored in the memory 1109 to perform arithmetic processing based on information acquired from each unit within the AI response output device 10010.
  • Control states by the control unit 1110 include, for example, a state in which a response from the large-scale language model of the local LLM processing unit 10028, or a response from the large-scale language model of the large-scale language model server 19001 or the multimodal large-scale language model of the multimodal large-scale language model server 20001 acquired via the communication unit 1132 is output via the display unit 10011 or the audio output unit 1140, such as a speaker.
  • the control unit 1110 can generate an instruction sentence based on the input, send it to the local large-scale language model of the local LLM processing unit 10028 provided in the AI response output device 10010, the large-scale language model provided in the large-scale language model server 19001, or the multimodal large-scale language model provided in the large-scale language model server 20001, and obtain a response from these large-scale language models.
  • the storage unit 1170 may also store a fixed response phrase database (which may be referred to as a fixed response phrase DB) for outputting fixed phrases in response to instruction sentences from the AI response output device 10010.
  • the control unit 1110 may perform control to generate responses to be output using data stored in the fixed response phrase database.
  • Figure 1C shows an example of a fixed response phrase database. In the example of Figure 1C, fixed response phrases output by the AI response output device 10010 are stored for each condition assigned a condition number.
  • a response may be output using the fixed response phrase “Good morning” or "Today is _____ day of _____ month, isn't it?"
  • the _____ part of "_____ day of ____ month” may be generated using information stored in the memory 1109 or the like of the AI response output device 10010.
  • the control unit 1110 can perform control so that one of the canned response phrases is selected at random using a random number or the like and output as a response. In this way, it is possible to eliminate and improve the situation where responses under the same conditions become monotonous.
  • condition number 1 is the same for the examples of condition numbers 2, 3, and 4.
  • the control unit 1110 can perform control so that the canned response phrase of each example shown in Figure 1C is output for the condition content of each example shown in Figure 1C.
  • Condition number 5 is an example of control in which, when the control unit 1110 is unable to understand the meaning of user input acquired via the touch panel, microphone 1139, or operation input unit 1107 as natural language, or when the user input contains a clear grammatical error, the control unit 1110 outputs a response using a standard response phrase such as "I didn't quite catch that” or "I might not know about that.” By responding in this way, the user can be prompted to input again, and corrected user input can be waited for.
  • a standard response phrase such as "I didn't quite catch that" or "I might not know about that.”
  • Condition number 6 is an example of when the control unit 1110 detects an error (abnormal state) in any of the components constituting the AI response output device 10010 shown in Figure 1B, and a user input is made via the touch panel, microphone 1139, or operation input unit 1107. In this case, the control unit 1110 performs control to output a response using the standard response phrase "It seems to be working poorly.” By responding in this manner, it is possible to explain to the user that the AI response output device 10010 is malfunctioning, and to prompt the user to take action to deal with the error, etc.
  • the AI response output device 10010 may output a response using the canned response database (canned response DB) described using Figure 1C. Alternatively, it may output a response that combines the responses from these large-scale language models with a response using the canned response database (canned response DB).
  • the fixed response phrase database (fixed response phrase DB) of Figure 1C described above is stored in the storage unit 1170, and is used by the control unit 1110 of the artificial intelligence response output device 10010.
  • the fixed response phrase database (fixed response phrase DB) shown in Figure 1C may be provided on the large-scale language model server 19001 side or the large-scale language model server 20001 side.
  • the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may generate a response using the fixed response phrase database (fixed response phrase DB).
  • the control unit of the large-scale language model server 19001 or the control unit of the large-scale language model server 20001 may send a response generated using the fixed response phrase database (fixed response phrase DB) to the artificial intelligence response output device 10010, instead of a response generated using the large-scale language model stored on the respective server.
  • a response generated using the fixed response phrase database fixed response phrase DB
  • the AI response output device 10010 does not have a fixed response phrase database (fixed response phrase DB)
  • the AI response output device 10010 has been described as having a display panel with a display screen that uses fixed pixels.
  • This concept may also include a projection-type image display device (projector) that has a projection optical system provided behind the display panel with a display screen that uses fixed pixels, and projects an optical image of the image on the display panel of the display screen onto a screen or wall.
  • projector projection-type image display device
  • the AI response output device 10010 is equipped with a display unit 10011.
  • the AI response output device 10010 does not necessarily have to be equipped with a display unit 10011.
  • it may be configured to accept input from a user to the AI via the audio signal input unit 1133 or microphone 1139, and output a response from the AI, such as a large-scale language model, in response to the user input via the audio output unit 1140.
  • the AI response output device and AI response output system are capable of accepting input from a user to an AI such as a large-scale language model, and outputting a response to the user input that is generated by AI inference, such as a large-scale language model held by a server device on a network or a local large-scale language model held by the AI response output device itself.
  • AI such as a large-scale language model
  • AI inference such as a large-scale language model held by a server device on a network or a local large-scale language model held by the AI response output device itself.
  • the technology according to this embodiment makes it possible to provide more suitable AI response output technology.
  • AI response output technology is expected to be introduced into higher quality, more reliable infrastructure.
  • the introduction of this technology into infrastructure will contribute to supporting economic development and human welfare, with a focus on affordable and fair access for all. This will contribute to "Build resilient infrastructure, promote inclusive and sustainable industrialization, inclusive and sustainable technological development," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.
  • SDGs Sustainable Development Goals
  • the technology according to this embodiment makes it possible to provide more suitable AI response output technology. It is expected that such AI response output technology will be introduced into public transportation facilities to improve access to transportation systems for vulnerable people. The introduction of this technology into public transportation will contribute to improving traffic safety through the expansion of public transportation, and realizing access to a safe, affordable, and easily usable sustainable transportation system for all people. This will contribute to "Sustainable cities and communities" - goal 11 of the Sustainable Development Goals (SDGs) advocated by the United Nations.
  • SDGs Sustainable Development Goals
  • Example 2 A second embodiment of the present invention is an example of an imaging device and an imaging support method to which an artificial intelligence response output technology is applied.
  • the imaging device and the imaging support method according to the second embodiment support a user in imaging by an AI assistant function using the artificial intelligence response output technology.
  • FIG. 2A is a diagram showing an example of a system including an imaging device and a generative artificial intelligence (hereinafter, artificial intelligence will also be referred to as AI) server according to Example 2.
  • the system according to Example 2 includes an imaging device 1200, a large-scale language model (hereinafter, large-scale language model will also be referred to as LLM) server 19001, a multimodal LLM server 20001, and a second server 19002 that is different from these servers.
  • the imaging device 1200, the LLM server 19001, the multimodal LLM server 20001, and the second server 19002 are connected via the Internet 19000, which is a communications network.
  • the LLM server 19001 and the multimodal LLM server 20001 are each an example of a generative AI server on which generative AI is built.
  • the imaging device 1200 has an imaging function unit 1201 that has an imaging unit 1180, and an AI response output device 10010.
  • the AI response output device 10010 has substantially the same configuration and functions as the AI response output device 10010 in Example 1.
  • FIG. 2B is a diagram showing an example of the configuration of an imaging device according to Example 2.
  • the imaging function unit 1201 has an imaging unit 1180, a captured image input unit 1181, an external image input unit 1182, an image input adjustment unit 1183, an image and audio memory unit 1184, an image encoding unit 1185, an audio encoding unit 1186, a data generation unit 1187, a data recording unit 1188, a recording medium 1189, a current position acquisition unit 1195, a video signal input unit 1131, and an audio signal input unit 1133.
  • the imaging unit 1180 is used to obtain image information corresponding to a captured image of a subject.
  • the imaging unit 1180 includes, for example, an optical system such as a lens, an aperture, a shutter, an image sensor, etc.
  • the captured image input unit 1181 accepts input of image information of the captured image obtained by the imaging unit 1180.
  • the external image input unit 1182 accepts input of image information of an external image via the video signal input unit 1131.
  • the image input adjustment unit 1183 adjusts whether image information of a captured image or an external image is to be input.
  • the image audio memory unit 1184 temporarily stores image information input from the image input adjustment unit 1183 and audio information input from the audio signal input unit 1133.
  • the image encoding unit 1185 converts and compresses the image information stored in the image and audio memory unit 1184 into data in a predetermined format.
  • the image encoding unit 1186 converts and compresses the audio information stored in the image and audio memory unit 1184 into data in a predetermined format.
  • the data generation unit 1187 generates image data or video data based on data from the image encoding unit 1185, or data from the image encoding unit 1185 and data from the image encoding unit 1186.
  • the data recording unit 1188 records the image data or video data generated by the data generation unit 1187 on a recording medium 1189.
  • the recording medium 1189 is, for example, an HDD, SSD, USB ROM, DVD, or CD.
  • the current position acquisition unit 1195 acquires location information indicating the current position of the imaging device 1200, such as a GPS system.
  • the AI response output device 10010 includes a control unit 1110, a display unit 10011, an external power input interface 1111, a power supply 1106, a secondary battery 1112, a storage unit 1170, a video control unit 1160, an operation input unit 1107, an attitude sensor 1113, a memory 1109, a local LLM processing unit 10028, a non-volatile memory 1108, a communication unit 1132, an audio output unit 1140, and a microphone 1139.
  • the configuration and functions of the AI response output device 10010 are substantially the same as those of the AI response output devices according to Examples 1 to 9.
  • the control unit 1110 is connected to communicate with each unit, including the captured image input unit 1181 and image input adjustment unit 1183 that constitute the imaging function unit 1201.
  • the control unit 1110 can perform various processes on the image information of the captured image obtained by the imaging unit 1180.
  • the storage unit 1170 or non-volatile memory 1108 stores a program for implementing the AI assistant function.
  • the control unit 1110 reads this program, expands it into memory 1109, and executes it to implement the AI assistant function and control its on/off state.
  • the control unit 1110 executes processing to send to the generation AI instruction information requesting imaging support information to assist the user in imaging using the imaging unit 1180.
  • the control unit 1110 also executes processing to receive response information including the imaging support information from the generation AI.
  • the control unit 1110 further executes processing to display on the display unit 10011 the imaging support information included in the received response information or information based on the imaging support information. Through this processing by the control unit 1110, the user is assisted in imaging using the imaging unit 1180.
  • the storage unit 1170, non-volatile memory 1108, memory 1109, and recording medium 1189 are each an example of a “storage unit” in this application.
  • the current location acquisition unit 1195 is an example of a "location information acquisition unit” in this application.
  • FIGS. 2C to 2F are diagrams showing an example of the processing flow in the imaging device and generation AI server according to the second embodiment.
  • the processing flows shown in FIGS. 2C to 2F chronologically describe the main processing executed by the control unit 1110 in the imaging device 1200 and the main processing executed by the generation AI server including the multimodal LLM server 20001.
  • the processing flows shown in FIGS. 2C to 2F are processing flows assuming that the processing is executed in parallel with processing in response to user operations in the imaging device 1200. It is assumed that in the imaging device 1200, imaging is performed in the background by the imaging unit 1180 at approximately regular intervals, and multiple time-series through images obtained over the most recent fixed period are stored in the recording medium 1189, storage unit 1170, non-volatile memory 1108, or memory 1109.
  • step S1001 a process is performed to determine whether an operation to activate the AI assistant function has been performed.
  • the control unit 1110 determines whether an operation to activate the AI assistant function has been performed.
  • the operation to activate the AI assistant function can be, for example, an operation of pressing a physical button via the operation input unit 1107 or a button displayed on a touch panel, an operation of inputting a predetermined command by voice via the microphone 1139, or an operation of inputting a command using hand or facial gestures. If it is determined in this determination that an operation to activate the AI assistant function has been performed (S1001: Yes), the process proceeds to step S1004. On the other hand, if it is determined that an operation to activate the AI assistant function has not been performed (S1001: No), the process proceeds to step S1002.
  • step S1002 a process is performed to determine whether or not a condition for initiating a query as to whether or not the AI assistant function needs to be activated has been met.
  • the control unit 1110 determines whether or not a condition for initiating a query as to whether or not the AI assistant function needs to be activated has been met.
  • the condition may be, for example, detecting that the user has spent a certain amount of time or more deciding on a composition.
  • a more specific example of the condition may be, for example, that the shutter button has not been pressed for a certain period of time since the imaging device 1200 was set to imaging mode.
  • step S1002 If it is determined that the conditions for initiating an inquiry about whether the AI assistant function needs to be activated are met (S1002: Yes), the processing proceeds to step S1003. On the other hand, if it is determined that the conditions for initiating an inquiry about whether the AI assistant function needs to be activated are not met (S1002: No), the processing returns to step S1001.
  • step S1003 a process is performed to inquire of the user as to whether or not to activate the AI assistant function.
  • the control unit 1110 causes the display unit 10011 to display text information in natural language inquiring as to whether or not to activate the AI assistant function.
  • This process can also be considered as a process of inquiring of the user as to whether or not imaging assistance is required.
  • Fig. 2G is a diagram showing an example of a display screen in the imaging device according to Example 2.
  • Fig. 2G shows an example of the display of text information representing an interactive expression between the generated AI and the user using the AI assistant function.
  • the generated AI's words 10571 are displayed on the right side
  • the user's words 10572 are displayed on the left side.
  • the respective words are displayed in chronological order from top to bottom.
  • step S1003 for example, as shown in FIG. 2G, text information such as "Do you want to start the AI assistant?" is displayed.
  • the control unit 1110 may cause the audio output unit 1140 to output a voice such as "Do you want to start the AI assistant?" In this way, by asking the user whether to start the AI assistant function, the user can start the AI assistant function in situations where they are unaware of the AI assistant function, have forgotten about it, do not know how to start it, or find the start-up operation troublesome.
  • the control unit 1110 accepts the user's response to the inquiry as to whether or not the AI assistant function needs to be activated, i.e., whether or not imaging assistance is required. If the control unit 1110 accepts a response indicating that the AI assistant function needs to be activated, for example, by the user inputting or selecting "Yes" in response to the above inquiry (S1003: Yes), the processing step proceeds to step S1004.
  • control unit 1110 accepts a response indicating that the AI assistant does not need to be activated, for example, by the user inputting or selecting "No" in response to the above inquiry (S1003: No)
  • the processing related to the AI assistant function is terminated, or the parameters related to the conditions for starting the inquiry are reset, and the processing step returns to step S1001.
  • the user may respond by pressing a part of the screen of the display unit 10011, which functions as a touch panel, where "Yes” or “No” is displayed, by pressing a button corresponding to "Yes” or “No”, by saying “Yes” or “No”, or by inputting a response using gestures with their hands or facial expressions.
  • step S1004 processing is performed to acquire current location information.
  • the control unit 1110 receives current location information indicating the current location of the imaging device 1200 from the current location acquisition unit 1195.
  • the current location information is, for example, information indicating coordinates on the Earth, and is acquired using GPS.
  • the processing of step 1004 may be as follows.
  • the control unit 1110 receives from the imaging unit 1180 a through image that includes a landmark that can identify the current location.
  • the control unit 1110 transmits instruction information to the generation AI server that includes through image information at this current location and instructs the imaging device 1200 to provide its current location.
  • the generation AI server Upon receiving this instruction information, the generation AI server generates response information that includes current location information based on the through image that includes the landmark, and transmits this to the imaging device 1200.
  • the control unit 1110 receives this response information and acquires current location information based on the received response information.
  • the control unit 1110 may also include information indicating the name of the landmark entered by the user in the instruction information, thereby increasing the accuracy of determining the current location or the imaging point, which will be described later.
  • the generation AI can handle text information including natural language, colloquial expressions, etc. If the generation AI includes a multimodal LLM rather than just an LLM, the generation AI can handle information other than text information, such as image information and audio information, in addition to text information including natural language, colloquial expressions, etc.
  • the instruction information sent to the generation AI server is also called an instruction sentence, prompt, etc.
  • the generation AI server generates and returns response information in a form that responds to the instructions represented by the received instruction information. The user can receive assistance with imaging in natural language, easily understand the content of the assistance, and easily imagine the image that will be captured.
  • step S1005 processing is performed to instruct the server to provide an imaging point.
  • the control unit 1110 generates instruction information that includes current location information and instructs the server to provide an imaging point where an image of a landmark corresponding to the current location indicated by the current location information can be captured, and transmits this information to the generation AI server.
  • this instruction information can also be considered information that requests, as imaging support information, imaging point information that indicates an imaging point where an image of a landmark corresponding to the current location can be captured.
  • an imaging point is an example of an "imaging position" in this application
  • imaging point information is an example of "imaging position information" in this application.
  • the imaging point may be, for example, a standing or sitting position of the photographer that is considered suitable for capturing an image of a famous landmark near the current location.
  • the imaging point may be, for example, a standing or crouching position of the user that fits the entire landmark within the imaging area.
  • the imaging point may be, for example, a standing or crouching position of the user that results in a composition in which the landmark is placed at a specific position or area within the imaging area.
  • the instruction information may include the focal length of the lens attached to the imaging device, the variable range of the focal length of the zoom lens, etc.
  • the generation AI may determine the imaging point based on information related to the focal lengths of these lenses.
  • step S1006 a process is performed to generate response information including the imaging point. Specifically, upon receiving this instruction information, the generation AI server generates response information including the imaging point corresponding to the current location and transmits it to the imaging device.
  • step S1007 a process is performed to determine whether the current position is an imaging point. Specifically, the control unit 1110 compares the current position indicated by the current position information with the imaging point corresponding to the current position, and determines whether the current position is an imaging point based on whether these positions substantially overlap.
  • the control unit 1110 transmits to the generation AI server instruction information that includes current position information or through-image information at the current position and instructs the generation AI server to determine whether or not the current position is an imaging point corresponding to the current position.
  • the generation AI server determines whether or not the current position is an imaging point, generates response information that includes determination result information that indicates the determination result, and transmits this to the imaging device 1200.
  • the control unit 1110 receives this response information and determines whether or not the current position is an imaging point based on the received response information.
  • step S1011 If it is determined that the current position is an imaging point (S1007: Yes), the processing proceeds to step S1011. On the other hand, if it is determined that the current position is not an imaging point (S1007: No), the processing proceeds to step S1008.
  • step S1008 processing is performed to instruct the user to be guided to the imaging point.
  • the control unit 1110 generates instruction information that instructs the user to be guided to the imaging point corresponding to the current position indicated by the current position information, and transmits the generated instruction information to the generation AI server.
  • this instruction information can also be said to be information that requests guidance information for guiding the user to the imaging point as imaging support information.
  • step S1009 a process is performed to generate response information including guidance information for guiding the user to the imaging point. Specifically, upon receiving the instruction information, the generation AI server generates response information including guidance information for guiding the user to the imaging point and transmits it to the imaging device 1200.
  • the guidance information is text information in natural language, such as "Is this Cologne Cathedral? If you turn 1 meter to the right and 3 meters back, you can capture the whole picture.”
  • the generation AI server transmits the generated response information to the imaging device 1200.
  • step S1010 a process is performed to output guidance information to guide the user.
  • the control unit 1110 receives response information from the generation AI server, including information to guide the user to the imaging point.
  • the control unit 1110 also causes the display unit 10011 to display guidance information to guide the user based on the received response information.
  • the control unit 1110 may cause the audio output unit 1140 to output the contents of the guidance information as audio.
  • the guidance information may be information received from the generation AI server as is, or information generated based on the received information may be output. In this way, when the guidance information is output, the user can move to a suitable imaging point in a short amount of time without having to search for an imaging point. Thereafter, the processing returns to step S1004.
  • the control unit 1110 may calculate the difference between the imaging point and the current position, and generate guidance information for guiding the user to the imaging point based on that difference.
  • step S1011 processing is performed to acquire at least one of orientation information and through-image information.
  • the control unit 1110 receives the acquired orientation information of the image capture device 1200 from the orientation sensor 1113.
  • the control unit 1110 receives the acquired through-image information from the image capture unit 1180.
  • the control unit 1110 may receive both the orientation information and the through-image information.
  • a through-image is an image captured by the image capture unit 1180 in the background when the image capture mode is set in the image capture device 1200 and the shutter button is not pressed.
  • a through-image is acquired, for example, at a predetermined interval.
  • step S1012 a process is performed to instruct the system to give advice to the user on how to capture images.
  • the control unit 1110 generates instruction information that includes at least one of posture information and through-image information and instructs the system to give advice to the user on how to capture images, and transmits the generated instruction information to the generation AI server.
  • this instruction information can also be considered information that requests advice information to advise the user on how to capture images as imaging support information.
  • the instruction information may include information indicating a landmark in the through image.
  • the control unit 1110 specifies an arbitrary position or area in the image area of the through image in response to a user input operation, and determines an object corresponding to the specified position or area as the landmark.
  • the instruction information may include, for example, an instruction to send a model image to serve as a model.
  • step S1013 a process is performed to generate response information including advice information to advise the user on an appropriate method of capturing images.
  • the generation AI server upon receiving the instruction information, the generation AI server generates response information including advice information to advise the user on an appropriate method of capturing images based on at least one of the posture information and the through-image information.
  • the advice information may include, for example, information to instruct the user on the position and orientation of the imaging device 1200 to obtain an appropriate composition.
  • the advice information may be text information in natural language, such as "Hold the camera from the waist down,” “Arranging people diagonally will create a sense of depth,” or “Pressing the shutter button slowly will prevent blur.”
  • the advice information may be text information including natural language, such as "Placing landmarks on one of the left and right halves and people on the other will result in a well-balanced composition.”
  • the advice information may be an image on which an image, symbol, mark, etc. indicating the area where a landmark should be placed, the area where the person being photographed should be placed, or both of these areas is superimposed on the through image.
  • the advice information may include a model image, which is an exemplary image.
  • the content of the advice may be varied depending on the level of the user's imaging skill.
  • the imaging device 1200 may preset the user's imaging skill level from among multiple levels. These multiple levels may be, for example, three levels: advanced, intermediate, and beginner.
  • the imaging device 1200 may record how the imaging device 1200 has been handled in the past in the storage unit 1170 or the data recording unit 1188, and send information representing this handling to the generation AI, causing the generation AI to determine the user's imaging skill level.
  • the control unit 1110 instructs the generation AI to provide advice according to the user's imaging skill level. In this case, the user can receive advice that matches their own imaging skill level.
  • the control unit 1110 in cooperation with the generation AI server, presents the user with hints for achieving the optimum position of the imaging device 1200, the optimum orientation of the imaging device 1200, the optimum composition, the optimum imaging conditions, and preventing (suppressing) camera shake. If the user is a beginner, the control unit 1110 presents, for example, hints on settings related to aperture, shutter speed, focus, and zoom lens. If the user is an intermediate user, the control unit 1110 presents, for example, hints on settings related to ISO sensitivity, white balance (WB), macro lens, and how to deal with fast-moving subjects (suppressing subject shake).
  • WB white balance
  • the settings of the imaging device 1200 may be automatically set by the control unit 1110 based on the response information received from the generation AI server.
  • step S1014 a process is performed to output advice information to advise the user on how to take images.
  • the control unit 1110 receives response information from the generation AI server, including advice information to advise the user on an appropriate way to take images. Furthermore, based on the response information, the control unit 1110 causes the display unit 10011 to display the advice information to advise the user on an appropriate way to take images. In addition to or instead of the display, the control unit 1110 may cause the audio output unit 1140 to output the advice information to advise the user on the appropriate way to take images as audio.
  • the control unit 1110 may output the information received from the generation AI server as is, or may output information generated (edited) based on the received information.
  • the advice information for advising the user on the best way to take images is to be specific, as described above, it is necessary to use posture information or through-image information.
  • the advice information is to be general, it is not necessary to use posture information or through-image information.
  • general advice information may be stored in advance in the storage unit 1170, non-volatile memory 1108, or recording medium 1189 of the imaging device 1200, and the control unit 1110 may read and output the stored advice information.
  • the control unit 1110 does not need to send instruction information to the generation AI server. In this case, it is possible to reduce the amount of information sent and received to the generation AI server and the processing in the generation AI server, thereby shortening the time required to output advice information and reducing the load on the generation AI server.
  • FIG. 2H is a diagram showing an example of the change in composition before and after receiving advice from the generation AI in Example 2.
  • the composition on the left is a pre-advice composition 251 before the user receives advice from the generation AI
  • the composition on the right is a post-advice composition 252 after the user receives advice from the generation AI.
  • portions of the first landmark 10501, the second landmark 10502, and the two people being imaged 10505 each extend outside the imaging area (angle of view range) 1050.
  • the imaging area 1050 of the pre-advice composition 251 includes a first car 1504, two people on bicycles 10506, and a second car 10508 as moving objects moving from left to right.
  • two other people 10507 are included in the pre-advice composition 251.
  • the imaging position where the user stands or the magnification of the zoom lens is adjusted so that the first landmark 10501 and the second landmark 10502 do not protrude outside the imaging area 1050. Furthermore, the user has guided the person to be imaged 10505 or adjusted the magnification of the zoom lens so that the person to be imaged 10505 fits within the imaging area 1050. Furthermore, the two people to be imaged 10505 have been guided by the user so that they are positioned diagonally so that the three-dimensional effect of the subjects is created in the captured image. Note that no particular guidance is given to moving objects such as the first car 10504, the two people on bicycles 10506, and the second car 10508, or other people 10507.
  • step S1015 a process is performed to determine whether an operation to obtain a temporary processed image has been performed.
  • the control unit 1110 determines whether an operation to obtain a temporary processed image has been performed.
  • the operation to obtain a temporary processed image may be, for example, an operation of half-pressing the shutter button.
  • the shutter button is a two-stage switch, the operation of pressing the shutter button to the first stage is called a "half-press," and the operation of pressing it further to the second stage is called a "full press.”
  • the process proceeds to step S1016.
  • the process returns to step S1004.
  • step S1016 a process is performed to instruct the extraction of unnecessary objects.
  • the control unit 1110 generates instruction information that instructs the extraction of unnecessary objects that are considered unsuitable as subjects in the through image, and transmits the generated instruction information to the generation AI server.
  • this instruction information can also be said to be information that requests, as imaging support information, extraction result information that represents the results of extracting unnecessary objects that are considered unsuitable as subjects in the through image, which is a captured image.
  • This instruction information may be, for example, information that instructs the extraction, as unnecessary objects, of at least one of moving objects and non-imaging people from among objects in the through image, excluding landmarks and imaging target people.
  • FIG. 2I is a diagram showing an example of how an unwanted object is extracted from a through image obtained by an imaging device according to Example 2.
  • the upper row shows multiple through images 1053A, 1053B, 1053C, and 1053D in time series obtained by processing executed in the background.
  • a moving object can be defined as an object whose position changes relative to a landmark or other stationary subject in multiple time-series through-images 1053A-1053D.
  • a moving object can be defined as an object that moves across multiple grid-like reference lines extending horizontally and vertically in the image area of the through-images when image stabilization is in operation, in multiple time-series through-images.
  • a person not to be imaged can be defined as, for example, a person included in any of the through images 1053A to 1053D whose facial orientation or line of sight deviates from the direction facing the image capture device 1200 by more than a certain level.
  • a person not to be imaged can be defined as, for example, a person whose proportion of the image area of any of the through images 1053A to 1053D is less than a certain level.
  • subjects marked with black arrows have been extracted as moving objects or people not to be imaged.
  • a first automobile 10504, a second automobile 10508, two bicycles, and their drivers, people 10506 have been extracted as moving objects whose position has changed based on the first landmark 10501 or the second landmark 10502.
  • two other people 10507, small and facing sideways have been extracted as people not to be imaged, who occupy an area of the imaging area 1050 that is less than a certain percentage, or whose face or line of sight deviates by more than a certain level from the direction facing the imaging device.
  • Moving objects are likely to be passersby, vehicles, etc. Furthermore, people not to be photographed are likely to be strangers from the user's perspective. Therefore, moving objects or people not to be photographed are likely not objects or people that the user actually wants to photograph as subjects. If the user wants to remove such unwanted objects or people from the captured image by processing the image, they can have these objects or people selected automatically to a certain extent, reducing the need for complicated operations.
  • step S1017 a process is performed to generate response information including extraction result information that indicates the results of the extraction of the unnecessary object.
  • the generation AI server upon receiving the instruction information, the generation AI server generates response information including the extraction result information and transmits the generated response information to the imaging device 1200.
  • the extraction result information is, for example, an extraction result image in which a mark indicating the extracted unnecessary object has been added to a through image.
  • the extraction result information is information that, in addition to the extraction result image, includes text information including natural language such as "Moving objects and other people in the camera image have been automatically extracted.”
  • step S1018 a process is performed to output the extraction result information.
  • the control unit receives response information including the extraction result information from the generation AI server via the interface, and displays the extraction result information on the display unit 10011.
  • the control unit 1110 may also cause the audio output unit 1140 to output information related to the extraction result as audio, in addition to displaying the extraction result information.
  • the extraction result information may be the information received from the generation AI server that is output as is, or information that has been generated (edited) based on the received information may be output.
  • the extraction result information may also include text information in natural language, such as, for example, "Moving objects and other people in the camera image have been automatically extracted.”
  • step S1019 a process for selecting an object to be processed is performed.
  • the control unit 1110 initially sets all extracted unnecessary objects as candidates for processing.
  • the user can remove objects that the user does not want to process from the candidates for processing, or can specify objects other than the candidates for processing as candidates for processing.
  • the control unit 1110 selects an object to be processed in the through image in response to these user operations.
  • the control unit 1110 may select the extracted unnecessary object as the object to be processed as is, or, in a situation where unnecessary objects are not being extracted, may select an object desired by the user in the through image as the object to be processed in response to the user's operation.
  • step S1020 a process is performed to instruct the processing target to be processed in the through image.
  • the control unit 1110 generates instruction information that includes the through image and instructs the processing target to be processed in the through image, and transmits the generated instruction information to the generation AI server.
  • this instruction information can also be considered information that requests, as imaging support information, a processed image in which the processing target has been removed from the through image, for example, at least one of a moving object and a person not to be captured.
  • This instruction information may be, for example, information instructing the system to remove an image representing the object to be processed from the through image and restore the area from which the image was removed using image data from the surrounding area. In this way, it is possible to use AI to generate a good-looking image while using a realistic background included in the composition.
  • this instruction information may be information that instructs the system to remove an image representing the object to be processed from the through image and then repair the removed area using desired image data specified by the user. In this way, it is possible to generate a good-looking image that reflects the user's preferences.
  • a process is performed to generate response information including processing result information.
  • the generation AI server receives the instruction information, it processes the processing target in the through image to obtain a processed image, generates response information including the processing result information, and transmits the generated response information to the imaging device.
  • the processing result information includes the processed image.
  • the processing result information may also include text information in natural language, such as "The specified processing target was excluded and processing was performed using surrounding image data," in addition to the processed image.
  • step S1022 processing is performed to output the processing result information.
  • the control unit 1110 receives response information including the processing result information from the generation AI server, and causes the display unit 10011 to display the processed image included in the processing result information as a temporary processed image.
  • the control unit 1110 also causes the display unit to display text information included in the processing result information, and causes the audio output unit 1140 to output it as audio.
  • the control unit 1110 may also display the temporary processed image and the original image, which is a through image before processing, side by side on the display screen of the display unit 10011. This allows the user to compare the temporary processed image with the original image, and design the processed image to suit their own preferences.
  • Figure 2J is a diagram showing an example in which an original image and a temporarily processed image are displayed in an imaging device according to Example 2.
  • an original image 1055 and a temporarily processed image are displayed side by side on the display screen of the display unit 10011.
  • the imaging area 1050 includes a first automobile 10504, a second automobile 10508, two bicycles and their drivers, person 10506, and another person 10507 who is not to be imaged.
  • the temporary processed image 1056 includes only the first landmark 10501, the second landmark 10502, and the person 10505 to be imaged within the image capture area 1050, and the moving objects, the first automobile 10504, the second automobile 10508, the two bicycles and their drivers 10506, and the other person 10507, who is not to be imaged, have been erased.
  • step S1023 a process is performed to determine whether the operation to obtain a temporary processed image is being performed continuously. Specifically, the control unit 1110 determines whether the operation to obtain a temporary processed image is being performed continuously. If it is determined that the operation to obtain a temporary processed image is being performed continuously (S1023: Yes), the process returns to step S1023. On the other hand, if it is determined that the operation to obtain a temporary processed image is not being performed continuously (S1023: No), the process proceeds to step S1024.
  • step S1024 a process is performed to determine whether an operation to save an image has been performed. Specifically, the control unit 1110 determines whether an operation to save an image has been performed by the user. An example of an operation to save an image is an operation to fully press the shutter button. If it is determined that an operation to save an image has been performed (S1024: Yes), the process proceeds to step S1025. On the other hand, if it is determined that an operation to save an image has not been performed (S1024: No), the process returns to step S1008.
  • step S1025 the original image and the edited image are stored in association with each other.
  • the control unit 1110 stores the edited image, which is a temporary edited image, and the original image, which is a live view image before editing, in the storage unit 1170 or the recording medium 1189 in association with each other.
  • the file name of the edited image includes the capture number and letters, symbols, numbers, extension, etc. indicating that the image has been edited
  • the file name of the original image includes the same capture number and letters, symbols, numbers, extension, etc. indicating that it is the original image.
  • the control unit 1110 also causes the display unit to display image save report information indicating that the image has been saved.
  • the control unit 1110 may cause the audio output unit 1140 to output the image save report information as audio, or to output a specified sound from the audio output unit 1140.
  • the image save report information may be text information in natural language, such as, for example, "The original image data, which has not been processed, and the processed image data have been recorded in a format that can be distinguished by a human.” Once the image has been saved, the processing proceeds to step S1026.
  • the user can compare the original image with the edited image when checking the captured image later, and can select the image to use according to the user's convenience or preference.
  • step S1026 a process is performed to determine whether or not to terminate the AI assistant function. Specifically, the control unit 1110 determines whether or not to terminate the AI assistant function based on factors such as whether an operation to terminate the AI assistant function has been performed, and whether a certain amount of time has passed without the user using the AI assistant function. If this determination determines that the AI assistant function should be terminated, the control unit 1110 terminates the AI assistant function. On the other hand, if it determines that the AI assistant function should not be terminated, the process returns to step S1004.
  • the user can obtain hints for taking suitable images when taking pictures, and can take suitable images without relying on their own geographical knowledge (local familiarity), imaging skills, etc. For example, in a place they are visiting for the first time, a user can achieve a suitable composition, a suitable angle of view, and suitable settings for the imaging device, even if their personal imaging skills are not particularly high.
  • the imaging device according to Example 2 even if a user has high imaging skills, when taking a selfie, the orientation of the imaging device relative to the subject, the way the imaging device is held, the location of the photographer, etc., are different from those used in normal imaging, and the user may not be able to capture the image as desired. Even in such cases, with the imaging device according to Example 2, the user can take a more suitable image by using the AI assistant function.
  • the user can receive advice about imaging in colloquial language, which increases the affinity between the imaging device and the user and allows the user to intuitively understand the content of the advice.
  • the imaging device by using generation AI, particularly multimodal LLM, it is possible to handle image information and text information simultaneously, eliminating the need to switch the communication destination generation AI server depending on the type of information, and simplifying the processing of the control unit.
  • generation AI particularly multimodal LLM
  • the user can receive advice that meets more detailed requests. For example, by creating instruction information such as "Please tell me the optimal imaging conditions for the imaging device without using technical terms," the user can receive advice that is understandable even if they are a beginner and do not understand technical terms.
  • the imaging device With the imaging device according to Example 2, it is possible to preserve the sense of realism and memories (memories) of the original image when it was captured, and by eliminating (processing) and recording any extraneous objects other than those you want to remember when capturing the image, it is possible to create a high-quality image without the need for subsequent editing.
  • Example 2 it is also possible to have the generating AI recognize an object that moves from outside the angle of view into the image as the desired subject, or as an unwanted subject.
  • the generating AI recognizes an object that moves from outside the angle of view into the image as the desired subject, or as an unwanted subject.
  • you want to capture a subject near the center of the image horizontally by setting this in advance, you can request assistance from the generating AI that predicts when the subject will approach the center and instructs the user on when to release the shutter.
  • Example 2 when the LLM dialogue window and the image display window are displayed on the display unit, if the display unit has only one screen, the windows may be displayed side by side on the screen, or the windows may be displayed by switching between them. If the display unit has two screens, for example, one on the top and one on the back of the imaging device, the LLM dialogue window may be displayed on one screen, and the image display window may be displayed on the other screen.
  • Example 2 when an original image and a processed image, or a through image and a model image, are displayed on the display unit, if the display unit has only one screen, these two images may be displayed by switching between them one by one, or two images may be displayed side by side on one screen. If the display unit has two screens, one image may be displayed on each screen.
  • the preferred composition may be specified as one, or the user may select from multiple.
  • preferred compositions include a composition in which the landmark fits within the angle of view, a composition in which the subject person fits within the angle of view, a composition in which both the landmark and the subject person fit within the angle of view, and a composition that creates a three-dimensional effect.
  • the generation AI may propose multiple preferred arrangements, etc.
  • Example 2 a function may be implemented that allows an image that has been processed by the generation AI (by converting a portion of the background) to be later reconverted into an image desired by the user.
  • a moving object that enters the field of view from outside may be identified as an unwanted object, and the unwanted object may be removed from the captured image by processing.
  • Example 2 is an example in which the assist function using the generated AI is applied to capturing still images, it can also be applied to continuous shooting or video capture.
  • continuous shooting means capturing still images continuously (for example, at approximately 3 to 30 frames per second) while the shutter button is pressed.
  • the generated AI may be constructed on a server connected to the imaging device via a network, or may include both one constructed on the server and one built into the imaging device.
  • the control unit in the imaging device may transmit and receive data to and from the generated AI constructed on the server, or may transmit and receive data to and from the local generated AI built into the imaging device.
  • the imaging device is, for example, a digital camera, a digital movie camera, a personal computer with imaging function, a smartphone, a tablet terminal, etc.
  • the imaging device may also be, for example, a combination of a digital camera and a smartphone, etc., connected to each other so that they can communicate with each other.
  • the imaging unit included in the digital camera is used as the imaging unit in Example 2
  • the control unit included in the digital camera and the control unit included in the smartphone, etc. are used as the control unit in Example 2.
  • Communication between the imaging device and the generation AI server may be performed via a communication unit included in the smartphone, etc.
  • the imaging support method based on the processing flow executed in the imaging device and generation AI of Example 2 is one embodiment of the present invention.
  • a program for causing one or more processors (CPU, MPU, etc.) constituting the control unit of the imaging device of Example 2 to execute each process so that the imaging support method is implemented is also one embodiment of the present invention.
  • a tangible, non-transitory recording medium on which the program is recorded is also one embodiment of the present invention.
  • the program may be downloaded from a server and used.
  • this embodiment makes it possible to provide more suitable AI imaging assistant technology.
  • this AI imaging assistant technology is expected to be introduced into higher quality, more reliable infrastructure. Furthermore, by introducing this technology into infrastructure, it can contribute to supporting economic development and human welfare, with a focus on affordable and fair access for all. This will contribute to "Build resilient infrastructure, promote inclusive and sustainable industrialization, inclusive and sustainable technological development," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.
  • SDGs Sustainable Development Goals
  • this embodiment makes it possible to provide more suitable artificial intelligence imaging assistant technology.
  • this type of artificial intelligence imaging assistant technology when introduced into public transportation, can contribute to improving traffic safety through the expansion of public transportation, and realizing access to a safe, affordable, and easily usable sustainable transportation system for all people. This will contribute to "Sustainable cities and communities," one of the Sustainable Development Goals (SDGs) advocated by the United Nations.
  • SDGs Sustainable Development Goals
  • the present invention is not limited to the above-described embodiments and includes various modifications.
  • the above-described embodiments are detailed descriptions of the entire system in order to clearly explain the present invention, and are not necessarily limited to systems that include all of the configurations described.
  • it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment and it is also possible to add the configuration of another embodiment to the configuration of one embodiment.
  • a processor includes transistors and other circuits and is considered a circuitry or processing circuitry.
  • a processor may also be a programmed processor that executes programs stored in memory.
  • a circuit, unit, or means refers to hardware that is programmed to realize or executes the described functions.
  • the hardware may be any hardware disclosed in this specification or any hardware known to be programmed to realize or execute the described functions.
  • the hardware is a processor, which is considered to be a type of circuitry
  • the circuitry, means, or unit is the combination of the hardware and the software used to configure the hardware and/or processor.
  • 10010...AI response output device 10011...display unit, 10028...local LLM processing unit, 1107...operation input unit, 1110...control unit, 1132...communication unit, 1140...audio output unit, 1139...microphone, 1160...video control unit, 1170...storage unit, 1180...imaging unit, 1108...non-volatile memory, 1109...memory, 1113...orientation sensor, 1189...recording medium, 1195...current position acquisition unit, 1200...imaging device, 1201...imaging function unit, 19000...Internet, 19001...large-scale language model server, 20001...multimodal large-scale language model server

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

より好適な人工知能応答出力技術を提供すること。本発明によれば、持続可能な開発目標(SDGs)の「9産業と技術革新の基盤をつくろう」、「11住み続けられるまちづくりを」に貢献する。 撮像部と、表示部と、制御部と、を備え、制御部は、ユーザによる撮像部を用いた撮像を支援するための撮像支援情報を要求する指示情報を生成人工知能に送信する処理と、生成人工知能から撮像支援情報を含む応答情報を受信する処理と、受信した応答情報に含まれる撮像支援情報または当該撮像支援情報に基づく情報を表示部に表示させる処理と、を実行する、撮像装置を提供する。

Description

撮像装置、撮像支援方法、およびプログラム
 本発明は、撮像装置、撮像支援方法、およびプログラムに関する。
 言語モデルなどの人工知能を用いた応答出力技術については、例えば、特許文献1に開示されている。
特表2019-528512号公報
 しかしながら、特許文献1の開示では、人工知能を用いた応答出力技術をユーザにより好適に提供するための構成などについての考慮は十分ではなかった。
 本発明の目的は、より好適な応答出力技術を提供することにある。
 上記課題を解決するために、例えば特許請求の範囲に記載の構成を採用する。本願は上記課題を解決する手段を複数含んでいるが、その一例を挙げるならば、撮像装置であって、撮像部と、表示部と、制御部と、を備え、前記制御部は、ユーザによる前記撮像部を用いた撮像を支援するための撮像支援情報を要求する指示情報を生成人工知能に送信する処理と、前記生成人工知能から前記撮像支援情報を含む応答情報を受信する処理と、前記受信した応答情報に含まれる撮像支援情報または当該撮像支援情報に基づく情報を前記表示部に表示させる処理と、を実行するように構成すればよい。
 また、他の一例を挙げるならば、撮像支援方法であって、撮像装置における制御部が、ユーザによる撮像部を用いた撮像を支援するための撮像支援情報を要求する指示情報を生成人工知能に送信する処理と、前記生成人工知能から前記撮像支援情報を含む応答情報を受信する処理と、前記受信した応答情報に含まれる撮像支援情報または当該撮像支援情報に基づく情報を表示部に表示させる処理と、を実行するように構成すればよい。
 また、他の一例を挙げるならば、プログラムであって、撮像装置における制御部を構成するプロセッサに、ユーザによる撮像部を用いた撮像を支援するための撮像支援情報を要求する指示情報を生成人工知能に送信する処理と、前記生成人工知能から前記撮像支援情報を含む応答情報を受信する処理と、前記受信した応答情報に含まれる撮像支援情報または当該撮像支援情報に基づく情報を表示部に表示させる処理と、を実行させるように構成すればよい。
 本発明によれば、より好適な応答出力技術を提供できる。これ以外の課題、構成および効果は、以下の実施形態の説明において明らかにされる。
本発明の一実施例に係る人工知能応答出力装置およびシステムの一例を示す図である。 本発明の一実施例に係る人工知能応答出力装置の一例を示す図である。 本発明の一実施例に係る人工知能応答出力装置およびシステムの動作の一例を示す図である。 本発明の一実施例に係る撮像装置と生成人工知能サーバとを含むシステムの一例を示す図である。 本発明の一実施例に係る撮像装置の構成の一例を示す図である。 本発明の一実施例に係る撮像装置および生成人工知能サーバにおける処理フローの一例を示す図である。 本発明の一実施例に係る撮像装置および生成人工知能サーバにおける処理フローの一例を示す図である。 本発明の一実施例に係る撮像装置および生成人工知能サーバにおける処理フローの一例を示す図である。 本発明の一実施例に係る撮像装置および生成人工知能サーバにおける処理フローの一例を示す図である。 本発明の一実施例に係る撮像装置における表示画面の一例を示す図である。 本発明の一実施例に係る生成人工知能からの助言を受ける前後における構図の変化の一例を示す図である。 本発明の一実施例に係る撮像装置により得られるスルー画像において不要物体が抽出される様子の一例を示す図である。 本発明の一実施例に係る撮像装置において元画像と仮加工済画像とが表示された例を示す図である。
 以下、本発明の実施の形態を図面に基づいて詳細に説明する。なお、本発明は実施例の説明に限定されるものではなく、本明細書に開示される技術的思想の範囲内において当業者による様々な変更および修正が可能である。また、本発明を説明するための全図において、同一の機能を有するものには、同一の符号を付与し、その繰り返しの説明は省略する場合がある。
 なお、本発明の各実施例に係る人工知能応答出力装置が表示画面を有する場合は、表示装置と呼んでもよい。人工知能応答出力装置が音声出力機能を有する場合は、音声出力装置と呼んでもよい。人工知能応答出力装置は単に情報処理装置と呼んでもよい。人工知能応答出力装置と、大規模言語モデルを保持する大規模言語モデルサーバを含むシステムを人工知能応答出力システムと呼んでもよい。また、人工知能応答出力装置が人工知能である大規模言語モデルの応答サービスをユーザに提供し、ユーザの助力になる場合は、人工知能応答出力装置または人工知能応答出力装置の表示出力は、ユーザにとって人工知能(AI)アシスタントとなることができる。よって、この場合、人工知能応答出力装置はAIアシスタント装置またはAIアシスタント表示装置と呼んでもよい。同様に、この場合、人工知能応答出力装置と、大規模言語モデルを保持する大規模言語モデルサーバを含むシステムをAIアシスタントシステムまたはAIアシスタント表示システムと呼んでもよい。また、この場合、人工知能応答出力装置は、ユーザと人工知能の間のインタフェースとなるので、人工知能インタフェース装置と呼んでもよい。この場合、人工知能応答出力装置と、大規模言語モデルを保持する大規模言語モデルサーバを含むシステムを人工知能インタフェースシステムと呼んでもよい。
 <実施例1>
 本発明の実施例1として、大規模言語モデル人工知能からの応答を出力する人工知能応答出力装置およびそのシステムについて、説明する。
 図1Aを用いて、本発明の人工知能応答出力装置10010の一例について説明する。また、当該人工知能応答出力装置10010が大規模言語モデルサーバ19001と通信などにより連携する場合について、人工知能応答出力装置10010が大規模言語モデルサーバ19001および/またはマルチモーダルな大規模言語モデルサーバ20001とを含むシステムの一例について説明する。
 図1Aの例では、人工知能応答出力装置10010は、表示部10011を有する。図1Aの例では、表示部10011は、平面ディスプレイでもよく、背面から映像を投影するスクリーンでもよく、光学像を空中に結像する空中浮遊映像でもよい。表示部10011が平面ディスプレイの場合は、液晶パネルとバックライトを有する液晶ディスプレイでもよい。また、表示部10011は、プラズマディスプレイでもよい。表示部10011は、画素が自発光する有機ELディスプレイでもよい。また、表示部10011はタッチ操作入力センサを設け、タッチパネルとして構成してもよい。
 図1Aの例では、人工知能応答出力装置10010が備える音声出力部1140はスピーカで構成されている。また、人工知能応答出力装置10010がマイク1139を備え、ユーザの声を収音できる。当該マイク1139からの音声入力や、後述する操作入力部を介したユーザの操作入力により、人工知能応答出力装置10010は、人工知能である大規模言語モデルへの指示文(プロンプト)の元となるユーザ入力を取得することができる。
 人工知能応答出力装置10010は、人工知能応答出力装置10010にローカルの大規模言語モデルを備えてもよい。この場合、当該大規模言語モデルの応答を上記表示部10011の表示出力、および/または、音声出力部1140の音声出力として出力してもよい。
 また、人工知能応答出力装置10010は、ローカルの大規模言語モデルを備えず、外部の大規模言語モデルサーバ19001と通信し、大規模言語モデルサーバ19001から受信する応答を上記表示部10011の表示出力、および/または、音声出力部1140の音声出力として出力してもよい。
 または、人工知能応答出力装置10010は、ローカルの大規模言語モデルも備え、さらに、大規模言語モデルを有する外部の大規模言語モデルサーバ19001またはマルチモーダルな大規模言語モデルを有する外部の大規模言語モデルサーバ20001と通信するように構成してもよい。この場合、当該ローカルの大規模言語モデルの応答と、大規模言語モデルサーバ19001の大規模言語モデルまたはマルチモーダルな大規模言語モデルサーバ20001のマルチモーダルな大規模言語モデルから受信する応答とを、切り替えて、いずれか一方を、上記表示部10011の表示出力、および/または、音声出力部1140の音声出力として出力してもよい。または、当該ローカルの大規模言語モデルの応答と、大規模言語モデルサーバ19001の大規模言語モデルまたはマルチモーダルな大規模言語モデルサーバ20001のマルチモーダルな大規模言語モデルから受信する応答の両者にもとづいて生成した応答を、上記表示部10011の表示出力、および/または、音声出力部1140の音声出力として出力してもよい。
 人工知能応答出力装置10010が外部の大規模言語モデルサーバ19001または大規模言語モデルサーバ20001と通信して連携する場合の構成は、以下のとおりである。人工知能応答出力装置10010は、通信部1132を介して、インターネット19000に接続された通信装置19011と通信可能である。図1Aの例では、通信部1132と通信装置19011との通信は無線の例を示しているが、有線通信でも構わない。通信部1132と通信装置19011までの通信経路において、有線の部分と無線の部分があってもよいし、ルータや中継器を経由してもよい。また、通信部1132からインターネット19000までの通信経路において、有線の部分と無線の部分があってもよいし、ルータや中継器を経由してもよい。人工知能応答出力装置10010は、通信装置19011およびインターネット19000を介して、大規模言語モデルサーバ19001と通信可能である。また、人工知能応答出力装置10010は、通信装置19011およびインターネット19000を介して、大規模言語モデルサーバ19001または大規模言語モデルサーバ20001、およびこれらのサーバと異なる第2のサーバ19002と通信可能である。人工知能応答出力装置10010と大規模言語モデルサーバ19001または大規模言語モデルサーバ20001とを含めた構成を一つのシステムとして考えてもよい。
 以下の説明において、単に「大規模言語モデル」と説明し特に断りがない場合は、人工知能応答出力装置10010が備えるローカルの大規模言語モデルと、大規模言語モデルサーバ19001が備える大規模言語モデル、および大規模言語モデルサーバ20001が備えるマルチモーダルな大規模言語モデルを含めた概念と考えてよい。
 図1Aの例では、表示部10011が、ユーザから人工知能である大規模言語モデルへの指示文(プロンプト)を入力する指示文表示領域10051と、大規模言語モデルからの応答を表示する人工知能応答表示領域10061の2つの表示領域に各要素を表示する例を示している。図1Aの例では、指示文表示領域10051には、ユーザを示すアイコン10052、指示文の構成要素としての自然言語やソフトウェアコードなどのテキスト10053、指示文の構成要素としての画像10054、指示文の構成要素としての動画10055、などを表示する例を示している。図1Aの例では、人工知能応答表示領域10061には、人工知能または人工知能アシスタントを示すアイコン10062、人工知能からの応答の構成要素としての自然言語やソフトウェアコードなどのテキスト10063、人工知能からの応答の構成要素としての画像10064、人工知能からの応答の構成要素としての動画10065、などを表示する例を示している。なお、図1Aに示す人工知能応答出力装置10010の表示部10011の表示例は、あくまで一例である。人工知能応答出力装置10010が用いられる実装例に応じて、図1Aに示す例とは異なる表示を行えばよい。
 ここで、大規模言語モデルについて説明する。大規模言語モデルはLLM(Large Language Model)とも表記される。具体的には、GPT-1、GPT-2、GPT-3、InstructGPT、ChatGPTなど様々なモデルが公開されている。本実施例においてもこれらの技術を用いればよい。なお、これらの大規模言語モデルは人間界に存在する数多くの文書、テキストに含まれる自然言語を対象に、大規模な事前学習が行われて生成された人工知能モデルである。人工知能モデルのパラメータ数は億を超える。さらに、これに加えて、人間からのフィードバックにもとづく強化学習を施したモデルもある。ベースとなるモデルの一例はTransformerと呼ばれるモデルなどである。これらのモデルの学習の一例として、例えば、参考文献1などが公開されている。
 [参考文献1]
 Long Ouyang, et. al. “Training language models to follow instructions with human feedback”, https://arxiv.org/pdf/2203.02155.pdf
 これらの大規模言語モデルは、自然言語を対象とした翻訳、自然言語を対象とした文章校正、自然言語を対象とした文章要約、などが可能である。そのうち、高度なものでは、自然言語による質問回答(対話または会話ともよばれる)、自然言語による提案生成、プログラミングコードの生成などが可能である。これらの人工知能モデルのパラメータ数は非常に大きいため、学習には膨大なデータ、計算資源が必要である。よって、特定の用途に限って、このレベルの人工知能の学習を行うことは非常に資源効率が悪い。そこで、様々な用途に応用できる基盤モデル(Foundation Model)として、大規模な事前学習を行ってモデルが生成されている。例えば、図1Aに示す大規模言語モデルサーバ19001は、このような大規模言語モデルを備え、API(Application Programming Interface)を介して、様々な端末で利用できるように構成してもよい。また、図1Aに示す人工知能応答出力装置10010がローカルの大規模言語モデルを備え、人工知能応答出力装置10010自身が利用するように構成してもよい。いずれの大規模言語モデルの学習自体は、別途、大規模な事前学習を行って生成し、生成した大規模言語モデルを複製して、大規模言語モデルサーバ19001や人工知能応答出力装置10010などに備えればよい。このように、用途ごとや端末ごとに事前学習を行うのではなく、大規模な事前学習を行って生成した基盤モデルである大規模言語モデルを複製して個々のサーバや端末で利用すれば、学習に用いる資源消費を共有することができるため資源効率がよい。
 なお、大規模な事前学習を行って生成した基盤モデルとしての大規模言語モデルであっても、個々のサーバや装置において、用途や目的に応じて転移学習などの追加学習を行うように構成してもよい。
 また、大規模言語モデルは、自然言語を事前学習し、自然言語を対象とした入出力処理を行うことができる。さらに、自然言語のテキスト情報に加えて自然言語のテキスト情報以外の種類の情報も併せて処理が可能なマルチモーダルな大規模言語モデル人工知能も本発明の実施例に適用可能である。図1Aにおいては、マルチモーダルな大規模言語モデルを有する大規模言語モデルサーバ20001を示す。例えば、マルチモーダルな大規模言語モデル人工知能の一例としては、具体的には、GPT-4(参考文献2参照)、Gato(参考文献3参照)などが公開されている。本実施例においてもこれらの技術を用いればよい。なお、これらのマルチモーダルな大規模言語モデルは人間界に存在する数多くの文書、テキストに含まれる自然言語および自然言語のテキスト情報以外の種類の情報(例えば、画像、動画、音声など)を対象に、大規模な事前学習が行われて生成された人工知能モデルである。さらに、これに加えて、人間からのフィードバックにもとづく強化学習を施したモデルもある。以下、画像、動画、音声などの自然言語のテキスト情報以外の種類の情報を非自然言語情報源と称してもよい。
 [参考文献2]
 Open AI “GPT-4 Technical Report”, https://cdn.openai.com/papers/gpt-4.pdf
 [参考文献3]
 Scott Reed, et. al. “A Generalist Agent”,
https://arxiv.org/pdf/2205.06175.pdf
 次に、図1Bを用いて、これらの大規模言語モデルなどの人工知能に対するユーザからの入力を受け付け、当該ユーザからの入力に対する大規模言語モデルなどの人工知能からの応答を出力する、人工知能応答出力装置10010の構成例について説明する。
 人工知能応答出力装置10010は、表示部10011、制御部1110、メモリ1109、不揮発性メモリ1108、外部電源入力インタフェース1111、操作入力部1107、電源1106、二次電池1112、ストレージ部1170、映像制御部1160、姿勢センサ1113、通信部1132、音声出力部1140、マイク1139、映像信号入力部1131、音声信号入力部1133、撮像部1180、等を備えている。人工知能応答出力装置10010は、例えば、いわゆるモニタやテレビなどの大画面を有するものでもよい。
 表示部10011は、平面ディスプレイでもよく、背面から映像を投影するスクリーンでもよく、光学像を空中に結像する空中浮遊映像を表示するものでもよい。表示部10011が平面ディスプレイの場合は、液晶パネルとバックライトを有する液晶ディスプレイでもよい。また、表示部10011は、プラズマディスプレイでもよい。表示部10011は、画素が自発光する有機ELディスプレイでもよい。表示部10011がパネルの場合は、表示パネルと称してもよい。表示部10011にタッチ操作入力センサを設け、ユーザ230の指によるタッチ操作入力を受け付けるように構成してもよい。この場合、表示部10011はタッチパネルとして構成してもよい。当該タッチパネルを介したユーザの操作入力により、人工知能応答出力装置10010は、人工知能である大規模言語モデルへの指示文(プロンプト)の元となるユーザ入力を取得することができる。
 通信部1132は、Wi―Fi方式の通信インタフェース、Bluetooth(登録商標)方式の通信インタフェース、4G、5Gなどの移動体通信インタフェースなどで構成すればよい。これらの通信方式を用いて、人工知能応答出力装置10010の通信部1132は、インターネット19000に接続された通信装置19011と通信可能である。なお、通信部1132と通信装置19011までの通信経路において、有線の部分と無線の部分があってもよいし、ルータや中継器を経由してもよい。有線の場合は、通信部1132は、ハードウェアとしてイーサネットの接続インタフェースを有してLAN方式の通信方式を用いて通信を行ってもよい。これにより、人工知能応答出力装置10010はインターネット19000に接続された各種サーバと通信可能である。
 人工知能応答出力装置10010にはCPUなどの制御部1110およびメモリ1109が備えられており、当該制御部1110は、表示部10011や通信部1132などを制御する。
 電源1106は、外部から外部電源入力インタフェース1111を介して入力されるAC電流をDC電流に変換し、人工知能応答出力装置10010の各部にそれぞれ必要なDC電流を供給する。二次電池1112は、電源1106から供給される電力を蓄電する。また、二次電池1112は、外部電源入力インタフェース1111を介して、外部から電力が供給されない場合に、電力を必要とする各部に対して電力を供給する。
 操作入力部1107は、例えば操作ボタンや、リモートコントローラ等の信号受信部または赤外光受光部であり、表示部10011のタッチ操作入力センサへのユーザによるタッチ操作とは異なる操作についての信号を入力する。表示部10011のタッチ操作入力センサをタッチ操作するユーザとは別に、操作入力部1107は、例えば管理者が人工知能応答出力装置10010を操作するために用いられてもよい。当該操作入力部1107を介したユーザの操作入力により、人工知能応答出力装置10010は、人工知能である大規模言語モデルへの指示文(プロンプト)の元となるユーザ入力を取得することができる。なお、表示部10011のタッチ操作入力センサも前記操作入力部1107の一部として含む構成とする変形例もあり得る。
 映像信号入力部1131は、外部の映像出力装置を接続して映像データを入力する。映像信号入力部1131は、様々なデジタル映像入力インタフェースが考えられる。例えば、HDMI(登録商標)(High―Definition Multimedia Interface)規格の映像入力インタフェース、DVI(Digital Visual Interface)規格の映像入力インタフェース、またはDisplayPort規格の映像入力インタフェースなどで構成すればよい。または、アナログRGBや、コンポジットビデオなどのアナログ映像入力インタフェースを設けてもよい。映像信号入力部1131は、各種USBインタフェースなどでもよい。
 音声信号入力部1133は、外部の音声出力装置を接続して音声データを入力する。音声信号入力部1133は、HDMI規格の音声入力インタフェース、光デジタル端子インタフェース、または、同軸デジタル端子インタフェース、などで構成すればよい。音声信号入力部1133は、各種USBインタフェースなどでもよい。HDMI規格のインタフェースの場合は、映像信号入力部1131と音声信号入力部1133とは、端子およびケーブルが一体化したインタフェースとして構成されてもよい。
 音声出力部1140は、音声信号入力部1133に入力された音声データに基づいた音声を出力することが可能である。音声出力部1140は、ストレージ部1170に格納されている音声データに基づいた音声を出力することも可能である。音声出力部1140は、スピーカで構成してもよい。また、音声出力部1140は、内蔵の操作音やエラー警告音を出力してもよい。または、HDMI規格に規定されるAudio Return Channel機能のように、外部機器にデジタル信号として音声信号を出力する構成を音声出力部1140としてもよい。または、ヘッドホンなどの外部機器にアナログ信号として音声信号を出力する構成を音声出力部1140としてもよい。
 マイク1039は、人工知能応答出力装置10010の周辺の音を収音し、信号に変換して音声信号を生成するマイクである。ユーザの声など人物の声をマイクが収録して、生成した音声信号を後述する制御部1110が音声認識処理を行って、当該音声信号から文字情報を取得するように構成してもよい。当該マイク1139からの音声入力により、人工知能応答出力装置10010は、人工知能である大規模言語モデルへの指示文(プロンプト)の元となるユーザ入力を取得することができる。
 撮像部1180は、イメージセンサを有するカメラである。人工知能応答出力装置10010の表示部10011側の前面にカメラを設けてもよく、表示部10011側の背面にカメラを設けてもよい。前面のカメラと背面のカメラの両者を設けてもよい。本実施例では、撮像部1180は、前面のカメラと背面のカメラの両者を有するものとして説明する。
 ストレージ部1170は、映像データ、画像データ、音声データ等の各種データなどの各種情報を記録する記憶装置である。ストレージ部1170は、ハードディスクドライブ(HDD)などの磁気記録媒体記録装置や、ソリッドステートドライブ(SSD)などの半導体素子メモリで構成してもよい。ストレージ部1170には、例えば、製品出荷時に予め映像データ、画像データ、音声データ等の各種データ等の各種情報が記録されていてもよい。また、ストレージ部1170は、通信部1132を介して外部機器や外部のサーバ等から取得した映像データ、画像データ、音声データ等の各種データ等の各種情報を記録してもよい。ストレージ部1170に記録された映像データ、画像データ等は、表示部10011に出力される。ストレージ部1170に記録された映像データ、画像データ等を、通信部1132を介して外部機器や外部のサーバ等に出力してもよい。
 映像制御部1160は、表示部10011に入力する映像信号に関する各種制御を行う。映像制御部1160は、映像処理回路と称してもよく、例えば、ASIC、FPGA、映像用プロセッサなどのハードウェアで構成されてもよい。なお、映像制御部1160は、映像処理部、画像処理部と称してもよい。映像制御部1160は、例えば、メモリ1109に記憶させる映像信号と、映像信号入力部1131に入力された映像信号(映像データ)等のうち、どの映像信号を表示部10011に入力するかといった映像切り替えの制御等を行う。また、映像制御部1160は、映像信号入力部1131から入力された映像信号やメモリ1109に記憶させる映像信号等に対して画像処理を行う制御を行ってもよい。画像処理としては、例えば、画像の拡大、縮小、変形等を行うスケーリング処理、輝度を変更するブライト調整処理、画像のコントラストカーブを変更するコントラスト調整処理、画像を光の成分に分解して成分ごとの重みづけを変更するレティネックス処理等がある。
 姿勢センサ1113は、重力センサまたは加速度センサ、またはこれらの組み合わせにより構成されるセンサであり、人工知能応答出力装置10010の姿勢を検出することができる。姿勢センサ1113の姿勢検出結果に基づいて、制御部1110が、接続される各部の動作を制御してもよい。
 不揮発性メモリ1108は、人工知能応答出力装置10010で用いる各種データを格納する。不揮発性メモリ1108に格納されるデータには、例えば、人工知能応答出力装置10010の表示部10011に表示する各種操作用のデータ、表示アイコン、ユーザの操作が操作するためのオブジェクトのデータやレイアウト情報等が含まれる。メモリ1109は、表示部10011に表示する映像データや装置の制御用データ等を記憶する。制御部1110がストレージ部1170から各種ソフトウェアを読み出して、メモリ1109に展開して記憶してもよい。
 ローカルLLM処理部10028は、大規模言語モデル(LLM)を保持できるメモリを備え、制御部1110の制御にもとづいて、大規模言語モデルの推論を実行できる。ハードウェアとしてはいわゆるGPU(Graphics Processing Unit)などで構成すればよい。ローカルLLM処理部10028は、推論のみならず、学習を行ってもよい。なお、人工知能応答出力装置10010のローカル環境での大規模言語モデルの推論の実行が不要な場合などは必ずしも、ローカルLLM処理部10028を要しない。
 制御部1110は、接続される各部の動作を制御する。また、制御部1110は、メモリ1109に記憶されるプログラムと協働して、人工知能応答出力装置10010内の各部から取得した情報に基づく演算処理を行ってもよい。制御部1110による制御状態には、例えば、ローカルLLM処理部10028の大規模言語モデルからの応答や、通信部1132を介して取得した、大規模言語モデルサーバ19001の大規模言語モデルまたはマルチモーダルな大規模言語モデルサーバ20001のマルチモーダルな大規模言語モデルからの応答を、表示部10011またはスピーカ等である音声出力部1140を介して出力する状態がある。
 なお、上述のタッチパネル、マイク1139または操作入力部1107を介してユーザから入力があった場合に、当該入力にもとづいて指示文を生成し、人工知能応答出力装置10010が備えるローカルLLM処理部10028のローカルの大規模言語モデル、大規模言語モデルサーバ19001が備える大規模言語モデル、または大規模言語モデルサーバ20001が備えるマルチモーダルな大規模言語モデルへ送信し、これらの大規模言語モデルから応答を取得する制御は、いずれも制御部1110が行えばよい。
 また、ストレージ部1170には、人工知能応答出力装置10010の指示文の応答として定型文を出力するための応答定型文データベース(応答定型文DBと表記してもよい)を格納してもよい。制御部1110が当該応答定型文データベースに格納されるデータを用いて出力する応答を生成する制御を行えばよい。図1Cに応答定型文データベースの一例を示す。図1Cの例では、条件番号が付された各条件について、人工知能応答出力装置10010が出力する定型文の応答が格納されている。例えば、条件番号1のように、上述のタッチパネル、マイク1139または操作入力部1107を介してユーザから「おはようございます」が入力された場合、応答定型文として「おはようございます」または「今日は〇月〇日ですね。」を用いて応答を出力すればよい。「〇月〇日」など〇の部分は、人工知能応答出力装置10010が有するメモリ1109などに格納される情報を用いて生成すればよい。
 また、図1Cに示すデータベースにおける応答定型文の例において、/で区切られた複数の応答定型文が格納されている場合は、制御部1110は、乱数などを用いてランダムに、いずれかの応答定型文を選択して応答が出力されるように制御すればよい。このようにすると、同一の条件における応答が単調になるという状況を解消して改善することができる。条件番号2、3、および4の例についても、条件番号1の例の説明と同様である。図1Cに示される各例の条件内容に対して、図1Cに示される各例の応答定型文を用いた出力をするように、制御部1110が制御を行えばよい。
 次に、図1Cに示される条件番号5の例について説明する。条件番号5は、タッチパネル、マイク1139または操作入力部1107を介して取得したユーザ入力について、制御部1110が自然言語として意味が理解できなかった場合またはユーザ入力に明らかに文法の誤りがある場合に、制御部1110が応答定型文として「ちょっと聞き取れませんでした」または「それについてはわからないかもしれません」を用いて応答を出力する制御を行う例である。このように応答することでユーザに再度の入力を促すことができ、修正されたユーザ入力を待つことができる。
 次に、図1Cに示される条件番号6の例について説明する。条件番号6は、図1Bに示す人工知能応答出力装置10010を構成する各部のいずれかの部においてエラー(異常状態)であることを制御部1110が検出している状態で、タッチパネル、マイク1139または操作入力部1107を介してユーザ入力があった場合の例である。この場合、制御部1110は応答定型文として「調子が悪いみたいです」を用いて応答を出力する制御を行う。このように応答することでユーザに人工知能応答出力装置10010が不調であることを説明することができ、ユーザにエラー対応などを促すことができる。
 人工知能応答出力装置10010は、人工知能応答出力装置10010が備えるローカルの大規模言語モデル、大規模言語モデルサーバ19001が備える大規模言語モデル、および大規模言語モデルサーバ20001が備えるマルチモーダルな大規模言語モデル、などの大規模言語モデルの応答に替えて、図1Cを用いて説明した応答定型文データベース(応答定型文DB)を用いた応答を出力してもよい。または、これらの大規模言語モデルの応答と応答定型文データベース(応答定型文DB)を用いた応答を組み合わせた応答を出力してもよい。
 なお、以上説明した図1Cの応答定型文データベース(応答定型文DB)はストレージ部1170に格納され、人工知能応答出力装置10010の制御部1110がこれを用いればよい。しかしながら、図1Cに示す応答定型文データベース(応答定型文DB)を大規模言語モデルサーバ19001側または大規模言語モデルサーバ20001側に備えてもよい。この場合は、大規模言語モデルサーバ19001が有する制御部または大規模言語モデルサーバ20001が有する制御部が、当該応答定型文データベース(応答定型文DB)を用いた応答を生成すればよい。大規模言語モデルサーバ19001が有する制御部または大規模言語モデルサーバ20001が有する制御部は、それぞれのサーバに格納される大規模言語モデルにより生成する応答に替えて、応答定型文データベース(応答定型文DB)を用いて生成した応答を、人工知能応答出力装置10010へ送信すればよい。このようにすれば、人工知能応答出力装置10010に応答定型文データベース(応答定型文DB)が備えられていない場合でも、応答定型文データベース(応答定型文DB)を用いた応答の生成が可能となる。
 なお、上記の説明では、人工知能応答出力装置10010は固定画素を用いた表示画面の表示パネルを有すると説明した。当該概念には、固定画素を用いた表示画面の表示パネルのあとに投射光学系を設けて、当該表示画面の表示パネルの映像の光学像をスクリーンや壁に投射する、投射型映像表示装置(プロジェクタ)を含んでもよい。
 なお、図1Aおよび図1Bの例では、人工知能応答出力装置10010が表示部10011を備える例を説明した。しかしながら、本発明の実施例にかかる人工知能応答出力装置10010は必ずしも表示部10011を備えていなくてもよい。例えば、表示部10011を備えなくとも、人工知能に対するユーザからの入力を音声信号入力部1133またはマイク1139を介して受け付け、当該ユーザからの入力に対する大規模言語モデルなどの人工知能からの応答を音声出力部1140を介して出力するように構成すればよい。
 以上説明した本発明の実施例1に係る、人工知能応答出力装置および人工知能応答出力システムによれば、大規模言語モデルなどの人工知能に対するユーザからの入力を受け付け、ネットワーク上のサーバ装置が有する大規模言語モデルまたは人工知能応答出力装置自体が有するローカルの大規模言語モデルなどの人工知能の推論により生成された、当該ユーザからの入力に対する応答を出力することが可能となる。
 また、本実施例に係る技術では、より好適な人工知能応答出力技術を提供することが可能となる。このような人工知能応答出力技術は、より質の高い、より信頼できるインフラへの導入が期待できる。当該技術がインフラへ導入されていくことにより、全ての人々に安価で公平なアクセスに重点を置いた経済発展と人間の福祉の支援に寄与できる。これにより、国連の提唱する持続可能な開発目標(SDGs:Sustainable Development Goals)の「9産業と技術革新の基盤をつくろう」に貢献する。
 また、本実施例に係る技術では、より好適な人工知能応答出力技術を提供することが可能となる。このような人工知能応答出力技術は、脆弱な立場にある人々の輸送システムへのアクセス性を向上させるために、公共交通機関の設備への導入が期待できる。当該技術が公共交通機関へ導入されていくことにより、公共交通機関の拡大などを通じた交通の安全性改善、および、全ての人々に安全かつ安価で容易に利用できる、持続可能な輸送システムへのアクセスの実現に寄与できる。これにより、国連の提唱する持続可能な開発目標(SDGs:Sustainable Development Goals)の「11住み続けられるまちづくりを」に貢献する。
 <実施例2>
 本発明の実施例2は、人工知能応答出力技術が適用された撮像装置および撮像支援方法の例である。実施例2に係る撮像装置および撮像支援方法は、人工知能応答出力技術を用いたAIアシスタント機能により、ユーザによる撮像を支援する。
 図2Aは、実施例2に係る撮像装置と生成人工知能(以下、人工知能のことをAIともいう)サーバとを含むシステムの一例を示す図である。図2Aに示すように、実施例2に係るシステムは、撮像装置1200と、大規模言語モデル(以下、大規模言語モデルのことをLLMともいう)サーバ19001と、マルチモーダルなLLMサーバ20001と、これらのサーバとは異なる第2のサーバ19002とを含む。撮像装置1200と、LLMサーバ19001、マルチモーダルなLLMサーバ20001、および第2のサーバ19002とは、通信ネットワークであるインターネット19000を介して、接続されている。なお、LLMサーバ19001と、マルチモーダルなLLMサーバ20001は、それぞれ、生成AIが構築された生成AIサーバの一例である。
 撮像装置1200は、撮像部1180を有する撮像機能部1201と、人工知能応答出力装置10010とを有している。人工知能応答出力装置10010は、実施例1における人工知能応答出力装置10010とほぼ同様の構成および機能を有している。
 図2Bは、実施例2に係る撮像装置の構成の一例を示す図である。撮像機能部1201は、撮像部1180、撮像画像入力部1181、外部画像入力部1182、画像入力調整部1183、画像音声メモリ部1184、画像エンコード部1185、音声エンコード部1186、データ生成部1187、データ記録部1188、記録媒体1189、現在位置取得部1195、映像信号入力部1131、および音声信号入力部1133を有している。
 撮像部1180は、被写体の撮像画像に対応した画像情報を得るためのものである。撮像部1180は、例えば、レンズなどの光学系、絞り、シャッタ、撮像素子などを含む。撮像画像入力部1181は、撮像部1180で得られた撮像画像の画像情報の入力を受け付けるものである。外部画像入力部1182は、映像信号入力部1131を通して外部画像の画像情報の入力を受け付けるものである。画像入力調整部1183は、撮像画像と外部画像の何れの画像情報を入力させるかを調整するものである。画像音声メモリ部1184は、画像入力調整部1183から入力された画像情報と、音声信号入力部1133から入力された音声情報とを、一時的に記憶するものである。
 画像エンコード部1185は、画像音声メモリ部1184に記憶されている画像情報を所定の形式によるデータに変換したり圧縮したりするものである。画像エンコード部1186は、画像音声メモリ部1184に記憶されている音声情報を所定の形式によるデータに変換したり圧縮したりするものである。データ生成部1187は、画像エンコード部1185からのデータ、あるいは、画像エンコード部1185からのデータおよび画像エンコード部1186からのデータに基づいて、画像データあるいは動画データを生成するものである。データ記録部1188は、データ生成部1187により生成された画像データあるいは動画データを記録媒体1189に記録するものである。記録媒体1189は、例えば、HHD、SSD、USBROM、DVD、CDなどである。現在位置取得部1195は、撮像装置1200の現在の位置を表す位置情報を取得するものであり、例えば、GPSシステムである。
 人工知能応答出力装置10010は、制御部1110、表示部10011、外部電源入力インタフェース1111、電源1106、二次電池1112、ストレージ部1170、映像制御部1160、操作入力部1107、姿勢センサ1113、メモリ1109、ローカルLLM処理部10028、不揮発性メモリ1108、通信部1132、音声出力部1140、およびマイク1139を有している。人工知能応答出力装置10010の構成および機能は、実施例1~9に係る人工知能応答出力装置とほぼ同様である。ただし、制御部1110は、撮像機能部1201を構成する撮像画像入力部1181、画像入力調整部1183等を含む各部と通信可能に接続されている。制御部1110は、撮像部1180によって得られた撮像画像の画像情報に対して種々の処理を実行することができる。
 ストレージ部1170あるいは不揮発性メモリ1108には、AIアシスタント機能を実現するためのプログラムが格納されている。制御部1110は、このプログラムを読み出し、メモリ1109に展開して実行することにより、AIアシスタント機能および当該機能のオンオフ制御を実現させる。
 制御部1110は、ユーザによる撮像部1180を用いた撮像を支援するための撮像支援情報を要求する指示情報を生成AIに送信する処理を実行する。また、制御部1110は、生成AIから撮像支援情報を含む応答情報を受信する処理を実行する。さらに、制御部1110は、受信した応答情報に含まれる撮像支援情報または当該撮像支援情報に基づく情報を表示部10011に表示させる処理を実行する。制御部1110のこのような処理により、ユーザによる撮像部1180を用いた撮像が支援される。
 なお、ストレージ部1170、不揮発性メモリ1108、メモリ1109、記録媒体1189は、それぞれ、本願における「記憶部」の一例である。また、現在位置取得部1195は、本願における「位置情報取得部」の一例である。
 図2Cから図2Fは、実施例2に係る撮像装置および生成AIサーバにおける処理フローの一例を示す図である。なお、図2Cから図2Fに示す処理フローは、撮像装置1200における制御部1110が実行する主な処理と、マルチモーダルなLLMサーバ20001を含む生成AIサーバが実行する主な処理とを、時系列的に記載したものである。また、図2Cから図2Fに示す処理フローは、撮像装置1200におけるユーザの操作に応じた処理と並列に実行されることを想定した場合の処理フローである。撮像装置1200において、バックグラウンドで撮像部1180による撮像が略一定の周期で行われ、直近の一定期間に得られた時系列的な複数のスルー画像が、記録媒体1189、ストレージ部1170、不揮発性メモリ1108、あるいはメモリ1109に格納される場合を想定する。
 ステップS1001では、AIアシスタント機能を起動する操作が行われたか否かを判定する処理が行われる。具体的には、制御部1110は、AIアシスタント機能を起動する操作が行われたか否かを判定する。AIアシスタント機能を起動する操作は、例えば、操作入力部1107を介した物理的なあるいはタッチパネルに表示されたボタンを押す操作、マイク1139を介した所定のコマンドを音声で入力する操作、あるいは手や顔の表情によるジェスチャーで入力する操作などである。この判定において、AIアシスタント機能を起動する操作が行われたと判定された場合には(S1001:Yes)、処理ステップはステップS1004に移行する。一方、AIアシスタント機能を起動する操作が行われていないと判定された場合には(S1001:No)、処理ステップはステップS1002に移行する。
 ステップS1002では、AIアシスタント機能の起動要否の問合せを開始する条件が満たされたか否かを判定する処理が行われる。具体的には、制御部1110は、AIアシスタント機能の起動要否の問合せを開始する条件が満たされたか否かを判定する。当該条件は、例えば、ユーザが構図の決定に一定以上の時間を費やしていることを検知すること、とすることができる。当該条件のより具体的な例としては、例えば、撮像装置1200が撮像モードに設定されてから一定期間の間、シャッターボタンが押されていないこと、という条件であってよい。
 この判定において、AIアシスタント機能の起動要否の問合せを開始する条件が満たされたと判定された場合には(S1002:Yes)、処理ステップはステップS1003に移行する。一方、AIアシスタント機能の起動要否の問合せを開始する条件が満たされていないと判定された場合には(S1002:No)、処理ステップはステップS1001に戻る。
 ステップS1003では、ユーザにAIアシスタント機能を起動するか否かを問い合せる処理が行われる。具体的には、制御部1110が、AIアシスタント機能を起動するか否かを問い合せる旨の自然言語によるテキスト情報を表示部10011に表示させる。なお、この処理は、撮像支援の要否をユーザに問い合せる処理であるともいえる。
 図2Gは、実施例2に係る撮像装置における表示画面の一例を示す図である。図2Gでは、AIアシスタント機能による生成AIとユーザとの対話的な表現を表すテキスト情報の表示例が示されている。図2Gの例では、表示画面1057内において、右側に生成AIの言葉10571が表示され、ユーザの言葉10572が左側に表示されている。また、図2Gの例では、それぞれの言葉が、上から下へ時系列的に並んで表示されている。
 ステップS1003において、例えば、図2Gに示すように、「AIアシスタントを起動しますか?」といったテキスト情報を表示させる。また、制御部1110は、上記の表示と併せて、あるいは、上記の表示に代えて、「AIアシスタントを起動しますか?」といった音声を音声出力部1140に出力させてもよい。このように、AIアシスタント機能を起動するかユーザに問い合わせることで、ユーザは、AIアシスタント機能の存在を知らない、忘れている、あるいは、起動のさせ方が分からない、起動の操作を面倒に思う、といった状況において、AIアシスタント機能を起動させることができる。
 制御部1110は、AIアシスタント機能の起動の要否、すなわち撮像支援の要否の問合せに対するユーザの回答を受け付ける。上記の問合せに対して、ユーザが「はい」を入力または選択するなどにより、制御部1110が、AIアシスタント機能の起動を要するする旨の回答を受け付けた場合には(S1003:Yes)、処理ステップはステップS1004に移行する。一方、上記の問合せに対して、ユーザが「いいえ」を入力または選択するなどにより、制御部1110が、AIアシスタントの起動を要しない旨の回答を受け付けた場合には(S1003:No)、AIアシスタント機能に係る処理が終了するか、または、問合せを開始する条件に係るパラメータがリセットされ、処理ステップがステップS1001に戻る。なお、ユーザは、タッチパネルとして機能する表示部10011の画面上で「はい」/「いいえ」と表示された箇所を指で押したり、「はい」/「いいえ」に対応するボタンを押したり、「はい」/「いいえ」と音声を発したりすることにより、回答してもよいし、手や顔の表情によるジェスチャーで入力してもよい。
 ステップS1004では、現在位置情報を取得する処理が行われる。具体的には、制御部1110は、現在位置取得部1195から、撮像装置1200の現在の位置を表す現在位置情報を受信する。現在位置情報は、例えば、地球上の座標を表す情報であり、GPSにより取得される。
 なお、ステップ1004の処理は、次のような処理であってもよい。例えば、制御部1110は、撮像部1180から、現在位置を特定し得るランドマークが含まれるスルー画像を受信する。制御部1110は、この現在位置におけるスルー画像情報を含み、現在位置を教えるように指示する指示情報を生成AIサーバに送信する。生成AIサーバは、この指示情報を受信すると、ランドマークが含まれるスルー画像を基に現在位置情報を含む応答情報を生成し、撮像装置1200に送信する。制御部1110は、この応答情報を受信し、受信した応答情報を基に現在位置情報を取得する。また、制御部1110は、ユーザによって入力されたランドマークの名称を表す情報を、指示情報に含めるようにし、現在位置あるいは後述する撮像ポイントの決定確度を高めるようにしてもよい。
 ちなみに、生成AIが大規模言語モデル(LLM)を含む場合には、生成AIは、自然言語、口語表現等を含むテキスト情報を扱うことができる。生成AIが、単なるLLMではなくマルチモーダルLLMを含む場合には、生成AIは、自然言語、口語表現等を含むテキスト情報に加えて、画像情報、音声情報などテキスト情報以外の情報も扱うことができる。生成AIサーバに送信する指示情報は、指示文、プロンプトなどとも呼ばれている。生成AIサーバは、受信した指示情報が表す指示に応答する形態で応答情報を生成し返信する。ユーザは、撮像に係るアシストを自然言語で受けることができ、アシストの内容を容易に理解し、撮像される画像を容易に想像することができる。
 ステップS1005では、撮像ポイントを教えるよう指示する処理が行われる。具体的には、制御部1110は、現在位置情報を含み、現在位置情報が表す現在位置に対応したランドマークが撮像可能な撮像ポイントを教えるよう指示する指示情報を生成し、生成AIサーバに送信する。すなわち、この指示情報は、現在位置に対応したランドマークが撮像可能な撮像ポイントを表す撮像ポイント情報を、撮像支援情報として要求する情報であるともいえる。なお、撮像ポイントは、本願における「撮像位置」の一例であり、撮像ポイント情報は、本願における「撮像位置情報」の一例である。
 撮像ポイントは、例えば、現在位置の付近で著名なランドマークを撮像するのに適していると考えられる撮像者の立ち位置もしくは座り位置であってもよい。撮像ポイントは、例えば、ランドマークの全体が撮像領域に収まるようなユーザの立ち位置またはしゃがみ位置であってもよい。撮像ポイントは、例えば、ランドマークが撮像領域内の特定の位置または領域に配置されるような構図が得られるユーザの立ち位置またはしゃがみ位置であってもよい。これらの場合において、指示情報には、撮像装置に取り付けられたレンズの焦点距離、ズームレンズの焦点距離の可変範囲などが含まれてもよい。生成AIは、これらレンズの焦点距離に関する情報に基づいて、撮像ポイントを決定してよい。
 ステップS1006では、撮像ポイントを含む応答情報を生成する処理が行われる。具体的には、生成AIサーバは、この指示情報を受信すると、現在位置に対応した撮像ポイントを含む応答情報を生成し、撮像装置に送信する。
 ステップS1007では、現在位置が撮像ポイントであるか否かを判定する処理が行われる。具体的には、制御部1110は、現在位置情報が表す現在位置と、現在位置に対応した撮像ポイントとを比較し、これらの位置が実質的に重なっているか否かにより、現在位置が撮像ポイントであるか否かを判定する。
 なお、ステップS1004~S1007の処理に代えて、次のような処理が行われてもよい。例えば、制御部1110は、現在位置情報または現在位置におけるスルー画像情報を含み、現在位置が当該現在位置に対応する撮像ポイントであるか否かを判定するように指示する指示情報を、生成AIサーバに送信する。生成AIサーバは、この指示情報を受信すると、現在位置が撮像ポイントである否かを判定し、その判定結果を表す判定結果情報を含む応答情報を生成し、撮像装置1200に送信する。制御部1110は、この応答情報を受信し、受信した応答情報を基に、現在位置が撮像ポイントであるか否かを判定する。
 この判定において、現在位置が撮像ポイントであると判定された場合には(S1007:Yes)、処理ステップはステップS1011に移行する。一方、現在位置が撮像ポイントでないと判定された場合には(S1007:No)、処理ステップはステップS1008に移行する。
 ステップS1008では、撮像ポイントまでユーザを誘導するよう指示する処理が行われる。具体的には、制御部1110は、現在位置情報が表す現在位置に対応した撮像ポイントまでユーザを誘導するよう指示する指示情報を生成し、生成した指示情報を生成AIサーバに送信する。なお、この指示情報は、撮像ポイントへとユーザを誘導するための誘導情報を、撮像支援情報として要求する情報であるともいえる。
 ステップS1009では、ユーザを撮像ポイントまで誘導するための誘導情報を含む応答情報を生成する処理が行われる。具体的には、生成AIサーバは、指示情報を受信すると、ユーザを撮像ポイントまで誘導するための誘導情報を含む応答情報を生成し、撮像装置1200に送信する。
 図2Gに示すように、誘導情報は、例えば、「ここはケルン大聖堂ですね?あと1m右へ、あと3m後ろならば全体を写せます。」といった自然言語によるテキスト情報である。生成AIサーバは、生成した応答情報を撮像装置1200に送信する。
 ステップS1010では、ユーザを誘導するための誘導情報を出力する処理が行われる。具体的には、制御部1110は、生成AIサーバから、ユーザを撮像ポイントまで誘導するための情報を含む応答情報を受信する。また、制御部1110は、受信した応答情報を基に、ユーザを誘導するための誘導情報を、表示部10011に表示させる。制御部1110は、誘導情報の表示と併せて、あるいは当該表示に代えて、誘導情報の内容を音声出力部1140に音声で出力させてもよい。なお、誘導情報は、生成AIサーバから受信した情報をそのまま出力してもよいし、受信した情報に基づいて生成された情報を出力するようにしてもよい。このように、誘導情報が出力されると、ユーザは、撮像ポイントを探索することなく、好適な撮像ポイントに短時間で移動することができる。その後、処理ステップはステップS1004に戻る。
 なお、ステップS1008~S1010の処理に代えて、次のような処理が行われてもよい。制御部1110は、撮像ポイントと現在位置との差分を算出し、その差分に基づいて、ユーザを撮像ポイントに誘導するための誘導情報を生成してもよい。
 ステップS1011では、姿勢情報およびスルー画像情報の少なくとも一方を取得する処理が行われる。具体的には、制御部1110は、姿勢センサ1113から、取得された撮像装置1200の姿勢情報を受信する。あるいは、制御部1110は、撮像部1180から、取得されたスルー画像情報を受信する。制御部1110は、姿勢情報とスルー画像情報の両方を受信してもよい。なお、スルー画像は、上述の通り、撮像装置1200において、撮像モードが設定されており、シャッターボタンが押されていない状況において、バックグラウンドで撮像部1180により撮像される画像である。スルー画像は、例えば、所定の周期で取得される。
 ステップS1012では、撮像の仕方をユーザに助言(アドバイス)するよう指示する処理が行われる。具体的には、制御部1110は、姿勢情報およびスルー画像情報の少なくとも一方の情報を含み、撮像の仕方をユーザに助言するよう指示する指示情報を生成し、生成した指示情報を生成AIサーバに送信する。なお、この指示情報は、撮像の仕方をユーザに助言するための助言情報を、撮像支援情報として要求する情報であるともいえる。
 指示情報には、スルー画像におけるランドマークを指し示す情報を含めるようにしてもよい。例えば、制御部1110は、ユーザの入力操作に応じてスルー画像の画像領域における任意の位置もしくは領域が指定され、指定された位置もしくは領域に対応した物体をランドマークとして決定する。なお、指示情報は、例えば、模範となる手本画像を送るよう指示する内容を含めてもよい。
 ステップS1013では、好適な撮像の仕方をユーザに助言するための助言情報を含む応答情報を生成する処理が行われる。具体的には、生成AIサーバは、指示情報を受信すると、姿勢情報およびスルー画像情報の少なくとも一方に基づいて、好適な撮像の仕方をユーザに助言するための助言情報を含む応答情報を生成する。なお、助言情報は、例えば、好適な構図が得られるような撮像装置1200の位置、向きなどをユーザに教示するための情報を含んでもよい。
 図2Gに示すように、助言情報は、例えば、「腰から下の位置でカメラを構えましょう。」、「人物が斜めに並べば奥行き感が出せます。」、「シャッターボタンをゆっくり押し込むとブレません。」といった自然言語によるテキスト情報であってもよい。また例えば、助言情報は、「左半分と右半分のうち一方にランドマークが配置され、他方に人物が配置されると、バランスの良い構図となります。」といった自然言語を含むテキスト情報であってもよい。また例えば、助言情報は、スルー画像において、ランドマークが配置されるべき領域、撮像対象人物が配置されるべき領域、あるいは、これら両方の領域などを示す画像・記号・マーク等が重畳された画像としてもよい。また、助言情報は、模範的な画像である手本画像を含んでもよい。
 なお、助言の内容は、ユーザの撮像スキルの高さに応じて変えるようにしてもよい。例えば、撮像装置1200において、ユーザの撮像スキルのレベルを、複数段階の中から予め設定しておく。この複数段階は、例えば、上級者、中級者、および初級者の3段階である。また例えば、撮像装置1200は、当該撮像装置1200の過去の扱われ方を、ストレージ部1170あるいはデータ記録部1188に記録しておき、その扱われ方を表す情報を生成AIに送信して、生成AIにユーザの撮像スキルのレベルを判定させてもよい。そして、制御部1110は、ユーザの撮像スキルのレベルに応じた助言をするよう生成AIに指示する。この場合、ユーザは、自身の撮像スキルのレベルに合った助言をもらうことができる。
 ユーザのレベルに応じた助言の具体例としては、制御部1110は、生成AIサーバとの協働により、撮像装置1200の好適な位置、撮像装置1200の好適な向き、好適な構図、好適な撮像条件、手ブレ防止(抑制)などを実現させるためのヒントをユーザに提示する。ユーザが初級者である場合には、制御部1110は、例えば、絞り、シャッタスピード、フォーカス、およびズームレンズなどに係る設定のヒントなどを提示する。ユーザが中級者である場合には、制御部1110は、例えば、ISO感度、ホワイトバランス(WB)、マクロレンズ、被写体の速い動きへの対応(被写体ブレ抑制)に係る設定のヒントなどを提示する。
 また、撮像装置1200のセッティング、具体的には、絞り、シャッタスピード、ISO感度、手振れ補正のオンオフ、ズームレンズの倍率、WBなどについては、制御部1110が、生成AIサーバから受信した応答情報に基づいて、自動的に設定するようにしてもよい。
 ステップS1014では、撮像の仕方をユーザに助言するための助言情報を出力する処理が行われる。具体的には、制御部1110は、生成AIサーバから、好適な撮像の仕方をユーザに助言するための助言情報を含む応答情報を受信する。また、制御部1110は、応答情報を基に、好適な撮像の仕方をユーザに助言するための助言情報を、表示部10011に表示させる。制御部1110は、当該表示と併せて、あるいは当該表示に代えて、上記の好適な撮像の仕方をユーザに助言するための助言情報を音声出力部1140に音声で出力させてもよい。制御部1110は、生成AIサーバから受信した情報をそのまま出力してもよいし、受信した情報に基づいて生成(編集)された情報を出力するようにしてもよい。
 なお、好適な撮像の仕方をユーザに助言するための助言情報を、個別具体的な内容にする場合には、上述の通り、姿勢情報あるいはスルー画像情報を用いる必要がある。一方、助言情報を、一般的な内容に留める場合には、姿勢情報あるいはスルー画像情報を用いなくてもよい。またこの場合は、撮像装置1200のストレージ部1170、不揮発性メモリ1108、あるいは記録媒体1189に、一般的な助言情報を予め格納しておき、制御部1110が、格納されている助言情報を読み出して出力させてもよい。すなわち、制御部1110は、生成AIサーバに指示情報を送信しなくてもよい。この場合、生成AIサーバに対する情報の送受信、生成AIサーバにおける処理を低減することができ、助言情報の出力に要する時間を短縮したり、生成AIサーバに対する負荷を軽減したりすることができる。
 図2Hは、実施例2に係る生成AIからの助言を受ける前後における構図の変化の一例を示す図である。図2Hにおいて、左側の構図は、生成AIからの助言をユーザが受ける前の助言前構図251であり、右側の構図は、生成AIからの助言をユーザが受けた後の助言後構図252である。助言前構図251では、図2Hに示すように、撮像領域(画角範囲)1050に対して、第1のランドマーク10501、第2のランドマーク10502、および二人の撮像対象人物10505のそれぞれの一部が外へはみ出している。また、助言前構図251の撮像領域1050内には、第1の自動車1504、二人の自転車に乗った人物10506、第2の自動車10508が、左から右に向けて移動する移動体として含まれている。さらに、助言前構図251内には、撮像対象人物10505の他に二人の他人10507が含まれている。
 一方、助言後構図252では、図2Hに示すように、撮像領域1050に対して、第1のランドマーク10501、および第2のランドマーク10502が外へはみ出さないように、ユーザが立つ撮像位置またはズームレンズの倍率が調整されている。また、撮像対象人物10505についても、撮像領域1050内に収まるように、ユーザによって撮像対象人物10505が誘導されたり、あるいは、ズームレンズの倍率が調整されたりしている。また、二人の撮像対象人物10505は、撮像画像において被写体の立体感が出るよう、斜めに配置されるように、ユーザによって誘導されている。なお、第1の自動車10504、二人の自転車に乗った人物10506、第2の自動車10508などの移動体、他人10507については、特に、誘導されない。
 ステップS1015では、仮加工済画像を得る操作が行われたか否かを判定する処理が行われる。具体的には、制御部1110は、仮加工済画像を得る操作が行われたか否かを判定する。仮加工済画像を得る操作は、例えば、シャッターボタンを半押しする操作であってもよい。一般的に、シャッターボタンが2段式のスイッチになっている場合において、シャッターボタンを1段目まで押す操作のことを「半押し」、2段目までさらに押し込む操作のことを「全押し」という。この判定において、仮加工済画像を得る操作が行われたと判定された場合には(S1015:Yes)、処理ステップはステップS1016に移行する。一方、仮加工済画像を得る操作が行われていないと判定された場合には(S1015:No)、処理ステップはステップS1004に戻る。
 ステップS1016では、不要物体を抽出するよう指示する処理が行われる。具体的には、制御部1110は、スルー画像において、被写体として好適ではないと考えられる不要物体を抽出するよう指示する指示情報を生成し、生成した指示情報を生成AIサーバに送信する。なお、この指示情報は、撮像画像であるスルー画像において、被写体として好適ではないと考えられる不要物体を抽出した結果を表す抽出結果情報を、撮像支援情報として要求する情報であるともいえる。この指示情報は、例えば、スルー画像において、ランドマークおよび撮像対象人物を除いた物体のうち、移動体および撮像非対象人物のうち少なくとも一方を、不要物体として抽出するよう指示する情報であってもよい。
 図2Iは、実施例2に係る撮像装置により得られるスルー画像において不要物体が抽出される様子の一例を示す図である。図2Iにおいて、上段には、バックグラウンドで実行される処理により得られた時系列的な複数のスルー画像1053A,1503B,1053C,1053Dが示されている。
 移動体は、例えば、図2Iに示すように、時系列的な複数のスルー画像1053A~1053Dにおいて、ランドマークあるいは他の静止している被写体を基準に、相対的に位置が変化している物体と規定することができる。あるいは、移動体は、例えば、手振れ補正が働いている状況の下、スルー画像の画像領域において水平方向および鉛直方向に延びる格子状の複数の基準線を設定し、時系列的な複数のスルー画像において、これら基準線を跨ぐように移動する物体と規定することができる。
 また、撮像非対象人物は、例えば、図2Iに示すように、スルー画像1053A~1053Dの何れかにおいて、スルー画像に含まれる人物のうち、顔の向きまたは視線の向きが、撮像装置1200に向かう向きと一定レベルを超えてずれている人物と規定することができる。あるいは、撮像非対象人物は、例えば、スルー画像1053A~1503Dの何れかの画像領域に対して占める割合が一定レベル未満である人物と規定することができる。
 図2Iに示すように、スルー画像1054の撮像領域1050内において、黒塗りの矢印が付された被写体が移動体または撮像非対象人物として抽出されている。具体的には、第1のランドマーク10501あるいは第2のランドマーク10502を基準に位置が変化している移動体として、第1の自動車10504、第2の自動車10508、および、二台の自転車とそれらの運転者である人物10506が抽出されている。また、撮像領域1050の面積に対して占める面積が一定割合以下である人物、あるいは、顔または視線の向きが撮像装置を向く向きから一定レベル以上ずれている人物である撮像非対象人物として、小さく写り込んだ横向きの二人の他人10507が抽出されている。
 移動体は、通行人、乗り物などである可能性が高い。また、撮像非対象人物は、ユーザから見て他人である可能性が高い。よって、移動体あるいは撮像非対象人物は、ユーザが本来被写体として撮像したい物体あるいは人物でない可能性が高い。ユーザは、そのような被写体として望まない物体・人物を、画像の加工処理によって撮像画像内から除去したい場合に、これらの物体・人物をある程度自動で選択してもらうことができ、煩雑な操作を低減することができる。
 ステップS1017では、不要物体が抽出された結果を表す抽出結果情報を含む応答情報を生成する処理が行われる。具体的には、生成AIサーバは、指示情報を受信すると、抽出結果情報を含む応答情報を生成し、生成した応答情報を撮像装置1200に送信する。抽出結果情報は、例えば、スルー画像に、抽出された不要物体を指し示すマーク等が付加された抽出結果画像である。また例えば、抽出結果情報は、上記抽出結果画像に加えて、「カメラ画像内の移動体と他人を自動的に抽出しました。」といった自然言語を含むテキスト情報が含まれる情報である。
 ステップS1018では、抽出結果情報を出力する処理が行われる。具体的には、制御部は、生成AIサーバから、抽出結果情報を含む応答情報を、インタフェースを介して受信し、抽出結果情報を表示部10011に表示させる。なお、制御部1110は、抽出結果情報の表示と併せて、抽出結果に係る情報を音声出力部1140に音声で出力させてもよい。また、抽出結果情報は、生成AIサーバから受信した情報をそのまま出力してもよいし、受信した情報に基づいて生成(編集)された情報を出力するようにしてもよい。また、抽出結果情報は、例えば、「カメラ画像内の移動体と他人を自動的に抽出しました。」といった自然言語によるテキスト情報を含んでもよい。
 ステップS1019では、加工対象を選択する処理が行われる。具体的には、制御部1110は、初期設定として、抽出されたすべての不要物体を加工対象候補として設定する。ユーザは、加工対象候補の中で加工したくない物体を加工対象から外す操作を行ったり、加工対象候補以外の物体を加工対象候補として指定する操作を行ったりする。制御部1110は、ユーザのこれらの操作に応じて、スルー画像内における加工対象を選択する。なお、制御部1110は、抽出された不要物体をそのまま加工対象として選択してもよいし、不要物体の抽出が行われない状況において、ユーザの操作に応じてスルー画像内におけるユーザが所望する物体を加工対象を選択してもよい。
 ステップS1020では、スルー画像において加工対象を加工するように指示する処理が行われる。具体的には、制御部1110は、スルー画像を含み、スルー画像において加工対象を加工するよう指示する指示情報を生成し、生成した指示情報を生成AIサーバに送信する。なお、この指示情報は、スルー画像において加工対象、例えば、移動体および撮像非対象人物の少なくとも一方を除去する加工が施された加工済画像を、撮像支援情報として要求する情報ともいえる。
 この指示情報は、例えば、スルー画像において、加工対象の物体を表す画像を取り除き、画像が取り除かれた領域を、その領域の周辺の画像データを用いて修復する加工を施すよう指示する情報であってもよい。このようにすることで、構図に含まれるリアルな背景を使いつつ、見映えの良い画像をAIを用いて生成することができる。
 また例えば、この指示情報は、スルー画像において、加工対象の物体を表す画像を取り除き、取り除かれた領域を、ユーザが指定した所望の画像データを用いて修復する加工を施すよう指示する情報であってもよい。このようにすることで、ユーザの好みが反映された見映えの良い画像を生成することができる。
 ステップS1021では、加工結果情報を含む応答情報を生成する処理が行われる。具体的には、生成AIサーバは、指示情報を受信すると、スルー画像における加工対象を加工して加工済画像を得、加工結果情報を含む応答情報を生成し、生成した応答情報を撮像装置に送信する。加工結果情報は、画像済画像を含んでいる。なお、加工結果情報は、例えば、加工済画像に加えて、「指定された加工対象を除外し周辺の画像データを用いて加工しました。」といった自然言語によるテキスト情報が含まれてもよい。
 ステップS1022では、加工結果情報を出力する処理が行われる。具体的には、制御部1110は、生成AIサーバから、加工結果情報を含む応答情報を受信し、加工結果情報に含まれる加工済画像を仮加工済画像として表示部10011に表示させる。また、制御部1110は、加工結果情報に含まれるテキスト情報を表示部に表示させたり、音声出力部1140に音声で出力させたりする。なお、制御部1110は、仮加工済画像と、加工前のスルー画像である元画像とを表示部10011の表示画面に並べて表示させてもよい。これにより、ユーザは、仮加工済画像と元画像と比較して見ることができ、自身の好みに合わせて加工済画像をデザインすることができる。
 図2Jは、実施例2に係る撮像装置において元画像と仮加工済画像とが表示された例を示す図である。例えば、図2Jに示すように、表示部10011の表示画面には、元画像1055と仮加工済画像とが並らんで表示される。元画像1055では、撮像領域1050内において、第1のランドマーク10501、第2のランドマーク10502、および撮像対象人物10505の他に、移動体である第1の自動車10504、第2の自動車10508、および二台の自転車とその運転者である人物10506、撮像非対象人物である他人10507が含まれている。一方、仮加工済画像1056では、撮像領域1050内において、第1のランドマーク10501、第2のランドマーク10502、および撮像対象人物10505のみが含まれており、移動体である第1の自動車10504、第2の自動車10508、および二台の自転車とその運転者である人物10506、撮像非対象人物である他人10507は消去されている。
 ステップS1023では、仮加工済画像を得る操作が継続して行われている状態であるか否かを判定する処理が行われる。具体的には、制御部1110は、仮加工済画像を得る操作が引き続き行われている状態であるか否かを判定する。この判定において、仮加工済画像を得る操作が継続して行われている状態であると判定された場合には(S1023:Yes)、処理ステップはステップS1023に戻る。一方、仮加工済画像を得る操作が継続して行われている状態ではないと判定された場合には(S1023:No)、処理ステップはステップS1024に移行する。
 ステップS1024では、画像を保存する操作が行われた否かを判定する処理が行われる。具体的には、制御部1110は、画像を保存する操作がユーザによって行われたか否かを判定する。画像を保存する操作は、例えば、シャッターボタンを全押しする操作である。この判定において、画像を保存する操作が行われたと判定された場合には(S1024:Yes)、処理ステップはステップS1025に移行する。一方、画像を保存する操作が行われていないと判定された場合には(S1024:No)、処理ステップはステップS1008に戻る。
 ステップS1025では、元画像と加工済画像とを対応付けて保存する処理が行われる。具体的には、制御部1110は、仮加工済画像としての加工済画像と、加工前のスルー画像である元画像とを、互いに対応付けて、ストレージ部1170あるいは記録媒体1189に保存する。例えば、加工済画像のファイル名には、撮像番号と加工済であることを表す文字・符号・数字・拡張子等とが含まれ、元画像のファイル名には、同じ撮像番号と元画像であることを表す文字・符号・数字・拡張子等が含まれる。
 また、制御部1110は、画像が保存された旨を表す画像保存報告情報を表示部に表示させる。制御部1110は、上記の画像保存報告情報の表示と併せて、あるいは、当該表示に代えて、画像保存報告情報を音声出力部1140に音声で出力させたり、定められた音を音声出力部1140に出力させたりしてもよい。画像保存報告情報は、例えば、「元の画像はデータ加工していないオリジナル画像データとデータ加工後の画像データは判別人により区別可能な形式で区別して記録しました。」といった自然言語によるテキスト情報であってもよい。画像が保存されると、処理ステップはステップS1026に移行する。
 このように、元画像と加工済画像とを対応付けて保存することにより、ユーザは、後から撮像した画像を確認する際に、元画像と加工済画像とを比較することができ、ユーザの都合あるいは好みに応じて、利用する画像を選択することができる。
 ステップS1026では、AIアシスタント機能を終了させるか否かを判定する処理が行われる。具体的には、制御部1110は、AIアシスタント機能を終了させる操作が行われたか否か、ユーザがAIアシスタント機能を使用しないで経過した時間が一定時間以上であるか否かなどに基づいて、AIアシスタント機能を終了させるか否かを判定する。この判定において、AIアシスタント機能を終了させると判定された場合には、制御部1110は、AIアシスタント機能を終了させる。一方、AIアシスタント機能を終了させないと判定された場合には、処理ステップはステップS1004に戻る。
 以上、本発明の実施例2に係る撮像装置によれば、ユーザは、撮像時に好適な撮像のためのヒントを得ることができ、自身の地理的な知識(土地勘)、撮像スキル等に依存することなく、好適な撮像を実施することができる。例えば、ユーザは、初めて来る場所において、個人の撮像スキルがそれほど高くなくても、好適な構図、好適な画角、撮像装置の好適な設定(セッティング)を実現させることができる。
 また、実施例2に係る撮像装置によれば、ユーザは、撮像スキルが高くても、自撮りを行う場合には、被写体に対する撮像装置の向き、撮像装置の構え方、撮像者の立つ場所などについて、通常の撮像とは感覚が異なるため、ユーザは思うように撮像できないことがある。このような場合であっても、実施例2に係る撮像装置によれば、ユーザは、AIアシスタント機能を利用することで、より好適な撮像を実施することが可能である。
 また、実施例2に係る撮像装置によれば、ユーザは、撮像に関するアドバイスを口語表現で得られるので、撮像装置とユーザとの親和性が高く、アドバイスの内容を直感的に理解することができる。
 また、実施例2に係る撮像装置によれば、生成AI、特にマルチモーダルLLMを用いることにより、画像情報とテキスト情報とを同時に扱うことができ、通信先の生成AIサーバを情報の種類によって切り替える必要がなく、制御部の処理を簡単にすることができる。
 また、実施例2に係る撮像装置によれば、指示情報の作成を工夫することで、ユーザは、より細かい要求に応えるアドバイスを受け取ることができる。例えば、指示情報を、「専門用語を使わずに撮像装置の好適な撮像条件を教えて」といった内容にすることで、ユーザが、仮に専門用語が分からない初心者であっても、理解できるアドバイスを得ることができる。
 また、実施例2に係る撮像装置によれば、元画像(オリジナル画像)によって撮像した際の臨場感、記憶(思い出)などを残すとともに、印象に残したい物体以外の余分な物体を撮影の際に排除(加工)して記録することにより、後で編集作業なく高品質な画像を作成することができる。
 なお、実施例2において、生成AIに、画角の外から内に入ってくるものを、目的の被写体であると認識させたり、不要な被写体であると認識させたりすることも可能である。前者の場合において、被写体を画像の左右方向の中央付近で撮りたい場合に、予めその設定をしておくことで、被写体が中央付近に差し掛かるタイミングを予測して、ユーザがシャッタを切るタイミングを教示してくれるアシストを、生成AIに要求することも考えられる。
 また、実施例2において、LLM対話ウィンドウと画像表示ウィンドウとを表示部に表示させる場合において、表示部が1画面のみ有している場合には、1画面に各ウィンドウを並列表示させてもよいし、ウィンドウを切り換えて表示させてもよい。表示部が2画面、例えば、撮像装置の上面と背面に1画面ずつ有している場合には、一方の画面にLLM対話ウィンドウを表示させ、他方の画面に画像表示ウィンドウを表示させてもよい。
 また、実施例2において、元画像と加工済画像、あるいは、スルー画像と手本画像を表示部に表示させる場合において、表示部が1画面のみ有している場合には、これら2画像を1画像ずつ切り換えて表示させてもよいし、1画面に2画像を並列に表示させてもよい。表示部が2画面を有している場合には、1画面に1画像ずつ表示させてもよい。
 また、実施例2において、好適な構図は、1つに特定してもよいし、複数の中からユーザが選択してもよい。好適な構図の例としては、ランドマークが画角内に収まる構図、被写体人物が画角内に収まる構図、ランドマークと被写体人物の両方が画角に収まる構図、立体感が出る構図などを考えることができる。また、好適な構図の考え方が一つであっても、生成AIが好適な配置などを複数提案してもよい。
 また、実施例2において、生成AIで加工(背景の一部分の変換)された画像に対し、後からユーザが所望する画像に再変換することができる機能が実現されてもよい。画角に外から内に入ってくる移動体を不要物体として特定し、撮像画像においてその不要物体を加工で取り除いてもよい。
 また、実施例2は、生成AIによるアシスト機能を、静止画の撮像に適用した例であるが、連写による撮像あるいは動画の撮像についても同様に適用することができる。ここで、連写による撮像とは、シャッタボタンを押している間、静止画を連続的に(例えば、3~30コマ/秒程度で)撮像することである。また、動画の撮像とは、静止画を所定のフレームレート(例えば、30~60fps)で撮像することである。
 また、実施例2において、生成AIは、撮像装置とネットワークを介して接続されたサーバに構築されたものであってもよいし、当該サーバに構築されたものと撮像装置に内蔵されたものとを含むものであってもよい。すなわち、状況あるいは撮像装置もしくはユーザの状態など種々の条件によって、撮像装置における制御部は、サーバに構築された生成AIと送受信を行ってもよいし、撮像装置に内蔵されたローカルな生成AIと送受信を行ってもよい。
 また、実施例2において、撮像装置は、例えば、デジタルカメラ、デジタルムービーカメラ、撮像機能付きのパーソナルコンピュータ、スマートフォン、タブレット端末などである。撮像装置は、例えば、互いに通信可能に接続されたデジタルカメラとスマートフォン等との組合せであってもよい。この場合、例えば、デジタルカメラに含まれる撮像部を、実施例2における撮像部として用い、デジタルカメラに含まれる制御部およびスマートフォン等に含まれる制御部を、実施例2における制御部として用いる。撮像装置における生成AIサーバとの通信は、スマートフォン等に含まれる通信部を介して行われるようにしてもよい。
 なお、実施例2に係る撮像装置および生成AIにおいて実行される処理の流れに基づく撮像支援方法は、本発明の実施例の一つである。また、当該撮像支援方法が実施されるよう、実施例2に係る撮像装置における制御部を構成する1以上のプロセッサ(CPU、MPUなど)に、各処理を実行させるためのプログラムも、本発明の実施例の一つである。さらに、当該プログラムが記録された有体の非一時的な記録媒体もまた、本発明の実施例の一つである。当該プログラムは、サーバからダウンロードして用いられてもよい。
 本実施例に係る技術では、より好適な人工知能撮像アシスタント技術を提供することが可能となる。このような人工知能撮像アシスタント技術は、実施例1と同様に、より質の高い、より信頼できるインフラへの導入が期待できる。また、当該技術がインフラへ導入されていくことにより、全ての人々に安価で公平なアクセスに重点を置いた経済発展と人間の福祉の支援に寄与できる。これにより、国連の提唱する持続可能な開発目標(SDGs)の「9産業と技術革新の基盤をつくろう」に貢献する。
 また、本実施例に係る技術では、より好適な人工知能撮像アシスタント技術を提供することが可能となる。このような人工知能撮像アシスタント技術は、実施例1と同様に、当該技術が公共交通機関へ導入されていくことにより、公共交通機関の拡大などを通じた交通の安全性改善、および、全ての人々に安全かつ安価で容易に利用できる、持続可能な輸送システムへのアクセスの実現に寄与できる。これにより、国連の提唱する持続可能な開発目標(SDGs)の「11住み続けられるまちづくりを」に貢献する。
 以上、種々の実施例について詳述したが、しかしながら、本発明は、上述した実施例のみに限定されるものではなく、様々な変形例が含まれる。例えば、上記した実施例は本発明を分かりやすく説明するためにシステム全体を詳細に説明したものであり、必ずしも説明したすべての構成を備えるものに限定されるものではない。また、ある実施例の構成の一部を他の実施例の構成に置き換えることが可能であり、また、ある実施例の構成に他の実施例の構成を加えることも可能である。また、各実施例の構成の一部について、他の構成の追加・削除・置換をすることが可能である。
 本明細書中に記載されている構成要素により実現される機能は、当該記載された機能を実現するようにプログラムされた、汎用プロセッサ、特定用途プロセッサ、集積回路、ASICs (Application Specific Integrated Circuits)、CPU (a Central Processing Unit)、従来型の回路、および/又はそれらの組合せを含む、circuitry又はprocessing circuitryにおいて実装されてもよい。プロセッサは、トランジスタやその他の回路を含み、 circuitry又はprocessing circuitryとみなされる。プロセッサは、メモリに格納されたプログラムを実行する、programmed processorであってもよい。
 本明細書において、circuitry、ユニット、手段は、記載された機能を実現するようにプログラムされたハードウェア、又は実行するハードウェアである。当該ハードウェアは、本明細書に開示されているあらゆるハードウェア、又は、当該記載された機能を実現するようにプログラムされた、又は、実行するものとして知られているあらゆるハードウェアであってもよい。
 当該ハードウェアがcircuitryのタイプであるとみなされるプロセッサである場合、当該circuitry、手段、又はユニットは、ハードウェアと、当該ハードウェア及び又はプロセッサを構成する為に用いられるソフトウェアの組合せである。
10010…人工知能応答出力装置、10011…表示部、10028…ローカルLLM処理部、1107…操作入力部、1110…制御部、1132…通信部、1140…音声出力部、1139…マイク、1160…映像制御部、1170…ストレージ部、1180…撮像部、1108…不揮発性メモリ、1109…メモリ、1113…姿勢センサ、1189…記録媒体、1195…現在位置取得部、1200…撮像装置、1201…撮像機能部、19000…インターネット、19001…大規模言語モデルサーバ、20001…マルチモーダルな大規模言語モデルサーバ

Claims (15)

  1.  撮像部と、表示部と、制御部と、を備え、
     前記制御部は、
     ユーザによる前記撮像部を用いた撮像を支援するための撮像支援情報を要求する指示情報を生成人工知能に送信する処理と、
     前記生成人工知能から前記撮像支援情報を含む応答情報を受信する処理と、
     前記受信した応答情報に含まれる撮像支援情報または当該撮像支援情報に基づく情報を前記表示部に表示させる処理と、を実行する、
     撮像装置。
  2.  請求項1に記載の撮像装置において、
     当該撮像装置の位置情報を取得する位置情報取得部を備え、
     前記指示情報は、前記位置情報取得部により取得された位置情報が表す位置に対応したランドマークが撮像可能な撮像位置を表す撮像位置情報を、前記撮像支援情報として要求する情報である、
     撮像装置。
  3.  請求項1に記載の撮像装置において、
     当該撮像装置の位置情報を取得する位置情報取得部を備え、
     前記指示情報は、前記位置情報取得部により取得された位置情報が表す位置に対応したランドマークが撮像可能な撮像位置へと前記ユーザを誘導するための誘導情報を、前記撮像支援情報として要求する情報である、
     撮像装置。
  4.  請求項3に記載の撮像装置において、
     前記撮像位置は、前記ランドマークの全体が撮像領域に収まるような前記ユーザの立ち位置またはしゃがみ位置である、
     撮像装置。
  5.  請求項1に記載の撮像装置において、
     前記指示情報は、撮像の仕方を前記ユーザに助言するための助言情報を、前記撮像支援情報として要求する情報である、
     撮像装置。
  6.  請求項1に記載の撮像装置において、
     前記指示情報は、前記撮像部により得られた撮像画像において、移動体および撮像非対象人物の少なくとも一方を抽出した結果を表す抽出結果情報を、前記撮像支援情報として要求する情報である、
     撮像装置。
  7.  請求項1に記載の撮像装置において、
     前記指示情報は、前記撮像部により得られた撮像画像において、移動体および撮像非対象人物の少なくとも一方を除去する加工が施された加工済画像を、前記撮像支援情報として要求する情報である、
     撮像装置。
  8.  請求項7に記載の撮像装置において、
     前記制御部は、
     前記得られた撮像画像における元画像と前記加工済画像とを前記表示部に表示させる処理と、
     ユーザによる操作に応答して、前記元画像と前記加工済画像とを互いに対応付けて記憶部に記憶させる処理と、を実行する、
     撮像装置。
  9.  請求項1に記載の撮像装置において、
     前記制御部は、
     前記ユーザが構図の決定に一定以上の時間を費やしていることを検知して、撮像支援の要否を前記ユーザに問い合せる処理と、
     前記撮像支援の要否の問合せに対する前記ユーザの回答を受け付ける処理と、
     前記撮像支援を要する旨の回答を受け付けた場合に、前記指示情報を前記生成人工知能に送信する処理と、を実行する、
     撮像装置。
  10.  請求項1に記載の撮像装置において、
     前記生成人工知能は、マルチモーダルな大規模言語モデルを含み、
     前記制御部は、前記受信した撮像支援情報または当該撮像支援情報に基づく情報として、自然言語によるテキスト情報が含まれる情報を、前記表示部に表示させる処理、を実行する、
     撮像装置。
  11.  請求項1に記載の撮像装置において、
     前記生成人工知能は、マルチモーダルな大規模言語モデルを含み、
     前記制御部は、前記受信した撮像支援情報または当該撮像支援情報に基づく情報として、対話的な表現を表すテキスト情報が含まれる情報を、前記表示部に表示させる処理、を実行する、
     撮像装置。
  12.  請求項1に記載の撮像装置において、
     前記生成人工知能は、当該撮像装置と通信ネットワークを介して接続されたサーバに構築されたものである、
     撮像装置。
  13.  請求項1に記載の撮像装置において、
     前記生成人工知能は、当該撮像装置と通信ネットワークを介して接続されたサーバに構築されたものと、当該撮像装置に内蔵されたものと、を含む、
     撮像装置。
  14.  撮像装置における制御部が、
     ユーザによる撮像部を用いた撮像を支援するための撮像支援情報を要求する指示情報を生成人工知能に送信する処理と、
     前記生成人工知能から前記撮像支援情報を含む応答情報を受信する処理と、
     前記受信した応答情報に含まれる撮像支援情報または当該撮像支援情報に基づく情報を表示部に表示させる処理と、を実行する、
     撮像支援方法。
  15.  撮像装置における制御部を構成するプロセッサに、
     ユーザによる撮像部を用いた撮像を支援するための撮像支援情報を要求する指示情報を生成人工知能に送信する処理と、
     前記生成人工知能から前記撮像支援情報を含む応答情報を受信する処理と、
     前記受信した応答情報に含まれる撮像支援情報または当該撮像支援情報に基づく情報を表示部に表示させる処理と、を実行させるための、
     プログラム。
PCT/JP2024/026516 2024-07-24 2024-07-24 撮像装置、撮像支援方法、およびプログラム Pending WO2026022985A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/JP2024/026516 WO2026022985A1 (ja) 2024-07-24 2024-07-24 撮像装置、撮像支援方法、およびプログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2024/026516 WO2026022985A1 (ja) 2024-07-24 2024-07-24 撮像装置、撮像支援方法、およびプログラム

Publications (1)

Publication Number Publication Date
WO2026022985A1 true WO2026022985A1 (ja) 2026-01-29

Family

ID=98565386

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2024/026516 Pending WO2026022985A1 (ja) 2024-07-24 2024-07-24 撮像装置、撮像支援方法、およびプログラム

Country Status (1)

Country Link
WO (1) WO2026022985A1 (ja)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019140561A (ja) * 2018-02-13 2019-08-22 オリンパス株式会社 撮像装置、情報端末、撮像装置の制御方法、および情報端末の制御方法
JP2022055656A (ja) * 2020-09-29 2022-04-08 オリンパス株式会社 アシスト装置およびアシスト方法
JP2023510430A (ja) * 2020-02-06 2023-03-13 三菱電機株式会社 シーンアウェア映像対話
JP2024035150A (ja) * 2022-08-30 2024-03-13 三菱電機株式会社 エンティティを制御するためのシステムおよび方法
WO2024149612A1 (en) * 2023-01-10 2024-07-18 Koninklijke Philips N.V. Distributed ultrasound including a wearable ultrasound device

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019140561A (ja) * 2018-02-13 2019-08-22 オリンパス株式会社 撮像装置、情報端末、撮像装置の制御方法、および情報端末の制御方法
JP2023510430A (ja) * 2020-02-06 2023-03-13 三菱電機株式会社 シーンアウェア映像対話
JP2022055656A (ja) * 2020-09-29 2022-04-08 オリンパス株式会社 アシスト装置およびアシスト方法
JP2024035150A (ja) * 2022-08-30 2024-03-13 三菱電機株式会社 エンティティを制御するためのシステムおよび方法
WO2024149612A1 (en) * 2023-01-10 2024-07-18 Koninklijke Philips N.V. Distributed ultrasound including a wearable ultrasound device

Similar Documents

Publication Publication Date Title
US8515728B2 (en) Language translation of visual and audio input
KR102668429B1 (ko) 전자장치 및 그 제어방법
US11429339B2 (en) Electronic apparatus and control method thereof
US20150237300A1 (en) On Demand Experience Sharing for Wearable Computing Devices
US11538278B2 (en) Electronic apparatus and control method thereof
JP2023549810A (ja) 動物顔スタイル画像の生成方法、モデルのトレーニング方法、装置及び機器
CN113052085A (zh) 视频剪辑方法、装置、电子设备以及存储介质
US12464217B2 (en) Method, electronic device, and storage medium for capturing a subject based on a posture template and capturing prompt information
JP2018528730A (ja) 動画提供装置、動画提供方法及びそのコンピュータプログラム
US20150177944A1 (en) Capturing objects in editable format using gestures
WO2024158893A1 (en) Systems and methods for capturing an image of a desired moment
CN117633703A (zh) 一种基于智能手表的多模态交互系统及方法
WO2025032913A1 (ja) 応答出力装置および応答出力システム
US20240127805A1 (en) Electronic apparatus and control method thereof
KR20180097040A (ko) 로봇 및 그 동작 방법
JP2025068754A (ja) 応答出力システムおよび応答出力装置
US11363084B1 (en) Methods and systems for facilitating conversion of content in public centers
CN119673014B (zh) 一种穿戴式技能教示系统
EP4498227A1 (en) Projection system, terminal device, projection device and control method thereof
US20260029987A1 (en) Electronic apparatus for controlling object included in screen based on user voice and controlling method
JP2025025082A (ja) 応答出力装置
WO2023182932A2 (zh) 目标物体的识别方法、装置、电子设备及存储介质
JP2025073085A (ja) システム
JP2026038158A (ja) システム
JP2026512389A (ja) ビデオ処理方法と装置、機器、記憶媒体及びプログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24948684

Country of ref document: EP

Kind code of ref document: A1