WO2018164932A1 - Zoom coding using simultaneous and synchronous multiple-camera captures - Google Patents
Zoom coding using simultaneous and synchronous multiple-camera captures Download PDFInfo
- Publication number
- WO2018164932A1 WO2018164932A1 PCT/US2018/020418 US2018020418W WO2018164932A1 WO 2018164932 A1 WO2018164932 A1 WO 2018164932A1 US 2018020418 W US2018020418 W US 2018020418W WO 2018164932 A1 WO2018164932 A1 WO 2018164932A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- camera
- video
- fov
- view
- field
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N23/00—Cameras or camera modules comprising electronic image sensors; Control thereof
- H04N23/45—Cameras or camera modules comprising electronic image sensors; Control thereof for generating image signals from two or more image sensors being of different type or operating in different modes, e.g. with a CMOS sensor for moving images in combination with a charge-coupled device [CCD] for still images
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N23/00—Cameras or camera modules comprising electronic image sensors; Control thereof
- H04N23/60—Control of cameras or camera modules
- H04N23/63—Control of cameras or camera modules by using electronic viewfinders
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N23/00—Cameras or camera modules comprising electronic image sensors; Control thereof
- H04N23/60—Control of cameras or camera modules
- H04N23/69—Control of means for changing angle of the field of view, e.g. optical zoom objectives or electronic zooming
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N23/00—Cameras or camera modules comprising electronic image sensors; Control thereof
- H04N23/60—Control of cameras or camera modules
- H04N23/62—Control of parameters via user interfaces
Definitions
- images are captured by lenses with different optical parameters, such as different visual zoom factors. Different camera optics affect the quality and appearance of the images captured by the cameras. These images may undergo further digital enhancement, such as digitally changing the zoom and field of view (FOV) parameters.
- FOV field of view
- multiple different video cameras can be used to capture the same event.
- the multiple different video cameras may have different optical capabilities, such as different optical zoom settings.
- the multiple video cameras may also be on the same device, such as a smart phone having two cameras oriented in the same direction.
- An exemplary method comprises capturing a first video of a scene with a first camera on a user device, where the first camera has a first field of view (e.g. a wide-angle field of view) with a first extent (e.g. a field of view between around 64° and 84° in the vertical and/or horizontal direction).
- a second camera e.g. a zoom camera
- the second camera has a second field of view having a second extent that is narrower than the first extent (e.g. less than around 64°).
- the second extent may be represented in terms of pixels within the first video.
- the second field of view is contained within the first field of view, such objects within the field of view of the second camera are also within the field of view of the first camera.
- the method further comprises tracking a position of an object of interest with respect to the first and second fields of view. This may be done using, for example, RFID or other wireless signals, or performing optical tracking using object recognition, among other possibilities.
- a determination is made of whether the object of interest is positioned entirely within the second field of view.
- a video output is provided. If the object of interest is positioned entirely within the second field of view, the video output comprises the second video.
- a cropped-and-scaled version of the first video is generated such that the video output includes the entirety of the object of interest and such that the extent of the output video corresponds to the second extent, and the cropped-and-scaled version of the first video is provided as the output.
- An exemplary method for providing an output video stream in a device having a first camera and a second camera, the second camera having a field of view that is narrower than and contained within a field of view of the first camera comprises tracking an object in video captured by the first camera. A determination, based on the tracking, is made on whether the object is entirely within the field of view of the second camera. If the object is within the field of view of the second camera, video captured by the second camera is outputted as the output video stream. If the object is outside the field of view of the second camera, video captured by the first camera is cropped and upscaled to generate a cropped-and-upscaled video of the object, and the generated video is outputted as the output video stream.
- the method is performed in real time.
- the method may also include, before tracking the object, displaying, on a screen of the device, video captured by the first camera and receiving user input selecting the object to be tracked in the displayed video.
- the output video stream may be displayed on the device.
- an indication of the position of the object may be displayed if the object is outside the field of view of the second camera.
- tracking the object may be done using RFID or other wireless protocol signals.
- FIG. 1 A is a functional block diagram of a client device that may be used in some embodiments.
- FIG. 1 B is a functional block diagram of a network entity that may be used in some embodiments.
- FIG. 2 depicts a client device with multiple video cameras, in accordance with an embodiment.
- FIG. 3A depicts a first system architecture for use of multiple video cameras, in accordance with an embodiment.
- FIG. 3B depicts a second system architecture for use of multiple video cameras, in accordance with an embodiment.
- FIG. 3C depicts a third system architecture for use of multiple video cameras, in accordance with an embodiment.
- FIG. 4 depicts a method, in accordance with an embodiment.
- FIG. 5A depicts selection between W-FOV and a N-FOV views at a first time, in accordance with an embodiment.
- FIG. 5B depicts selection between W-FOV and a N-FOV views at a second time, in accordance with an embodiment.
- FIGs. 6A to 6C depict views of a user interface display on a video capture device, in accordance with embodiments.
- FIGs. 7A and 7B depict embodiments of views of a user interface display on a video capture device.
- FIGs. 8A and 8B depict embodiments of views of a user interface display on a video capture device.
- FIGs. 9A and 9B depict embodiments of views of a user interface display on a video capture device.
- modules that carry out (i.e., perform, execute, and the like) various functions that are described herein in connection with the respective modules.
- a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) deemed suitable by those of skill in the relevant art for a given implementation.
- ASICs application-specific integrated circuits
- FPGAs field programmable gate arrays
- Each described module may also include instructions executable for carrying out the one or more functions described as being carried out by the respective module, and it is noted that those instructions could take the form of or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and/or the like, and may be stored in any suitable non-transitory computer-readable medium or media, such as commonly referred to as RAM, ROM, etc.
- modules that carry out (i.e., perform, execute, and the like) various functions that are described herein in connection with the respective modules.
- a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) deemed suitable by those of skill in the relevant art for a given implementation.
- ASICs application-specific integrated circuits
- FPGAs field programmable gate arrays
- Each described module may also include instructions executable for carrying out the one or more functions described as being carried out by the respective module, and it is noted that those instructions may take the form of or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and/or the like, and may be stored in any suitable non-transitory computer-readable medium or media, such as commonly referred to as RAM, ROM, etc.
- Exemplary embodiments disclosed herein are implemented using one or more wired and/or wireless network nodes, such as a wireless transmit/receive unit (WTRU) or other network entity.
- WTRU wireless transmit/receive unit
- FIG. 1 A is a system diagram of an exemplary WTRU 102, which may be employed as a client device, video capture device, or other components in embodiments described herein.
- the WTRU 102 may include a processor 118, a communication interface 119 including a transceiver 120, a transmit/receive element 122, a speaker/microphone 124, a keypad 126, a display/touchpad 128, a nonremovable memory 130, a removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and sensors 138.
- GPS global positioning system
- the processor 118 may be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Array (FPGAs) circuits, any other type of integrated circuit (IC), a state machine, and the like.
- the processor 118 may perform signal coding, data processing, power control, input/output processing, and/or any other functionality that enables the WTRU 102 to operate in a wireless environment.
- the processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit/receive element 122. While FIG. 1A depicts the processor 118 and the transceiver 120 as separate components, it will be appreciated that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
- the transmit/receive element 122 may be configured to transmit signals to, or receive signals from, a base station over the air interface 116.
- the transmit/receive element 122 may be an antenna configured to transmit and/or receive RF signals.
- the transmit/receive element 122 may be an emitter/detector configured to transmit and/or receive IR, UV, or visible light signals, as examples.
- the transmit/receive element 122 may be configured to transmit and receive both RF and light signals. It will be appreciated that the transmit/receive element 122 may be configured to transmit and/or receive any combination of wireless signals.
- the WTRU 102 may include any number of transmit/receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit/receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
- the transceiver 120 may be configured to modulate the signals that are to be transmitted by the transmit/receive element 122 and to demodulate the signals that are received by the transmit/receive element 122.
- the WTRU 102 may have multi-mode capabilities.
- the transceiver 120 may include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs, such as UTRA and IEEE 802.11 , as examples.
- the processor 118 of the WTRU 102 may be coupled to, and may receive user input data from, the speaker/microphone 124, the keypad 126, and/or the display/touchpad 128 (e.g., a liquid crystal display (LCD) display unit or organic light-emitting diode (OLED) display unit).
- the processor 118 may also output user data to the speaker/microphone 124, the keypad 126, and/or the display/touchpad 128.
- the processor 118 may access information from, and store data in, any type of suitable memory, such as the non-removable memory 130 and/or the removable memory 132.
- the non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device.
- the removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like.
- SIM subscriber identity module
- SD secure digital
- the processor 118 may access information from, and store data in, memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
- the processor 118 may receive power from the power source 134, and may be configured to distribute and/or control the power to the other components in the WTRU 102.
- the power source 134 may be any suitable device for powering the WTRU 102.
- the power source 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium- ion (Li-ion), and the like), solar cells, fuel cells, and the like.
- the processor 118 may also be coupled to the GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102.
- location information e.g., longitude and latitude
- the WTRU 102 may receive location information over the air interface 116 from a base station and/or determine its location based on the timing of the signals being received from two or more nearby base stations. It will be appreciated that the WTRU 102 may acquire location information by way of any suitable location-determination method while remaining consistent with an embodiment.
- the processor 118 may further be coupled to other peripherals 138, which may include one or more software and/or hardware modules that provide additional features, functionality and/or wired or wireless connectivity.
- the peripherals 138 may include sensors such as an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photographs or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands free headset, a Bluetooth® module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, and the like.
- sensors such as an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photographs or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands free headset, a Bluetooth® module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module
- FIG. 1 B depicts an exemplary network entity 190 that may be used in embodiments of the present disclosure, for example as an encoder, transport packager, origin server, edge streaming server, web server, or client device as described herein.
- network entity 190 includes a communication interface 192, a processor 194, and non-transitory data storage 196, all of which are communicatively linked by a bus, network, or other communication path 198.
- Communication interface 192 may include one or more wired communication interfaces and/or one or more wireless-communication interfaces. With respect to wired communication, communication interface 192 may include one or more interfaces such as Ethernet interfaces, as an example. With respect to wireless communication, communication interface 192 may include components such as one or more antennae, one or more transceivers/chipsets designed and configured for one or more types of wireless (e.g., LTE) communication, and/or any other components deemed suitable by those of skill in the relevant art. And further with respect to wireless communication, communication interface 192 may be equipped at a scale and with a configuration appropriate for acting on the network side— as opposed to the client side— of wireless communications (e.g., LTE communications, Wi-Fi communications, and the like). Thus, communication interface 192 may include the appropriate equipment and circuitry (perhaps including multiple transceivers) for serving multiple mobile stations, UEs, or other access terminals in a coverage area.
- wireless communication interface 192 may include the appropriate equipment and circuitry (perhaps including multiple transceivers)
- Processor 194 may include one or more processors of any type deemed suitable by those of skill in the relevant art, some examples including a general-purpose microprocessor and a dedicated DSP.
- Data storage 196 may take the form of any non-transitory computer-readable medium or combination of such media, some examples including flash memory, read-only memory (ROM), and random- access memory (RAM) to name but a few, as any one or more types of non-transitory data storage deemed suitable by those of skill in the relevant art could be used.
- data storage 196 contains program instructions 197 executable by processor 194 for carrying out various combinations of the various network-entity functions described herein. Exemplary Video Systems and Methods
- FIG. 2 depicts a client device with two video cameras, in accordance with an embodiment.
- a client device may have a greater number of video cameras.
- client device 202 may be a WTRU such as WTRU 102 of FIG. 1A.
- the client device 202 includes a front side opposing the back side, the front side having a display and a user interface.
- the client device 202 includes two cameras as peripherals on the back side, a first camera 206 having a relatively wider field of view (FOV) and a second camera 204 having a relatively narrower FOV that is contained entirely within the FOV of the first camera.
- FOV field of view
- the second camera 204 may be a narrow field of view (FOV) camera, and the first camera may be a wide FOV camera.
- the client device 202 may be a smart phone, a tablet, a dedicated video recording device, or the like.
- the wider FOV (W-FOV) camera 206 includes a 28-mm equivalent lens with optical settings of f/1.8 and the narrower FOV (N-FOV) camera 204 includes a 56-mm equivalent lens with optical settings of f/2.8, thus having a doubling (2x) zoom function over the W-FOV camera 206.
- either one or both of the N-FOV and W-FOV camera optics are variable, with each lens able to adjust the zoom within an optical zoom range.
- the different cameras may also have different zoom capabilities, spatial orientations, or different spatial resolutions.
- the device provides a streaming video output of a selected object of interest.
- the first camera captures relatively wider FOV (W-FOV) video, and the position of the selected object is tracked within the W-FOV video. Based on this tracking information, the device determines when the selected object is positioned within the field of view of the second camera, which has a relatively narrower FOV (N-FOV), and when the selected object is positioned only within the field of view of the first, wider-FOV camera.
- W-FOV relatively wider FOV
- the N-FOV video is used to generate the streaming output video
- the W-FOV video is used to generate the streaming output video. While the W-FOV video is used to generate the streaming output video, the W-FOV video is processed by cropping the video down to a zoom region that contains the object of interest and by upscaling the zoom region to match a zoom factor of the output provided from the second field of view camera.
- the entire spatial extent of the N-FOV video is used as the output video (when the tracked object is within the N-FOV).
- the N-FOV video is cropped and/or is upscaled or downscaled for use as the output video (e.g. as a result of image stabilization or to provide a digital zoom effect).
- the W-FOV video is cropped and upscaled to match the output generated using the N-FOV video. For example, if the output generated using the N-FOV video is ⁇ ⁇ ⁇ pixels, then, in an exemplary embodiment, the W-FOV video is cropped and upscaled to ⁇ ⁇ ⁇ pixels.
- the N-FOV camera has a zoom factor of two as compared to the W-FOV camera.
- the output of the W-FOV camera may be cropped down to a 1 / 2 M ⁇ 1 ⁇ 2N zoom region and upscaled by a factor of two to give an MxN output.
- FIG. 3A depicts a first system architecture for use of two video cameras, in accordance with an embodiment.
- FIG. 3A depicts the system 300 that includes the N-FOV camera 204, 302 and the W-FOV camera 206, 304 from FIG. 2, and also includes a W-FOV decompression module 306, a tracker module 312, a zoom region selection module 310, a crop and upscale module 314, an output selection module 318, a N-FOV decompression module 308, an alignment and disparity compensation module 316, a filter and smooth module 320, and an output video 322.
- both the W-FOV camera 302 and the N-FOV camera 304 produce compressed video output.
- the W-FOV camera 302 provides a compressed video output to the W-FOV decompression module 306 and the N-FOV camera 304 provides a compressed video output to the N-FOV decompression module 308.
- the decompression modules 306, 308 may decompress the compressed video into baseband frames.
- the decompressed video data obtained from the W-FOV camera 302 is provided to both the tracker module 312 and the crop and upscale module 314.
- the decompressed video data obtained from the N-FOV camera 304 is provided to the alignment and disparity compensation module 316.
- FIG. 3B depicts a second system architecture for use of two video cameras, in accordance with an embodiment.
- FIG. 3B depicts the system architecture 330 that is similar to the system 300 of FIG. 3A; however, the W-FOV camera 332 and the N-FOV camera 334 output uncompressed video, and the respective decompression modules 304, 314 are not used to decompress the video into baseband frames.
- the W-FOV camera 332 provides uncompressed video to the tracker module 312 and the crop and upscale module 314 without using a W-FOV decompression module.
- the N- FOV camera 334 may provide uncompressed video to the alignment and disparity compensation module 316 without using an N-FOV decompression module.
- the system 330 may be modified to include a decompression module 304, 314 for the camera providing the compressed video, as disclosed in FIG. 3A.
- the tracker module 312 determines location coordinates of the selected object within the W-FOV camera 302 view as the object moves within the camera view.
- the tracking of the object may be based on image recognition or other object tracking techniques.
- the position of the object is determined on a frame-by-frame basis. In other embodiments, the position of the object is determined less frequently.
- the tracker module 312 is assisted in tracking the selected object based on inputs from a radio frequency identification device (RFID) tag on the object. For example, an athlete may be wearing an RFID tag in communication with an RFID location system. Based on the determined location of the athlete's RFID tag and the orientation of the cameras, the location of the athlete may be determined.
- RFID radio frequency identification device
- the object to be tracked may be selected using a variety of techniques.
- the tracked object may be determined by a user interacting with a touchscreen display, for example by tapping onto an object or icon displayed on a touchscreen display.
- Example tracked objects may be a particular person (e.g., a sports player), or a particular region (e.g., a goal, a basketball hoop, a sports ball). Additional user interfaces may be used to select the zoom regions.
- a voice recognition module interprets instructions to select a certain person or a certain team's goal.
- the tracked object is automatically selected based on preset criteria for tracking objects of interest.
- the zoom region selection module 310 determines a region within the W-FOV video to be used as the zoom region.
- the zoom region is a region that has the same aspect ratio as the N-FOV output video and that is smaller (in linear dimensions) than the N-FOV output video by an amount equal to the relative zoom factor of the N-FOV output video. For example, if the output generated using the N-FOV video is ⁇ ⁇ ⁇ pixels and the relative zoom factor of the N-FOV camera is z (including any digital zoom performed on the N-FOV prior to output), then, in an exemplary embodiment, the zoom region has a size of M/z ⁇ N/z.
- the zoom region selection module 310 further operates to determine a position of the zoom region within the W-FOV video. In some embodiments, the position of the zoom region is selected such that the tracked object is substantially centered in the zoom region.
- the zoom region selection module 310 then provides data related to the zoom region to the tracker module 312 and to the crop and upscale module 314.
- the crop and upscale module 314 crops the video obtained by the W-FOV camera 302 to include the zoom region.
- the cropping of the video stream is updated based on inputs received from the tracker module 312.
- the position of the zoom region may be provided to the crop and upscale module as coordinate points of a bounding box around the zoom region.
- each frame from the W-FOV baseband video is cropped.
- the cropped video is then upscaled by a factor proportional to the ratio of zoom factors, which may be a ratio of the W-FOV focal length and the N-FOV focal length.
- the upscale factor would be a factor of two, doubling the size of the W-FOV cropped image.
- the crop and upscale module 314 provides the cropped and upscaled video of the zoom region to the output selection module 318.
- the alignment and disparity compensation module 316 determines the time alignment and location disparity of the N-FOV video data relative to the W-FOV image in order to overlay onto the spatial coordinate system of the W-FOV video.
- the N-FOV video may be overlaid on the cropped and upscaled W-FOV video after being aligned. This permits the transition between the selected videos.
- the output selection module 318 makes a determination of whether to output the cropped and upscaled video captured from the W-FOV camera 302 or to output the video captured from the N-FOV camera 304.
- the N-FOV output is selected to be the output and provided to the filter and smooth module 320. If the zoom region is not captured within the N-FOV camera 304, the cropped and upscaled video output of the zoom region captured by the W-FOV camera 302 is selected to be the output video and provided to the filter and smooth module 320.
- the output selection module 318 may operate in different ways depending on the type of information available from the tracker module.
- the tracker module may operate to provide a detailed perimeter (e.g. in the form of a polygon) of the tracked object.
- the object may be considered to be outside the N-FOV if any portion of the polygon is outside the N-FOV.
- the object may be considered to be outside the N-FOV if at least a threshold percentage of the polygon is outside the N-FOV.
- the tracker module may provide only coordinates for the tracked object. In some such embodiments, the object may be considered to be outside the N-FOV if the coordinates are outside the N-FOV.
- a bounding box may be defined and centered on the coordinates, and the object may be considered to be outside the N- FOV if any part of the bounding box is outside the N-FOV.
- the dimensions of the bounding box may be predefined dimensions, they may be dimensions based on an estimate of the object's apparent size (e.g. generated by the tracker module), or they may be set by a user, among other options.
- the filter and smooth module 320 provides an output video 322 to be displayed.
- the output video 322 may be generated using either the video captured by the W-FOV camera 302, or the video captured by the N-FOV camera 304, as determined by the output selection module 318.
- the filter and smooth module 320 processes the video to provide a smooth transition between the video from the W-FOV camera 302 and the N-FOV camera 304.
- the output video 322 is displayed on a screen of the video recording device.
- different video may be displayed on the user-interface display of the video recording device.
- the video recording device may display the video obtained from the W-FOV camera 302. Additional annotations may be displayed over the W-FOV video on the display device.
- the zoom region may be indicated with a rectangle box around the zoom region, and the location of the N-FOV may be similarly indicated.
- the output video 322 may be saved on the video capture device for later editing, may be streamed to the Internet or a remote party (or device), or may be output in other ways.
- the W-FOV video is displayed on a screen of the device, and a visual indicator is also displayed as an overlay on the video to indicate when the tracked object is not within the N- FOV.
- the indicator may be in the format of a highlighted box or some other visual cue. The indicator prompts the user to move the camera in the appropriate direction such that the tracked object may be recaptured, or within the field of view, of the N-FOV camera 304.
- the client device capturing the video with both the N-FOV and W-FOV cameras is a single device, such as a smartphone, and the smartphone's display is displaying a representation of video captured from the W-FOV camera, with the indications being displayed on the smartphone's touchscreen display of the zoom region and the N-FOV view.
- a higher frame rate camera is used with a lower frame rate camera.
- the cameras may be offset and the output videos of the cameras are combined to create a higher frame rate signal.
- the combined video is corrected for spatial disparity and lens variability.
- FIG. 3C depicts a third system architecture for use of multiple video cameras, in accordance with an embodiment.
- FIG. 3C depicts the system 360 that is similar to the systems 300 and 330 but includes the fine crop and upscale module 362.
- video captured by the N-FOV camera 304 is provided to the fine crop and upscale module 362.
- the fine crop and upscale module 362 receives data from the tracker module 312 regarding the location of the tracked object, similar to the data received by the crop and upscale module 314.
- the fine crop and upscale module 362 crops and upscales the video captured by the N-FOV camera 304 to more closely zoom in on the tracked object.
- the fine crop and upscale module 362 then provides a cropped and zoomed version of the N-FOV video to the output selection module 312 for later output.
- FIG. 4 depicts a method, in accordance with an embodiment.
- FIG. 4 depicts the method 400.
- the steps of method 400 may be accomplished with the systems depicted in FIGs. 3A-C, or other similar system, as known by those with skill in the art.
- a W-FOV image of a scene is captured.
- the W-FOV image may be a frame of video captured by the W-FOV camera 206.
- the location of the zoom region is tracked within the W-FOV view at 404, and portion of the W-FOV view is determined to be a zoom region (e.g., by the zoom region selection module 308).
- a N-FOV image of the scene is captured at 406.
- the N-FOV image may be a frame of video captured by the N-FOV camera 204.
- a process 408 determines if the tracked object is within the N-FOV image. If the tracked object is within the N-FOV image, the N-FOV image is output at 410. If the zoom region is not within the N-FOV image, a cropped and upscaled W-FOV image is generated in step 411 and output at 412. While FIG. 4 illustrates an embodiment in which the W-FOV image is cropped and upscaled only if the object is not within the N-FOV, in other embodiments, the W-FOV video is cropped and upscaled regardless of whether the object is within the N-FOV (although that cropped and upscaled video may not be output).
- a filtered and smoothed video is output during a transition between outputting the N- FOV video and the W-FOV video.
- the N-FOV camera capturing the video may include a variable zoom.
- the zoom settings of the N-FOV camera may be changed to attempt to keep the tracked object within the N- FOV. For example, the N-FOV camera may zoom out if the tracked object has moved outside of the N-FOV. If the N-FOV camera optics are unable to change any further to obtain the tracked object within the N-FOV view, a cropped and upscaled W-FOV video may then be output.
- the size of the zoom region may be adjusted for different zoom settings of the N-FOV camera.
- the methods 300, 330, 360, 400 of FIGs. 3A, 3B, 3C, and 4 may be performed in real time as the video is being captured by the client device, without intermediate storage of the W-FOV and N-FOV videos.
- the methods 300, 330, 360, 400 of FIGs. 3A, 3B, 3C, and 4 may be performed within milli-seconds.
- the methods 300, 330, 360, 400 of FIGs. 3A, 3B, 3C, and 4 may capture video with a N-FOV camera and a W-FOV camera and store the video in memory for analysis (which may include cropping, upscaling, filtering, and smoothing) at a later point in time.
- a non-transitory, computer- readable medium capable of storing video captured by a N-FOV camera and video captured by a W-FOV camera may be used.
- outputting the output video stream comprises recording the output video stream.
- FIGs. 5A and 5B illustrate an embodiment in which a basketball player 501 has been selected as a tracked object.
- a W-FOV camera 582 captures W-FOV video of a scene
- a N-FOV camera 584 captures N-FOV video of the scene at a first time t1.
- Frame 500 of W-FOV video and frame 504 of N-FOV video are captured substantially simultaneously.
- Frame 500 of W-FOV video, on the left includes images of two basketball players, and frame 504 of the N-FOV video, on the right, encompasses only one of those basketball players, tracked player 501.
- the location of the N-FOV within the W-FOV is depicted at 502 with a dashed rectangle.
- the W-FOV camera 582 and the N-FOV camera 584 respectively capture a frame 550 of W-FOV video and a frame 554 of N-FOV video at a second time t2.
- Frame 550 has been captured shortly after frame 500, with the player 501 who was within the N-FOV in FIG. 5A having moved towards the left within the W-FOV image, such that the tracked player 501 is no longer within the N- FOV 502 and thus is not visible in frame 554.
- the player's movement is tracked in the W-FOV video using video object tracking to determine the player's position in frame 550. While the players and the camera are in the relative positions illustrated in FIG.
- FIGs. 6A to 6C depict views of a user interface display on a video capture device 650, which may be a WTRU.
- the video capture device 650 includes both a N-FOV camera and a W-FOV camera.
- FIG. 6A depicts a first view of the user interface, in accordance with an embodiment.
- the video capture device 650 has a screen 652 on which the W-FOV video is being displayed in real time (as the video is captured). Also displayed on the screen 652 is a rectangle 662 that indicates the field of view of the N-FOV camera, which is narrower than and contained within the field of view of the W-FOV camera.
- the N-FOV camera has a zoom factor of two compared to the W- FOV camera, and the two cameras have the same aspect ratio, so the rectangle 662 subtends half the width and half the height of the W-FOV video.
- the N-FOV camera and W-FOV camera have fixed orientations relative to one another, so the rectangle 662 remains in the same position with respect to the W- FOV video.
- Two basketball players are displayed on the screen 652 in the W-FOV video.
- the user has selected one of those players, namely player 601 , to be tracked.
- the selection of player 601 may be made by the user touching the position of player 601 on the screen 652.
- a rectangle 602 is displayed to indicate to the user the position of the zoom region. This indication allows the user to turn the device to better align the zoom region 602 with the N-FOV indicated by rectangle 662.
- the tracked object 601 is within the N-FOV, so the N-FOV camera is used to generate the output video (e.g. for storage or for streaming to another device).
- the tracked player 601 has moved to the left (and the user of the device has failed to follow by turning the device).
- the position of the rectangle 602 indicating the zoom region has moved accordingly.
- video from the W-FOV camera is used to generate the output video.
- the zoom region may be cropped and upscaled to provide the output video.
- the switching of the source of the output video may be transparent to the user. In such embodiments, the user can try to follow the tracked object while the device automatically and transparently selects which camera to use for the output video.
- FIG. 7A and 7B depict views of a user interface display on a video capture device 750, which may be a WTRU.
- the video capture device 750 includes both a N-FOV camera and a W-FOV camera.
- the video capture device 750 has a screen 752 on which the W-FOV video is being displayed in real time (as it is captured).
- a rectangle 762 that indicates the field of view of the N-FOV camera.
- the user has selected player 601 to be tracked.
- a tracking module generates a bounding box representing an approximation of the extent of the tracked object (in this case player 601) around the tracked player, and that bounding box is displayed on the screen as rectangle 702.
- the N-FOV camera is used to generate the output video.
- the N-FOV camera is used to generate the output video as long as at least a threshold percentage (e.g., 90%) of the bounding box is within the N-FOV.
- the tracked player 601 has moved to the left.
- the position of the rectangle 702 indicating the bounding box has been updated accordingly.
- the extent of the bounding box has changed as the posture of the tracked player 601 has changed, but in some embodiments, the extent of the bounding box may be fixed (e.g., using predetermined or user- defined dimensions). Because the bounding box is now at least partially outside the N-FOV, video from the W-FOV camera is used to generate the output video. In particular, the zoom region (not displayed in FIGs. 7A-7B) may be cropped and upscaled to provide the output video.
- FIGs. 8A and 8B illustrate an embodiment under the same video recording conditions as those in FIGs. 7A and 7B.
- the zoomed output video is displayed on the user interface.
- tracked player 601 is within the N-FOV, and high- quality zoomed video from the N-FOV camera is displayed on the screen 852 of the video recording device 850.
- the tracked player moves out of the N-FOV, and, in response, the device automatically switches to output of cropped and upscaled video on the screen 852.
- this output video may have a lower quality than the video captured using the native resolution of the N-FOV camera (represented schematically by the pixelation in FIG. 8B).
- the tracked player 601 remains within the output video.
- an indicator such as the arrow 854 may be displayed on the screen 852 to indicate to the user which way to turn the device such that the tracked player 601 will again be within the N-FOV.
- the display of the video capture device may return to displaying a representation of the N-FOV camera video after the user reacquires the zoom region within the view of the N-FOV camera.
- FIGs. 9A and 9B illustrate an embodiment under the same video recording conditions as those in FIGs. 7A and 7B.
- the device 950 displays the N-FOV video on the display 952 while the tracked object is within the N-FOV, and the device switches to display of the W-FOV video when the tracked object is not within the W-FOV. This allows the user to more easily reorient the device such that the tracked object is again within the N-FOV.
- tracked player 601 is within the N-FOV
- high-quality zoomed video from the N-FOV camera is displayed on the screen 952 of the video recording device 950.
- the device further displays a rectangle 962 representing the N-FOV.
- the device may further display an indication, such as rectangle 702, to identify which object is being tracked.
- the device may display an indication of the zoom region (such as the rectangle 602 of FIGs. 6B-6C).
- the device 602 may switch back to displaying the N-FOV video if the tracked object is reacquired within the N- FOV.
- a device with more than two cameras may be used.
- an object may be tracked to identify which cameras out of a plurality of cameras have that object within their respective fields of view.
- An output video may be generated by selecting which one of those cameras will be used to provide the output any particular time.
- the output is selected from the camera with the narrowest field of view that still includes the tracked object within its field of view.
- Video from cameras with a wider field of view may be cropped and upscaled as appropriate to match the field of view of cameras with a narrower field of view.
- ROM read only memory
- RAM random access memory
- register cache memory
- semiconductor memory devices magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks, and digital versatile disks (DVDs).
- a processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Human Computer Interaction (AREA)
- Studio Devices (AREA)
Abstract
Systems and methods described herein disclose use of simultaneous and synchronous multiple-camera captures for a zoom region. An exemplary method using two field of view (FOV) video streams of a scene, where the second FOV is narrower than the first, comprises: tracking an object captured within the first FOV; responsive to determining that the object is entirely within the second FOV, outputting video corresponding to the second FOV; and responsive to determining that the object is outside the second FOV, outputting a cropped and up-scaled representation of video corresponding to the first FOV. Systems and methods disclosed herein, prior to tracking the object, display video captured for the first FOV and receive user input indicating selection of an object to be tracked in the displayed video for the first FOV.
Description
ZOOM CODING USING SIMULTANEOUS AND SYNCHRONOUS MULTIPLE-CAMERA CAPTURES
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001 ] The present application is a non-provisional filing of, and claims benefit under 35 U.S.C. §119(e) from, U.S. Provisional Patent Application Serial No. 62/468,773, entitled "Zoom Coding Using Simultaneous and Synchronous Multiple Camera Captures," filed March 8, 2017, the entirety of which is incorporated herein by reference.
BACKGROUND
[0002] In many video distribution systems, images are captured by lenses with different optical parameters, such as different visual zoom factors. Different camera optics affect the quality and appearance of the images captured by the cameras. These images may undergo further digital enhancement, such as digitally changing the zoom and field of view (FOV) parameters.
[0003] At times, multiple different video cameras can be used to capture the same event. The multiple different video cameras may have different optical capabilities, such as different optical zoom settings. The multiple video cameras may also be on the same device, such as a smart phone having two cameras oriented in the same direction.
SUMMARY
[0004] An exemplary method comprises capturing a first video of a scene with a first camera on a user device, where the first camera has a first field of view (e.g. a wide-angle field of view) with a first extent (e.g. a field of view between around 64° and 84° in the vertical and/or horizontal direction). A second camera (e.g. a zoom camera) on the user device simultaneously captures a video of a portion of the same scene. The second camera has a second field of view having a second extent that is narrower than the first extent (e.g. less than around 64°). The second extent may be represented in terms of pixels within the first video. The second field of view is contained within the first field of view, such objects within the field of view of the second camera are also within the field of view of the first camera. The method further comprises tracking a position of an object of interest with respect to the first and second fields of view. This may be done using, for example, RFID or other wireless signals, or performing optical tracking using object recognition, among other possibilities. In the exemplary method, a determination is made of whether the object of interest is positioned
entirely within the second field of view. In the method, a video output is provided. If the object of interest is positioned entirely within the second field of view, the video output comprises the second video. Otherwise, a cropped-and-scaled version of the first video is generated such that the video output includes the entirety of the object of interest and such that the extent of the output video corresponds to the second extent, and the cropped-and-scaled version of the first video is provided as the output.
[0005] An exemplary method for providing an output video stream in a device having a first camera and a second camera, the second camera having a field of view that is narrower than and contained within a field of view of the first camera, comprises tracking an object in video captured by the first camera. A determination, based on the tracking, is made on whether the object is entirely within the field of view of the second camera. If the object is within the field of view of the second camera, video captured by the second camera is outputted as the output video stream. If the object is outside the field of view of the second camera, video captured by the first camera is cropped and upscaled to generate a cropped-and-upscaled video of the object, and the generated video is outputted as the output video stream.
[0006] For one embodiment, the method is performed in real time. The method may also include, before tracking the object, displaying, on a screen of the device, video captured by the first camera and receiving user input selecting the object to be tracked in the displayed video. For one embodiment, the output video stream may be displayed on the device. For one embodiment, an indication of the position of the object may be displayed if the object is outside the field of view of the second camera. For one embodiment, tracking the object may be done using RFID or other wireless protocol signals.
BRIEF DESCRIPTION OF THE DRAWINGS
[0007] FIG. 1 A is a functional block diagram of a client device that may be used in some embodiments.
[0008] FIG. 1 B is a functional block diagram of a network entity that may be used in some embodiments.
[0009] FIG. 2 depicts a client device with multiple video cameras, in accordance with an embodiment.
[0010] FIG. 3A depicts a first system architecture for use of multiple video cameras, in accordance with an embodiment.
[0011] FIG. 3B depicts a second system architecture for use of multiple video cameras, in accordance with an embodiment.
[0012] FIG. 3C depicts a third system architecture for use of multiple video cameras, in accordance with an embodiment.
[0013] FIG. 4 depicts a method, in accordance with an embodiment.
[0014] FIG. 5A depicts selection between W-FOV and a N-FOV views at a first time, in accordance with an embodiment.
[0015] FIG. 5B depicts selection between W-FOV and a N-FOV views at a second time, in accordance with an embodiment.
[0016] FIGs. 6A to 6C depict views of a user interface display on a video capture device, in accordance with embodiments.
[0017] FIGs. 7A and 7B depict embodiments of views of a user interface display on a video capture device.
[0018] FIGs. 8A and 8B depict embodiments of views of a user interface display on a video capture device.
[0019] FIGs. 9A and 9B depict embodiments of views of a user interface display on a video capture device.
DETAILED DESCRIPTION
[0020] A detailed description of illustrative embodiments will now be provided with reference to the various Figures. Although this description provides detailed examples of possible implementations, it should be noted that the provided details are intended to be by way of example and in no way limit the scope of the application.
[0021] Note that various hardware elements of one or more of the described embodiments are referred to as "modules" that carry out (i.e., perform, execute, and the like) various functions that are described herein in connection with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) deemed suitable by those of skill in the relevant art for a given implementation. Each described module may also include instructions executable for carrying out the one or more functions described as being carried out by the respective module, and it is noted that those instructions could take the form of or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and/or the like, and may be stored in any suitable non-transitory computer-readable medium or media, such as commonly referred to as RAM, ROM, etc.
Exemplary Client and Server Hardware
[0022] Note that various hardware elements of one or more of the described embodiments are referred to as "modules" that carry out (i.e., perform, execute, and the like) various functions that are described herein
in connection with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) deemed suitable by those of skill in the relevant art for a given implementation. Each described module may also include instructions executable for carrying out the one or more functions described as being carried out by the respective module, and it is noted that those instructions may take the form of or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and/or the like, and may be stored in any suitable non-transitory computer-readable medium or media, such as commonly referred to as RAM, ROM, etc.
[0023] Exemplary embodiments disclosed herein are implemented using one or more wired and/or wireless network nodes, such as a wireless transmit/receive unit (WTRU) or other network entity.
[0024] FIG. 1 A is a system diagram of an exemplary WTRU 102, which may be employed as a client device, video capture device, or other components in embodiments described herein. As shown in FIG. 1 A, the WTRU 102 may include a processor 118, a communication interface 119 including a transceiver 120, a transmit/receive element 122, a speaker/microphone 124, a keypad 126, a display/touchpad 128, a nonremovable memory 130, a removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and sensors 138. It will be appreciated that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with an embodiment.
[0025] The processor 118 may be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Array (FPGAs) circuits, any other type of integrated circuit (IC), a state machine, and the like. The processor 118 may perform signal coding, data processing, power control, input/output processing, and/or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit/receive element 122. While FIG. 1A depicts the processor 118 and the transceiver 120 as separate components, it will be appreciated that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0026] The transmit/receive element 122 may be configured to transmit signals to, or receive signals from, a base station over the air interface 116. For example, in one embodiment, the transmit/receive element 122 may be an antenna configured to transmit and/or receive RF signals. In another embodiment, the transmit/receive element 122 may be an emitter/detector configured to transmit and/or receive IR, UV, or visible light signals, as examples. In yet another embodiment, the transmit/receive element 122 may be
configured to transmit and receive both RF and light signals. It will be appreciated that the transmit/receive element 122 may be configured to transmit and/or receive any combination of wireless signals.
[0027] In addition, although the transmit/receive element 122 is depicted in FIG. 1 A as a single element, the WTRU 102 may include any number of transmit/receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit/receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0028] The transceiver 120 may be configured to modulate the signals that are to be transmitted by the transmit/receive element 122 and to demodulate the signals that are received by the transmit/receive element 122. As noted above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs, such as UTRA and IEEE 802.11 , as examples.
[0029] The processor 118 of the WTRU 102 may be coupled to, and may receive user input data from, the speaker/microphone 124, the keypad 126, and/or the display/touchpad 128 (e.g., a liquid crystal display (LCD) display unit or organic light-emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker/microphone 124, the keypad 126, and/or the display/touchpad 128. In addition, the processor 118 may access information from, and store data in, any type of suitable memory, such as the non-removable memory 130 and/or the removable memory 132. The non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 may access information from, and store data in, memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0030] The processor 118 may receive power from the power source 134, and may be configured to distribute and/or control the power to the other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. As examples, the power source 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium- ion (Li-ion), and the like), solar cells, fuel cells, and the like.
[0031] The processor 118 may also be coupled to the GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or in lieu of, the information from the GPS chipset 136, the WTRU 102 may receive location information over the air interface 116 from a base station and/or determine its location based on the timing of the signals being received from two or more nearby base stations. It will be appreciated that the WTRU
102 may acquire location information by way of any suitable location-determination method while remaining consistent with an embodiment.
[0032] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and/or hardware modules that provide additional features, functionality and/or wired or wireless connectivity. For example, the peripherals 138 may include sensors such as an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photographs or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands free headset, a Bluetooth® module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, and the like.
[0033] FIG. 1 B depicts an exemplary network entity 190 that may be used in embodiments of the present disclosure, for example as an encoder, transport packager, origin server, edge streaming server, web server, or client device as described herein. As depicted in FIG. 1 B, network entity 190 includes a communication interface 192, a processor 194, and non-transitory data storage 196, all of which are communicatively linked by a bus, network, or other communication path 198.
[0034] Communication interface 192 may include one or more wired communication interfaces and/or one or more wireless-communication interfaces. With respect to wired communication, communication interface 192 may include one or more interfaces such as Ethernet interfaces, as an example. With respect to wireless communication, communication interface 192 may include components such as one or more antennae, one or more transceivers/chipsets designed and configured for one or more types of wireless (e.g., LTE) communication, and/or any other components deemed suitable by those of skill in the relevant art. And further with respect to wireless communication, communication interface 192 may be equipped at a scale and with a configuration appropriate for acting on the network side— as opposed to the client side— of wireless communications (e.g., LTE communications, Wi-Fi communications, and the like). Thus, communication interface 192 may include the appropriate equipment and circuitry (perhaps including multiple transceivers) for serving multiple mobile stations, UEs, or other access terminals in a coverage area.
[0035] Processor 194 may include one or more processors of any type deemed suitable by those of skill in the relevant art, some examples including a general-purpose microprocessor and a dedicated DSP.
[0036] Data storage 196 may take the form of any non-transitory computer-readable medium or combination of such media, some examples including flash memory, read-only memory (ROM), and random- access memory (RAM) to name but a few, as any one or more types of non-transitory data storage deemed suitable by those of skill in the relevant art could be used. As depicted in FIG. 1 B, data storage 196 contains program instructions 197 executable by processor 194 for carrying out various combinations of the various network-entity functions described herein.
Exemplary Video Systems and Methods
[0037] FIG. 2 depicts a client device with two video cameras, in accordance with an embodiment. In alternative embodiments, a client device may have a greater number of video cameras. In some embodiments, client device 202 may be a WTRU such as WTRU 102 of FIG. 1A. The client device 202 includes a front side opposing the back side, the front side having a display and a user interface. The client device 202 includes two cameras as peripherals on the back side, a first camera 206 having a relatively wider field of view (FOV) and a second camera 204 having a relatively narrower FOV that is contained entirely within the FOV of the first camera. The second camera 204 may be a narrow field of view (FOV) camera, and the first camera may be a wide FOV camera. The client device 202 may be a smart phone, a tablet, a dedicated video recording device, or the like. In an exemplary embodiment, the wider FOV (W-FOV) camera 206 includes a 28-mm equivalent lens with optical settings of f/1.8 and the narrower FOV (N-FOV) camera 204 includes a 56-mm equivalent lens with optical settings of f/2.8, thus having a doubling (2x) zoom function over the W-FOV camera 206. In some embodiments, either one or both of the N-FOV and W-FOV camera optics are variable, with each lens able to adjust the zoom within an optical zoom range. The different cameras may also have different zoom capabilities, spatial orientations, or different spatial resolutions.
[0038] In an exemplary embodiment, the device provides a streaming video output of a selected object of interest. In one such embodiment, the first camera captures relatively wider FOV (W-FOV) video, and the position of the selected object is tracked within the W-FOV video. Based on this tracking information, the device determines when the selected object is positioned within the field of view of the second camera, which has a relatively narrower FOV (N-FOV), and when the selected object is positioned only within the field of view of the first, wider-FOV camera. While the selected object is entirely within the field of view of the second camera, the N-FOV video is used to generate the streaming output video, and while the selected object is not entirely within the field of view of the second camera, the W-FOV video is used to generate the streaming output video. While the W-FOV video is used to generate the streaming output video, the W-FOV video is processed by cropping the video down to a zoom region that contains the object of interest and by upscaling the zoom region to match a zoom factor of the output provided from the second field of view camera.
[0039] In some embodiments, the entire spatial extent of the N-FOV video is used as the output video (when the tracked object is within the N-FOV). In other embodiments, the N-FOV video is cropped and/or is upscaled or downscaled for use as the output video (e.g. as a result of image stabilization or to provide a digital zoom effect). In either case, the W-FOV video is cropped and upscaled to match the output generated using the N-FOV video. For example, if the output generated using the N-FOV video is Μ χΝ pixels, then, in an exemplary embodiment, the W-FOV video is cropped and upscaled to ΜχΝ pixels. In one such case, the N-FOV camera has a zoom factor of two as compared to the W-FOV camera. In that case, the output of the
W-FOV camera may be cropped down to a 1/2M χ ½N zoom region and upscaled by a factor of two to give an MxN output.
[0040] FIG. 3A depicts a first system architecture for use of two video cameras, in accordance with an embodiment. In particular, FIG. 3A depicts the system 300 that includes the N-FOV camera 204, 302 and the W-FOV camera 206, 304 from FIG. 2, and also includes a W-FOV decompression module 306, a tracker module 312, a zoom region selection module 310, a crop and upscale module 314, an output selection module 318, a N-FOV decompression module 308, an alignment and disparity compensation module 316, a filter and smooth module 320, and an output video 322.
[0041] In the embodiment of FIG. 3A, both the W-FOV camera 302 and the N-FOV camera 304 produce compressed video output. The W-FOV camera 302 provides a compressed video output to the W-FOV decompression module 306 and the N-FOV camera 304 provides a compressed video output to the N-FOV decompression module 308. The decompression modules 306, 308 may decompress the compressed video into baseband frames. The decompressed video data obtained from the W-FOV camera 302 is provided to both the tracker module 312 and the crop and upscale module 314. The decompressed video data obtained from the N-FOV camera 304 is provided to the alignment and disparity compensation module 316.
[0042] FIG. 3B depicts a second system architecture for use of two video cameras, in accordance with an embodiment. FIG. 3B depicts the system architecture 330 that is similar to the system 300 of FIG. 3A; however, the W-FOV camera 332 and the N-FOV camera 334 output uncompressed video, and the respective decompression modules 304, 314 are not used to decompress the video into baseband frames. In such an embodiment, the W-FOV camera 332 provides uncompressed video to the tracker module 312 and the crop and upscale module 314 without using a W-FOV decompression module. Additionally, the N- FOV camera 334 may provide uncompressed video to the alignment and disparity compensation module 316 without using an N-FOV decompression module. In some embodiments, only one of the W-FOV camera 332 and the N-FOV camera 334 provides compressed video output, and the other one provides uncompressed video output. In such an embodiment, the system 330 may be modified to include a decompression module 304, 314 for the camera providing the compressed video, as disclosed in FIG. 3A.
[0043] Returning to FIG. 3A, the tracker module 312 determines location coordinates of the selected object within the W-FOV camera 302 view as the object moves within the camera view. The tracking of the object may be based on image recognition or other object tracking techniques. In some embodiments, the position of the object is determined on a frame-by-frame basis. In other embodiments, the position of the object is determined less frequently. In some embodiments, the tracker module 312 is assisted in tracking the selected object based on inputs from a radio frequency identification device (RFID) tag on the object. For example, an athlete may be wearing an RFID tag in communication with an RFID location system. Based on
the determined location of the athlete's RFID tag and the orientation of the cameras, the location of the athlete may be determined.
[0044] The object to be tracked may be selected using a variety of techniques. For example, the tracked object may be determined by a user interacting with a touchscreen display, for example by tapping onto an object or icon displayed on a touchscreen display. Example tracked objects may be a particular person (e.g., a sports player), or a particular region (e.g., a goal, a basketball hoop, a sports ball). Additional user interfaces may be used to select the zoom regions. In some embodiments, a voice recognition module interprets instructions to select a certain person or a certain team's goal. In other embodiments, the tracked object is automatically selected based on preset criteria for tracking objects of interest.
[0045] The zoom region selection module 310 determines a region within the W-FOV video to be used as the zoom region. In an exemplary embodiment, the zoom region is a region that has the same aspect ratio as the N-FOV output video and that is smaller (in linear dimensions) than the N-FOV output video by an amount equal to the relative zoom factor of the N-FOV output video. For example, if the output generated using the N-FOV video is ΜχΝ pixels and the relative zoom factor of the N-FOV camera is z (including any digital zoom performed on the N-FOV prior to output), then, in an exemplary embodiment, the zoom region has a size of M/z χ N/z. The zoom region selection module 310 further operates to determine a position of the zoom region within the W-FOV video. In some embodiments, the position of the zoom region is selected such that the tracked object is substantially centered in the zoom region.
[0046] The zoom region selection module 310 then provides data related to the zoom region to the tracker module 312 and to the crop and upscale module 314.
[0047] The crop and upscale module 314 crops the video obtained by the W-FOV camera 302 to include the zoom region. The cropping of the video stream is updated based on inputs received from the tracker module 312. The position of the zoom region may be provided to the crop and upscale module as coordinate points of a bounding box around the zoom region. In some embodiments, each frame from the W-FOV baseband video is cropped. The cropped video is then upscaled by a factor proportional to the ratio of zoom factors, which may be a ratio of the W-FOV focal length and the N-FOV focal length. In the example using the W-FOV camera 302 and the N-FOV camera 304, the upscale factor would be a factor of two, doubling the size of the W-FOV cropped image. The crop and upscale module 314 provides the cropped and upscaled video of the zoom region to the output selection module 318.
[0048] The alignment and disparity compensation module 316 determines the time alignment and location disparity of the N-FOV video data relative to the W-FOV image in order to overlay onto the spatial coordinate system of the W-FOV video. Thus, the N-FOV video may be overlaid on the cropped and upscaled W-FOV video after being aligned. This permits the transition between the selected videos.
[0049] The output selection module 318 makes a determination of whether to output the cropped and upscaled video captured from the W-FOV camera 302 or to output the video captured from the N-FOV camera 304. In response to a determination that the zoom region is captured within the field of view of the N-FOV camera 304, the N-FOV output is selected to be the output and provided to the filter and smooth module 320. If the zoom region is not captured within the N-FOV camera 304, the cropped and upscaled video output of the zoom region captured by the W-FOV camera 302 is selected to be the output video and provided to the filter and smooth module 320.
[0050] The output selection module 318 may operate in different ways depending on the type of information available from the tracker module. For example, in some embodiments, the tracker module may operate to provide a detailed perimeter (e.g. in the form of a polygon) of the tracked object. In such an embodiment, the object may be considered to be outside the N-FOV if any portion of the polygon is outside the N-FOV. In another embodiment, the object may be considered to be outside the N-FOV if at least a threshold percentage of the polygon is outside the N-FOV. In some embodiments, the tracker module may provide only coordinates for the tracked object. In some such embodiments, the object may be considered to be outside the N-FOV if the coordinates are outside the N-FOV. In other such embodiments, a bounding box may be defined and centered on the coordinates, and the object may be considered to be outside the N- FOV if any part of the bounding box is outside the N-FOV. The dimensions of the bounding box may be predefined dimensions, they may be dimensions based on an estimate of the object's apparent size (e.g. generated by the tracker module), or they may be set by a user, among other options.
[0051] The filter and smooth module 320 provides an output video 322 to be displayed. The output video 322 may be generated using either the video captured by the W-FOV camera 302, or the video captured by the N-FOV camera 304, as determined by the output selection module 318. During a transition between the two different video sources, the filter and smooth module 320 processes the video to provide a smooth transition between the video from the W-FOV camera 302 and the N-FOV camera 304.
[0052] In some embodiments, the output video 322 is displayed on a screen of the video recording device. In other embodiments, different video may be displayed on the user-interface display of the video recording device. In one such example, the video recording device may display the video obtained from the W-FOV camera 302. Additional annotations may be displayed over the W-FOV video on the display device. For example, the zoom region may be indicated with a rectangle box around the zoom region, and the location of the N-FOV may be similarly indicated. The output video 322 may be saved on the video capture device for later editing, may be streamed to the Internet or a remote party (or device), or may be output in other ways.
[0053] In some embodiments, the W-FOV video is displayed on a screen of the device, and a visual indicator is also displayed as an overlay on the video to indicate when the tracked object is not within the N-
FOV. The indicator may be in the format of a highlighted box or some other visual cue. The indicator prompts the user to move the camera in the appropriate direction such that the tracked object may be recaptured, or within the field of view, of the N-FOV camera 304. In one such embodiment, the client device capturing the video with both the N-FOV and W-FOV cameras is a single device, such as a smartphone, and the smartphone's display is displaying a representation of video captured from the W-FOV camera, with the indications being displayed on the smartphone's touchscreen display of the zoom region and the N-FOV view.
[0054] In some embodiments, a higher frame rate camera is used with a lower frame rate camera. The cameras may be offset and the output videos of the cameras are combined to create a higher frame rate signal. The combined video is corrected for spatial disparity and lens variability.
[0055] FIG. 3C depicts a third system architecture for use of multiple video cameras, in accordance with an embodiment. In particular, FIG. 3C depicts the system 360 that is similar to the systems 300 and 330 but includes the fine crop and upscale module 362. In some embodiments, video captured by the N-FOV camera 304 is provided to the fine crop and upscale module 362. The fine crop and upscale module 362 receives data from the tracker module 312 regarding the location of the tracked object, similar to the data received by the crop and upscale module 314. The fine crop and upscale module 362 then crops and upscales the video captured by the N-FOV camera 304 to more closely zoom in on the tracked object. The fine crop and upscale module 362 then provides a cropped and zoomed version of the N-FOV video to the output selection module 312 for later output.
[0056] FIG. 4 depicts a method, in accordance with an embodiment. In particular, FIG. 4 depicts the method 400. The steps of method 400 may be accomplished with the systems depicted in FIGs. 3A-C, or other similar system, as known by those with skill in the art. At 402, a W-FOV image of a scene is captured. The W-FOV image may be a frame of video captured by the W-FOV camera 206. The location of the zoom region is tracked within the W-FOV view at 404, and portion of the W-FOV view is determined to be a zoom region (e.g., by the zoom region selection module 308). Additionally, a N-FOV image of the scene is captured at 406. The N-FOV image may be a frame of video captured by the N-FOV camera 204.
[0057] For one embodiment, a process 408 determines if the tracked object is within the N-FOV image. If the tracked object is within the N-FOV image, the N-FOV image is output at 410. If the zoom region is not within the N-FOV image, a cropped and upscaled W-FOV image is generated in step 411 and output at 412. While FIG. 4 illustrates an embodiment in which the W-FOV image is cropped and upscaled only if the object is not within the N-FOV, in other embodiments, the W-FOV video is cropped and upscaled regardless of whether the object is within the N-FOV (although that cropped and upscaled video may not be output). In some embodiments, a filtered and smoothed video is output during a transition between outputting the N- FOV video and the W-FOV video.
[0058] In some embodiments, the N-FOV camera capturing the video may include a variable zoom. The zoom settings of the N-FOV camera may be changed to attempt to keep the tracked object within the N- FOV. For example, the N-FOV camera may zoom out if the tracked object has moved outside of the N-FOV. If the N-FOV camera optics are unable to change any further to obtain the tracked object within the N-FOV view, a cropped and upscaled W-FOV video may then be output. The size of the zoom region may be adjusted for different zoom settings of the N-FOV camera.
[0059] The methods 300, 330, 360, 400 of FIGs. 3A, 3B, 3C, and 4 may be performed in real time as the video is being captured by the client device, without intermediate storage of the W-FOV and N-FOV videos. For one embodiment, the methods 300, 330, 360, 400 of FIGs. 3A, 3B, 3C, and 4 may be performed within milli-seconds. For one embodiment, the methods 300, 330, 360, 400 of FIGs. 3A, 3B, 3C, and 4 may capture video with a N-FOV camera and a W-FOV camera and store the video in memory for analysis (which may include cropping, upscaling, filtering, and smoothing) at a later point in time. A non-transitory, computer- readable medium capable of storing video captured by a N-FOV camera and video captured by a W-FOV camera may be used. For one embodiment, outputting the output video stream comprises recording the output video stream.
[0060] FIGs. 5A and 5B illustrate an embodiment in which a basketball player 501 has been selected as a tracked object. In FIG. 5A, a W-FOV camera 582 captures W-FOV video of a scene, and a N-FOV camera 584 captures N-FOV video of the scene at a first time t1. Frame 500 of W-FOV video and frame 504 of N-FOV video are captured substantially simultaneously. Frame 500 of W-FOV video, on the left, includes images of two basketball players, and frame 504 of the N-FOV video, on the right, encompasses only one of those basketball players, tracked player 501. The location of the N-FOV within the W-FOV is depicted at 502 with a dashed rectangle. While the players and the camera are in the relative positions illustrated in FIG. 5A, with player 501 being within the N-FOV, video frames captured by the N-FOV camera are used to generate the output video 506. While a cropped and upscaled image 508 of tracked player 501 may be generated from the W-FOV frame, that image may be expected to have lower quality (represented schematically by pixelation) than the image captured using the N-FOV camera.
[0061] In the example of FIG. 5B, the W-FOV camera 582 and the N-FOV camera 584 respectively capture a frame 550 of W-FOV video and a frame 554 of N-FOV video at a second time t2. Frame 550 has been captured shortly after frame 500, with the player 501 who was within the N-FOV in FIG. 5A having moved towards the left within the W-FOV image, such that the tracked player 501 is no longer within the N- FOV 502 and thus is not visible in frame 554. The player's movement is tracked in the W-FOV video using video object tracking to determine the player's position in frame 550. While the players and the camera are in the relative positions illustrated in FIG. 5B, with player 501 being outside the N-FOV, video frames captured by the W-FOV camera are used to generate the output video. In using frame 550 to generate the W-FOV, a
zoom region is determined based on the position of player 501 , and areas of frame 550 outside the zoom region are cropped away. The remaining zoom region is upscaled to match the zoom-characteristics of the N-FOV camera, resulting in cropped-and-upscaled frame 558. The cropping and upscaling processes generate cropped-and-upscaled frame 558, which is included as frame 560 of the output video. Although cropped and upscaled frame 558 may be of a lower quality than frame 554, frame 558 is chosen for the output video because it includes the tracked object of interest.
[0062] FIGs. 6A to 6C depict views of a user interface display on a video capture device 650, which may be a WTRU. The video capture device 650 includes both a N-FOV camera and a W-FOV camera.
[0063] FIG. 6A depicts a first view of the user interface, in accordance with an embodiment. In the embodiment of FIG. 6A, the video capture device 650 has a screen 652 on which the W-FOV video is being displayed in real time (as the video is captured). Also displayed on the screen 652 is a rectangle 662 that indicates the field of view of the N-FOV camera, which is narrower than and contained within the field of view of the W-FOV camera. In this example, the N-FOV camera has a zoom factor of two compared to the W- FOV camera, and the two cameras have the same aspect ratio, so the rectangle 662 subtends half the width and half the height of the W-FOV video. In this example, the N-FOV camera and W-FOV camera have fixed orientations relative to one another, so the rectangle 662 remains in the same position with respect to the W- FOV video.
[0064] Two basketball players are displayed on the screen 652 in the W-FOV video. In this example, as illustrated in FIG. 6B, the user has selected one of those players, namely player 601 , to be tracked. The selection of player 601 may be made by the user touching the position of player 601 on the screen 652. While the player 601 is being tracked, in this embodiment, a rectangle 602 is displayed to indicate to the user the position of the zoom region. This indication allows the user to turn the device to better align the zoom region 602 with the N-FOV indicated by rectangle 662. Under the conditions illustrated in FIG. 6B, the tracked object 601 is within the N-FOV, so the N-FOV camera is used to generate the output video (e.g. for storage or for streaming to another device).
[0065] Under the conditions illustrated in FIG. 6C, however, the tracked player 601 has moved to the left (and the user of the device has failed to follow by turning the device). The position of the rectangle 602 indicating the zoom region has moved accordingly. Because the tracked player 601 is now at least partially outside the N-FOV, video from the W-FOV camera is used to generate the output video. In particular, the zoom region may be cropped and upscaled to provide the output video. In some embodiments, the switching of the source of the output video may be transparent to the user. In such embodiments, the user can try to follow the tracked object while the device automatically and transparently selects which camera to use for the output video.
[0066] FIGs. 7A and 7B depict views of a user interface display on a video capture device 750, which may be a WTRU. The video capture device 750 includes both a N-FOV camera and a W-FOV camera. In the embodiment of FIG. 7A, the video capture device 750 has a screen 752 on which the W-FOV video is being displayed in real time (as it is captured). Also displayed on the screen 752 is a rectangle 762 that indicates the field of view of the N-FOV camera. The user has selected player 601 to be tracked. In this embodiment, a tracking module generates a bounding box representing an approximation of the extent of the tracked object (in this case player 601) around the tracked player, and that bounding box is displayed on the screen as rectangle 702. In this embodiment, while the bounding box is included entirely within the N- FOV, the N-FOV camera is used to generate the output video. In some embodiments, the N-FOV camera is used to generate the output video as long as at least a threshold percentage (e.g., 90%) of the bounding box is within the N-FOV.
[0067] Under the conditions illustrated in FIG. 7B, however, the tracked player 601 has moved to the left. The position of the rectangle 702 indicating the bounding box has been updated accordingly. In this example, the extent of the bounding box has changed as the posture of the tracked player 601 has changed, but in some embodiments, the extent of the bounding box may be fixed (e.g., using predetermined or user- defined dimensions). Because the bounding box is now at least partially outside the N-FOV, video from the W-FOV camera is used to generate the output video. In particular, the zoom region (not displayed in FIGs. 7A-7B) may be cropped and upscaled to provide the output video.
[0068] FIGs. 8A and 8B illustrate an embodiment under the same video recording conditions as those in FIGs. 7A and 7B. In the embodiment of FIGs. 8A and 8B, however, the zoomed output video is displayed on the user interface. Initially, as illustrated in FIGs. 7A, tracked player 601 is within the N-FOV, and high- quality zoomed video from the N-FOV camera is displayed on the screen 852 of the video recording device 850. Subsequently, the tracked player moves out of the N-FOV, and, in response, the device automatically switches to output of cropped and upscaled video on the screen 852. Due to the upscaling, this output video may have a lower quality than the video captured using the native resolution of the N-FOV camera (represented schematically by the pixelation in FIG. 8B). However, the tracked player 601 remains within the output video. In some embodiments, an indicator such as the arrow 854 may be displayed on the screen 852 to indicate to the user which way to turn the device such that the tracked player 601 will again be within the N-FOV. The display of the video capture device may return to displaying a representation of the N-FOV camera video after the user reacquires the zoom region within the view of the N-FOV camera.
[0069] FIGs. 9A and 9B illustrate an embodiment under the same video recording conditions as those in FIGs. 7A and 7B. In the embodiment of FIGs. 9A and 9B, however, the device 950 displays the N-FOV video on the display 952 while the tracked object is within the N-FOV, and the device switches to display of the W-FOV video when the tracked object is not within the W-FOV. This allows the user to more easily
reorient the device such that the tracked object is again within the N-FOV. Initially, as illustrated in FIG. 7A, tracked player 601 is within the N-FOV, and high-quality zoomed video from the N-FOV camera is displayed on the screen 952 of the video recording device 950. Subsequently, the tracked player moves out of the N- FOV, and, in response, the device automatically switches to display of the W-FOV video on the screen. In this embodiment, the device further displays a rectangle 962 representing the N-FOV. The device may further display an indication, such as rectangle 702, to identify which object is being tracked. In other embodiments, the device may display an indication of the zoom region (such as the rectangle 602 of FIGs. 6B-6C). The device 602 may switch back to displaying the N-FOV video if the tracked object is reacquired within the N- FOV.
[0070] The foregoing examples have focused on embodiments using a device that has two cameras with different fields of view. In alternative embodiments, a device with more than two cameras may be used. For example, an object may be tracked to identify which cameras out of a plurality of cameras have that object within their respective fields of view. An output video may be generated by selecting which one of those cameras will be used to provide the output any particular time. In some such embodiments, the output is selected from the camera with the narrowest field of view that still includes the tracked object within its field of view. Video from cameras with a wider field of view may be cropped and upscaled as appropriate to match the field of view of cameras with a narrower field of view.
[0071] Although features and elements are described above in particular combinations, one of ordinary skill in the art will appreciate that each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, a read only memory (ROM), a random access memory (RAM), a register, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks, and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A method of providing an output video stream in a device having a first camera and a second camera, the second camera having a field of view that is narrower than and contained within a field of view of the first camera, the method comprising:
tracking an object in video captured by the first camera;
determining, based on the tracking, whether the object is entirely within the field of view of the second camera;
while the object is within the field of view of the second camera, outputting video captured by the second camera as the output video stream; and
while the object is outside the field of view of the second camera, (i) cropping and upscaling the video captured by the first camera to generate a cropped-and-upscaled video of the object and (ii) outputting the cropped-and-upscaled video as the output video stream.
2. The method of claim 1 , further comprising:
displaying, on a screen of the device, the video captured by the first camera; and
in the displayed video, displaying an indication of a position of the tracked object and an indication of the field of view of the second camera.
3. The method of claim 1 , wherein the method is performed in real time.
4. The method of claim 1 , further comprising, before tracking the object:
displaying, on a screen of the device, the video captured by the first camera; and
receiving user input selecting the object to be tracked in the displayed video.
5. The method of claim 1 , wherein outputting video captured by the second camera as the output video stream occurs while at least a threshold percentage of a bounding box around the object is within the field of view of the second camera.
6. The method of claim 1 , wherein the up-scaling is performed to match an optical zoom of the second camera.
7. The method of claim 1 , further comprising:
while the object is within the field of view of the second camera, displaying the video captured by the second camera on a screen of the device; and
while the object is outside the field of view of the second camera, displaying on the screen (i) the video captured by the first camera and (ii) an indication of the field of view of the second camera.
8. The method of 7 further comprising, while the object is outside the field of view of the second camera, displaying on the screen an indication of the position of the object.
9. The method of claim 1 , wherein outputting the output video stream comprises recording the output video stream.
10. The method of claim 1 , wherein outputting the output video stream comprises displaying the output video stream on a screen of the device.
11. A video capture device, comprising:
a first camera;
a second camera having a field of view that is narrower than and contained within a field of view of the first camera;
a processor; and
a non-transitory, computer-readable medium storing instructions that are operative, when executed on the processor, to perform the functions of:
tracking an object in video captured by the first camera;
determining, based on the tracking, whether the object is entirely within the field of view of the second camera;
while the object is within the field of view of the second camera, outputting video captured by the second camera as an output video stream; and
while the object is outside the field of view of the second camera, (i) cropping and upscaling the video captured by the first camera and (ii) outputting the cropped and upscaled video as the output video stream.
12. The device of claim 11 , further comprising:
a display; and
a user interface capable of receiving user selections of objects,
wherein the processor is further operative, before tracking the object, to perform the functions of: displaying, on the display, the video captured by the first camera; and
receiving user input selecting the object to be tracked in the displayed video.
13. The device of claim 11 , further comprising a Radio Frequency Identification (RFID) receiver, wherein the processor is further operative to track the object using the RFID receiver.
14. The device of claim 11 , further comprising a transceiver, wherein outputting video includes transmitting the output video stream to a remote device.
15. The device of claim 11 , wherein outputting the output video stream comprises recording the output video stream.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US201762468773P | 2017-03-08 | 2017-03-08 | |
| US62/468,773 | 2017-03-08 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018164932A1 true WO2018164932A1 (en) | 2018-09-13 |
Family
ID=61627200
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2018/020418 Ceased WO2018164932A1 (en) | 2017-03-08 | 2018-03-01 | Zoom coding using simultaneous and synchronous multiple-camera captures |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2018164932A1 (en) |
Cited By (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2020068580A1 (en) * | 2018-09-29 | 2020-04-02 | Apple Inc. | Devices, methods, and graphical user interfaces for depth-based annotation |
| US11727650B2 (en) | 2020-03-17 | 2023-08-15 | Apple Inc. | Systems, methods, and graphical user interfaces for displaying and manipulating virtual objects in augmented reality environments |
| US11797146B2 (en) | 2020-02-03 | 2023-10-24 | Apple Inc. | Systems, methods, and graphical user interfaces for annotating, measuring, and modeling environments |
| US11808562B2 (en) | 2018-05-07 | 2023-11-07 | Apple Inc. | Devices and methods for measuring using augmented reality |
| EP4300937A1 (en) * | 2022-06-29 | 2024-01-03 | Vestel Elektronik Sanayi ve Ticaret A.S. | Multi-image capture and storage |
| US11941764B2 (en) | 2021-04-18 | 2024-03-26 | Apple Inc. | Systems, methods, and graphical user interfaces for adding effects in augmented reality environments |
| US12020380B2 (en) | 2019-09-27 | 2024-06-25 | Apple Inc. | Systems, methods, and graphical user interfaces for modeling, measuring, and drawing using augmented reality |
| GB2632272A (en) * | 2023-07-28 | 2025-02-05 | Raytheon Systems Ltd | Systems and methods for collecting data from an observed field |
| US12395608B2 (en) | 2023-06-23 | 2025-08-19 | Adeia Guides Inc. | Systems and methods for enabling improved video conferencing |
| US12417594B2 (en) * | 2018-12-05 | 2025-09-16 | Tencent Technology (Shenzhen) Company Limited | Method for observing virtual environment, device, and storage medium |
| US12462498B2 (en) | 2021-04-18 | 2025-11-04 | Apple Inc. | Systems, methods, and graphical user interfaces for adding effects in augmented reality environments |
| US12469207B2 (en) | 2022-05-10 | 2025-11-11 | Apple Inc. | Systems, methods, and graphical user interfaces for scanning and modeling environments |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100238262A1 (en) * | 2009-03-23 | 2010-09-23 | Kurtz Andrew F | Automated videography systems |
| US20120300051A1 (en) * | 2011-05-27 | 2012-11-29 | Daigo Kenji | Imaging apparatus, and display method using the same |
| US20130230293A1 (en) * | 2012-03-02 | 2013-09-05 | H4 Engineering, Inc. | Multifunction automatic video recording device |
| US20150116501A1 (en) * | 2013-10-30 | 2015-04-30 | Sony Network Entertainment International Llc | System and method for tracking objects |
| US20160381289A1 (en) * | 2015-06-23 | 2016-12-29 | Samsung Electronics Co., Ltd. | Digital photographing apparatus and method of operating the same |
-
2018
- 2018-03-01 WO PCT/US2018/020418 patent/WO2018164932A1/en not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100238262A1 (en) * | 2009-03-23 | 2010-09-23 | Kurtz Andrew F | Automated videography systems |
| US20120300051A1 (en) * | 2011-05-27 | 2012-11-29 | Daigo Kenji | Imaging apparatus, and display method using the same |
| US20130230293A1 (en) * | 2012-03-02 | 2013-09-05 | H4 Engineering, Inc. | Multifunction automatic video recording device |
| US20150116501A1 (en) * | 2013-10-30 | 2015-04-30 | Sony Network Entertainment International Llc | System and method for tracking objects |
| US20160381289A1 (en) * | 2015-06-23 | 2016-12-29 | Samsung Electronics Co., Ltd. | Digital photographing apparatus and method of operating the same |
Cited By (23)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11808562B2 (en) | 2018-05-07 | 2023-11-07 | Apple Inc. | Devices and methods for measuring using augmented reality |
| US12174006B2 (en) | 2018-05-07 | 2024-12-24 | Apple Inc. | Devices and methods for measuring using augmented reality |
| WO2020068580A1 (en) * | 2018-09-29 | 2020-04-02 | Apple Inc. | Devices, methods, and graphical user interfaces for depth-based annotation |
| US11303812B2 (en) | 2018-09-29 | 2022-04-12 | Apple Inc. | Devices, methods, and graphical user interfaces for depth-based annotation |
| KR102522079B1 (en) * | 2018-09-29 | 2023-04-14 | 애플 인크. | Devices, methods, and graphical user interfaces for depth-based annotation |
| US11632600B2 (en) | 2018-09-29 | 2023-04-18 | Apple Inc. | Devices, methods, and graphical user interfaces for depth-based annotation |
| KR20210038619A (en) * | 2018-09-29 | 2021-04-07 | 애플 인크. | Devices, methods, and graphical user interfaces for depth-based annotation |
| US11818455B2 (en) | 2018-09-29 | 2023-11-14 | Apple Inc. | Devices, methods, and graphical user interfaces for depth-based annotation |
| US12131417B1 (en) | 2018-09-29 | 2024-10-29 | Apple Inc. | Devices, methods, and graphical user interfaces for depth-based annotation |
| US10785413B2 (en) | 2018-09-29 | 2020-09-22 | Apple Inc. | Devices, methods, and graphical user interfaces for depth-based annotation |
| US12417594B2 (en) * | 2018-12-05 | 2025-09-16 | Tencent Technology (Shenzhen) Company Limited | Method for observing virtual environment, device, and storage medium |
| US12406451B2 (en) | 2019-09-27 | 2025-09-02 | Apple Inc. | Systems, methods, and graphical user interfaces for modeling, measuring, and drawing using augmented reality |
| US12020380B2 (en) | 2019-09-27 | 2024-06-25 | Apple Inc. | Systems, methods, and graphical user interfaces for modeling, measuring, and drawing using augmented reality |
| US11797146B2 (en) | 2020-02-03 | 2023-10-24 | Apple Inc. | Systems, methods, and graphical user interfaces for annotating, measuring, and modeling environments |
| US12307067B2 (en) | 2020-02-03 | 2025-05-20 | Apple Inc. | Systems, methods, and graphical user interfaces for annotating, measuring, and modeling environments |
| US12592043B2 (en) | 2020-03-17 | 2026-03-31 | Apple Inc. | Systems, methods, and graphical user interfaces for displaying and manipulating virtual objects in augmented reality environments |
| US11727650B2 (en) | 2020-03-17 | 2023-08-15 | Apple Inc. | Systems, methods, and graphical user interfaces for displaying and manipulating virtual objects in augmented reality environments |
| US11941764B2 (en) | 2021-04-18 | 2024-03-26 | Apple Inc. | Systems, methods, and graphical user interfaces for adding effects in augmented reality environments |
| US12462498B2 (en) | 2021-04-18 | 2025-11-04 | Apple Inc. | Systems, methods, and graphical user interfaces for adding effects in augmented reality environments |
| US12469207B2 (en) | 2022-05-10 | 2025-11-11 | Apple Inc. | Systems, methods, and graphical user interfaces for scanning and modeling environments |
| EP4300937A1 (en) * | 2022-06-29 | 2024-01-03 | Vestel Elektronik Sanayi ve Ticaret A.S. | Multi-image capture and storage |
| US12395608B2 (en) | 2023-06-23 | 2025-08-19 | Adeia Guides Inc. | Systems and methods for enabling improved video conferencing |
| GB2632272A (en) * | 2023-07-28 | 2025-02-05 | Raytheon Systems Ltd | Systems and methods for collecting data from an observed field |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2018164932A1 (en) | Zoom coding using simultaneous and synchronous multiple-camera captures | |
| US12276899B2 (en) | Image pickup device and method of tracking subject thereof | |
| US9159169B2 (en) | Image display apparatus, imaging apparatus, image display method, control method for imaging apparatus, and program | |
| US11496666B2 (en) | Imaging apparatus with phase difference detecting element | |
| EP3054414B1 (en) | Image processing system, image generation apparatus, and image generation method | |
| CN108399349B (en) | Image recognition method and device | |
| US10334162B2 (en) | Video processing apparatus for generating panoramic video and method thereof | |
| US8350931B2 (en) | Arrangement and method relating to an image recording device | |
| JP6328255B2 (en) | Multi-imaging device, multi-imaging method, program, and recording medium | |
| US9269191B2 (en) | Server, client terminal, system and program for presenting landscapes | |
| US20190253747A1 (en) | Systems and methods for integrating and delivering objects of interest in video | |
| CN115147492B (en) | An image processing method and related equipment | |
| US10674066B2 (en) | Method for processing image and electronic apparatus therefor | |
| US11722762B2 (en) | Face tracking dual preview system | |
| CN101388981A (en) | Video image processing device and video image processing method | |
| JPWO2018003124A1 (en) | Imaging device, imaging method and imaging program | |
| JP2013038668A (en) | Display device, display method and program | |
| US20170034431A1 (en) | Method and system to assist a user to capture an image or video | |
| US20140210941A1 (en) | Image capture apparatus, image capture method, and image capture program | |
| US20210258505A1 (en) | Image processing apparatus, image processing method, and storage medium | |
| WO2017069902A1 (en) | Multiple camera autofocus synchronization | |
| WO2018222532A1 (en) | Rfid based zoom lens tracking of objects of interest | |
| US12464236B2 (en) | Imaging apparatus, focus control method, and focus control program | |
| WO2018164930A1 (en) | Improving motion vector accuracy by sharing cross-information between normal and zoom views | |
| US20110235856A1 (en) | Method and system for composing an image based on multiple captured images |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18710983 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18710983 Country of ref document: EP Kind code of ref document: A1 |