EP4690769A1 - Video conferencing device, system, and method using two-dimensional acoustic fence - Google Patents

Video conferencing device, system, and method using two-dimensional acoustic fence

Info

Publication number
EP4690769A1
EP4690769A1 EP23719164.8A EP23719164A EP4690769A1 EP 4690769 A1 EP4690769 A1 EP 4690769A1 EP 23719164 A EP23719164 A EP 23719164A EP 4690769 A1 EP4690769 A1 EP 4690769A1
Authority
EP
European Patent Office
Prior art keywords
sound
camera
image
distance
acoustic boundary
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23719164.8A
Other languages
German (de)
French (fr)
Inventor
Rajen Bhatt
Varun Kulkarni
Poojan Patel
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hewlett Packard Development Co LP
Original Assignee
Hewlett Packard Development Co LP
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hewlett Packard Development Co LP filed Critical Hewlett Packard Development Co LP
Publication of EP4690769A1 publication Critical patent/EP4690769A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N7/00Television systems
    • H04N7/14Systems for two-way working
    • H04N7/141Systems for two-way working between two video terminals, e.g. videophone
    • H04N7/147Communication arrangements, e.g. identifying the communication as a video-communication, intermediate storage of the signals

Definitions

  • An issue in videoconferencing is the intrusion of external noise sources, be it environmental noise or other individuals.
  • Various techniques provide what is called an acoustic fence around the video-conference area. Tn on variation, microphones are arranged in the form of a perimeter and used to detect background or far field noise which can be subtracted or used to mute or unmute the primary microphone audio. This technique uses multiple microphones located in various places and it is difficult to determine if the sound is inside or outside of the perimeter of the acoustically fenced area.
  • an acoustic fence is set to be within an angle of the centerline or an angle of the sensing microphone array. If the microphone array is located in the camera body, the centerlines of the camera and the microphone array can be matched. This results in an acoustic fence occurring for areas outside of an angle of the array centerline, which is an angle relating to the camera field of view.
  • the desired capture angle of the sound source localization can be varied manually. This technique operates in the one width (z.e., horizontal) dimension.
  • FIG. 1 is a top view of an example conference room including a camera with a microphone array, according to some aspects of the present disclosure.
  • FIG. 2 is a schematic perspective view of another example conference room with three individuals located at different coordinate positions in relation to a video conference camera.
  • FIG. 3 is a top view of the conference room of FIG. 2.
  • FIG. 4 is a table of pan angle values for a video conference camera corresponding to different coordinate positions in an example conference room.
  • FIG. 5 is top view of yet another example conference room, according to some aspects of the present disclosure.
  • FTG. 6 is a front view of still another example conference room with three individuals located at different coordinate positions, according to some examples of the present disclosure.
  • FIG. 7 is a top view of the conference room of FIG. 1 with one individual speaking from within an acoustic fence.
  • FIG. 8 is a top view of the conference room of FIGS. 1 and 7 with one individual speaking from outside of an acoustic fence.
  • FIG. 9 is a top view of the conference room of FIGS. 1, 7, and 8 with individuals speaking on either side of an acoustic fence.
  • FIG. 10 is a top view of the conference room of FIGS. 1 and 7-9 one individual speaking from within another acoustic fence, according to an example of the present disclosure.
  • FIG. 11 is a top view of the conference room of FIGS. 1 and 7-10 with one individual speaking from within yet another acoustic fence, according to an example of the present disclosure.
  • FIG. 12 is a side view of an example vertical dimension determination for a human head of a subject, according to an example of the present disclosure
  • FIG. 13 is a top view of an example horizontal dimension determination for a human head of a subject, according to an example of the present disclosure.
  • FIG. 14 is a perspective view of subjects in a conference room with different vertical dimensions based on the determinations of FIGS. 12 and 13.
  • FIG. 15 is a schematic diagram of a camera and a two-dimensional image plane with an example determination of room coordinates for a head bounding box, according to an example of the present disclosure.
  • FIG. 16 is a front view of a camera according to an example of the present disclosure.
  • FIG. 17 is a flowchart of a method of implementing a two-dimensional acoustic fence in a conference room, according to an example of the present disclosure.
  • FIG. 18 is a flowchart of a method of determining image plane coordinates for a detected subject, according to an example of the present disclosure.
  • FIG. 19 is a schematic of an example codec, according to an example of the present disclosure.
  • Videoconferencing systems typically connect people at a videoconferencing endpoint, such as a videoconference room, with people at other videoconferencing endpoints.
  • the framing of specific groups or individuals in the videoconference room can be improved by determining the location of individual participants in the room. For example, if Person A is sitting at 2.5 meters from the camera and Person B is sitting at 4 meters from the camera, the ability to detect this location information can enable various advanced framing and tracking experiences.
  • participant location information can be used to design a framing or bounding box that excludes people located more than a certain distance from the camera from framing and tracking.
  • a microphone array of a videoconference system When a microphone array of a videoconference system is used in a public place or a large conference room with two or more participants, background sounds, side conversations, or distracting noise may be present in the audio signal that the microphone array records and outputs to other participants in the videoconference. This is particularly true when the background sounds, side conversations, or distracting noises originate from within a field of view (FoV) of a camera used to record visual data for the videoconference system.
  • FoV field of view
  • the microphone array When the microphone array is being used to capture a user’s voice as audio for use in a teleconference, another participant or participants in the conference may hear the background sounds, side conversations, or distracting noise on their respective audio devices or speakers.
  • no industry standard or specification has been developed to reduce unwanted sounds in a videoconferencing system based on the distance from which the unwanted sounds are determined to originate from a videoconferencing camera.
  • the present disclosure provides a method of implementing a two-dimensional acoustic fence to remove or reduce the background sounds, side conversations, or distracting noises.
  • the communication between participants in the teleconference may be clearer, and the overall videoconferencing experience may be more enjoyable for the participants.
  • FIG. 1 illustrates a conference room 40 for use in videoconferencing.
  • the conference room 40 includes a conference table 42 and a series of chairs 44.
  • a videoconferencing device 46 includes a camera 48 and a microphone array 50 which are connected to a monitor 52 or television that is provided to display the far end conference site or sites and generally to provide the loudspeaker output.
  • the camera 48 is provided in the conference room 40 to view individuals (as shown in FIG. 7) seated in the various chairs 44, and the camera 48 has a FoV, horizontal and vertical, and an axis or centerline 56 extending in a direction that corresponds to the direction in which the camera 48 is pointing (z.e., the camera’s 48 line of sight that is straight at 90-degrees from its focal point).
  • the microphone array 50 is housed on or within a housing for the camera 48, and the microphone array 50 can be used to record and transmit audio data in the videoconference using sound source localization (SSL).
  • SSL is used in a way that is similar to the uses described in U.S. Patent App. Pub. No. 2023/0053202, which is incorporated herein by reference in its entirety.
  • the centerline 56 of the camera 48 is centered along the conference table 42.
  • a central microphone 60 is provided to capture the speaker (i.e., the person speaking) for transmission to a far end of the videoconference.
  • the camera 48 and the microphone array 50 are used in combination to provide a two-dimensional acoustic fence 62 or boundary so that audio signals originating from within the acoustic fence 62 are unmuted and transmitted to the far end of the videoconference via the microphone array 50. In this way, audio signals originating outside of the acoustic fence 62 are muted by the microphone array 50 This is accomplished using SSL and subject detection processes as described below.
  • a conference room 64 is illustrated with three videoconference participants 66, 68, 70 located at different coordinate positions.
  • the camera 48 has horizontal and vertical FoV and camera location with respect to 72 of the room is denoted by the three-dimensional (3D) coordinates ⁇ 0, 0, 0 ⁇ .
  • the camera 48 captures a view of all three participants 66, 68, 70 having locations that can be characterized in terms of a pan angle ⁇ bPAN and a distance measure between the camera 48 and each participant 66, 68, 70.
  • a first participant 66 has a location defined by a first pan angle 76 and a first distance 78.
  • a second participant 68 has a location defined by pan angle 80 and a second distance 82
  • a third participant 70 has a location defined by pan angle 84 and a third distance measure 86.
  • each participant 66, 68, 70 may be characterized in terms of the pan angles 76, 80, 84 and distances 78, 82, 86 that are derived from an XROOM dimension or axis 88 and yROOM dimension or axis 90.
  • the first participant 66 has a location defined by the first pan angle 76 and a first distance measure 78 which is characterized by two-dimensional room distance parameters ⁇ -0.5, 1 ⁇ to indicate that the participant is located at a vertical distance of 1 meter, measured from the camera 48 to along a ynooM axis 90, and at a horizontal distance of -0.5 meters, measured along an XROOM axis 88 that is perpendicular to the yROOM axis 90.
  • the second participant 68 has a location defined by the second pan angle 80 and a second distance measure 82 which is characterized by two-dimensional room distance parameters ⁇ 0, 3 ⁇ to indicate that the participant is located at a vertical distance of 3 meters (measured along the yROOM axis 90) and at a horizontal distance of 0 meters (measured along the XROOM axis 88) to indicate that the second person is located along the centerline 56 of the camera 48.
  • the third participant 70 has a location defined by the third pan angle 84 and a third distance measure 86 which is characterized by two-dimensional room distance parameters ⁇ 1, 2.5 ⁇ to indicate that the participant is located at a vertical distance of 2.5 meters (measured along the yROOM axis 90) and at a horizontal distance of 1 meter (measured along the X OOM axis 88).
  • a third distance measure 86 which is characterized by two-dimensional room distance parameters ⁇ 1, 2.5 ⁇ to indicate that the participant is located at a vertical distance of 2.5 meters (measured along the yROOM axis 90) and at a horizontal distance of 1 meter (measured along the X OOM axis 88).
  • pan angle OP AN values 94 for a video conference camera 48 are computed for meeting participants located at different coordinate positions ⁇ XROOM, YROOM ⁇ in the example conference room 64 of FIGS. 2 and 3.
  • An identical table (not shown) of negative pan angle OP AN values (e.g., -OP AN) would be computed for coordinate positions ⁇ - XROOM, YROOM ⁇ in the example conference room 64.
  • pan angle OP AN 45
  • the pan angle OP AN alone is not sufficient information for determining the two-dimensional ROOM distance parameters ⁇ XROOM, YROOM ⁇ for the location of an individual.
  • FIG. 5 depicts a top view of another example conference room 96.
  • the conference room 96 has a video conference camera 48 with meeting participants 98, 100, 102, 104, 106, 108 in different locations to illustrate how the perspective projection on the camera image sensor 110 changes to make an object appear smaller to the videoconferencing system as the object moves further from the camera 48.
  • any geometrical shape or object that is located along the centerline 56 of the camera 48 when the camera 48 is pointed at a far wall 112 of the conference room 96 is represented by straight horizontal and vertical edges (not shown).
  • the object appears to decrease in size as it moves further away from the camera 48 in the same perspective or viewing angle.
  • a depth dimension of the object measured from the camera 48 along the YROOM direction may be determined.
  • the height and width of the participant become smaller to the videoconferencing system, and when projected to camera image sensor 110, meeting participants are represented with a smaller number of pixels compared to participants that are nearer to the camera 48.
  • any meeting participant who is not determined to be located directly along the centerline 56 of the camera 48 (e.g., P AN 0) will create perspective projection on the camera image sensor 110.
  • a fifth meeting participant 106 who is not located along the centerline 56 of sight of camera 48 (e.g., pan angle P AN 0) will appear smaller than a third participant 102 that is located at the same vertical distance as the fifth meeting participant 106 measured along the yRooM axis 90.
  • a sixth meeting participant 108 who is not located along the centerline 56 of the camera 48 may appear to have the same size as the third meeting participant 102 even though they are located at different vertical distances as measured along the ynooM axis 90.
  • two heads are seen by the camera 48 has having the same size, they are not necessarily located at the same distance, and their locations in a two-dimensional XR00M-yR00M plane 114 may be different due to the pan angle OP AN and distortion in the height and width.
  • FIG. 6 illustrates a picture image 116 of another example conference room 118.
  • Three subj ects or meeting participants 120, 122, 124 are located in the conference room 118 at different coordinate positions and with corresponding head frames or bounding boxes 126, 128, 130 identified in terms of the coordinate positions for each of the participants 120, 122, 124.
  • the coordinate positions may be measured with reference to a room width dimension XROOM and a room depth dimension yRooM.
  • the room width dimension XROOM extends across a width of the conference room 118 from the centerline 56 (see FIG. 2) of the camera 48 (see FIG.
  • the room depth dimension y ooM extends down a length of the room 118 across the centerline 56 of the camera 48 (see FIG. 2).
  • the statistical distribution of human head height and width measurements may be used to determine a min-median-max measure for the head size in centimeters.
  • the measured angular extent of each head can be used to compute the percentage of the overall frame occupied by the head and the number of pixels for the head height and width measures.
  • an artificial intelligence (Al) human head detector model can be applied to detect the location of each head in a two-dimensional viewing plane with specified image plane coordinates and associated width and height measures for a head frame or bounding box (e.g., ⁇ xbox, ybox, width, height ⁇ ).
  • a head frame or bounding box e.g., ⁇ xbox, ybox, width, height ⁇ .
  • the present disclosure provides a method, device, system, and computer readable medium to accurately determining if a source of a sound originates within an acoustic fence 62 or acoustic boundary.
  • the location of meeting subject is determined using an Al human head detector model using room distance parameters, as discussed above.
  • the room distance parameters of human heads are then compared to room parameters that correspond to a two-dimensional acoustic fence 62. In this way, it becomes possible to determine if a particular sound recorded by the microphone array 50 has originated from within an area delimited by the acoustic fence 62 or from outside of the area delimited by the acoustic fence 62.
  • the microphone array 50 unmutes the sound and transmits the sound to the videoconference. However, if the sound is determined to have originated from outside of the acoustic fence 62, the microphone array 50 mutes the sound and does not transmit the sound to other participants in the videoconference.
  • FIG. 7 illustrates a top view of the conference room 40 of FIG. 1 with one person speaking within the acoustic fence 62.
  • four subjects or individuals 166, 168, 170, 172 are seated in the chairs 44 and two individuals 174, 176 are standing outside of a door 178 of the conference room 40.
  • individual 166 is speaking, as indicated by the shading of individual 166 and an SSL line 180 directed towards individual 166.
  • any individual detected within the FoV of the camera 48 who is speaking or otherwise creating sound will also be detected using SSL, and subsequent SSL lines will be used to detect the other individuals who are creating sound.
  • the FoV of the camera 48 is between 10 degrees and 120 degrees, or the FoV of the camera 48 is dependent on the camera 48.
  • the pan angle P AN i.e., the angle defined between the centerline 56 of the microphone array 50 and the SSL line
  • an Al human head detector process as described above in order to define spatial locations for each of the individuals 166, 168, 170, 172 in the FoV of the camera 48.
  • the Al human head detector model applies head bounding boxes 182, 184, 186, 188 to each detected individual 166, 168, 170, 172, and sizes of the bounding boxes 182, 184, 186, 188 are dependent upon an individual’s distance (e.g., distance 190) from the camera 48 and pan angle OP AN from the centerline 56 of the camera 48.
  • Individuals 174, 176 are not within the FoV of the camera 48 and are therefore not captured by the camera 48 or provided with bounding boxes.
  • the FoV of the camera 48 is wider than dimensions of the conference room 40, so portions of walls of the conference room 40 are also captured by the camera 48.
  • the camera 48 may have a wider or narrower FoV than illustrated in FIG. 7.
  • the acoustic fence 62 is created during a calibration phase in which a moderator or participant (not shown) draws the acoustic fence 62 or sets the boundaries of the acoustic fence 62 manually.
  • the videoconferencing system automatically detects dimensions for a room 40 and locations of meeting participants, and the videoconferencing system sets the acoustic fence 62 boundaries based on the detected room 40 dimensions and the locations of the meeting participants.
  • the acoustic fence 62 is arranged in a generally triangular configuration defined by a first side 192 extending from the camera 48 to a first wall 198 of the conference room 40, a second side 194 extending from the camera 48 to a second wall 200 of the conference room 40, and a base side 196 that extends between the first side 192 and the second side 194 in a direction that is perpendicular with respect to the centerline 56 of the camera 48.
  • the acoustic fence 62 defines a first distance 202 from the videoconferencing device 46 and a second distance 204 from the videoconferencing device 46 that is greater than the first distance 202. Sounds that are determined to have originated between the first distance 202 and the second distance 204 are unmuted. In this way, sounds that originate at distances that are less than the first distance 202 are muted, and sounds that originate at distances that are greater than the second distance 204 are muted.
  • the acoustic fence 62 defines a first pan angle 206 from the centerline 56 of the camera 48 and a second pan angle 208 from the centerline 56 of the camera 48. Sounds that are determined to have originated between the first pan angle 206 and the second pan angle 208 are unmuted. In other words, an origin point of the sound is determined based on the pan angle, the centerline, the room coordinates, and dimensions of the acoustic fence.
  • the acoustic fence 62 may be arranged in any suitable shape or configuration that is particularly desirable to videoconference meeting participants, such as configurations based on an environment of a room 40, or the acoustic fence 62 can be manually set by a user using room coordinates.
  • a moderator can change the boundaries of the acoustic fence 62 during a videoconference. Examples of different configurations in which the acoustic fence 62 can be arranged are described below with respect to FIGS. 10 and 11.
  • the comparison of the locations of the individuals with the boundaries of the acoustic fence 62 determines if a sound has originated within an area 210 delimited by the acoustic fence 62 or if a sound has originated from outside the area 210 delimited by the acoustic fence 62.
  • individuals 166, 170 are within the acoustic fence 62 while individuals 168, 172, 174, and 176 are outside of the acoustic fence 62. Since individual 170 (z.e., the active speaker as indicated by the SSL line) is within the acoustic fence 62, the microphone array 50 is unmuted and the sound created by individual 170 is transmitted to the videoconference. If individual 168 begins speaking, the microphone array 50 will also record and transmit the sound created by individual 168 since individual 168 is also within the acoustic fence 62.
  • FIG. 8 illustrates the conference room 40 of FIGS. 1 and 7 with one person speaking from outside of the acoustic fence 62.
  • individual 168 is speaking, and the videoconferencing device 46 has used SSL to locate the sound originating from individual 168 as indicated by the SSL line. Since individual 168 is outside of the area 210 delimited by the acoustic fence 62, the microphone array 50 mutes sound originating from individual 168 and does not transmit the sound originating from individual 168 to the videoconference. However, if individual 168 moves within the acoustic fence 62, the microphone array 50 will unmute sound that individual 168 creates while within the acoustic fence 62.
  • FIG. 9 illustrates the conference room 40 of FIGS. 1, 7, and 8 with multiple individuals speaking on either side of the acoustic fence 62.
  • individuals 168, 170, 174 are speaking, and the videoconferencing device 46 has used SSL to locate the sound originating from individuals 168, 170 as indicated by the SSL lines, 180a, 180b, respectively. Since individual 174 is located outside of the FoV of the camera 48, sound that is created by individual 174 is automatically muted and is not transmitted to the far end of the videoconference. Since individual 170 is within the area 210 delimited by the acoustic fence 62, the microphone array 50 is unmuted and the sound created by individual 170 is transmitted to the videoconference.
  • the microphone array 50 mutes sound originating from individual 168 and does not transmit the sound originating from individual 168 to the videoconference.
  • the videoconferencing device 46 is capable of identifying any number of individuals that are speaking within the FoV of the camera 48, but the microphone array 50 unmutes sound and transmits the sound to the videoconference if the sound is determined to have originated within the acoustic fence 62.
  • FIG. 10 illustrates the conference room 40 of FIGS. 1 and 7-9 with another acoustic fence 262 that is arranged as a rectangle.
  • individuals 168, 170 are located within the acoustic fence 262
  • individuals 166, 172 are located outside of the acoustic fence 262.
  • the microphone array 50 unmutes any speech or noise that is determined to have originated from individuals 168, 170, and the microphone array 50 mutes any speech or noise that is determined to have originated from individuals 166, 172.
  • muting is particularly useful for sounds that are closer to the videoconferencing device 46 than active participants of the videoconference , such as situations in which a prompter or moderator is located toward a front 264 of the room 40 while another participant is actively presenting or speaking. Any noise created by the moderator is muted and thus is not transmitted to the videoconference.
  • FIG. 11 illustrates the conference room 40 of FIGS. 1 and 7-10 with one individual speaking from within another acoustic fence 362.
  • the acoustic fence 362 may be arranged in any suitable shape or configuration that is particularly desirable to videoconference meeting participants. In some examples, more than one acoustic fence 362 is used in order to delimit particular areas of the videoconferencing room 40.
  • a first acoustic fence 364 is arranged as a rectangle around individual 170
  • a second acoustic fence 366 is arranged as a rectangle around individual 172.
  • the microphone array 50 unmutes noise created by individual 170 and/or individual 172 and the noise is transmitted to the videoconference.
  • this arrangement is useful in other environments in addition to conference rooms or enclosed rooms, such as public places or open concept workspaces, and hybrid workspaces.
  • the two-dimensional acoustic fence 362 can be used to isolate individual participants in a videoconference to prevent unwanted background noise from being recorded by the microphone array 50 and being transmitted to other participants in the videoconference.
  • a side view is illustrated of an example vertical dimension determination 400 for a human head 402.
  • a camera 48 is positioned to capture an image of the human head 402 so that a vertical head height measure V that can be calculated based on an angular extent angle 0FRAME V/2 of the upper half of a vertical head height V/2 and a distance d between the camera 48 and head 402.
  • the human head 402 has a head height V which corresponds to the vertical dimension of a head bounding box (not shown).
  • the vertical head height V makes an angle 0FRAME V extending from the bottom to the top of the head 402.
  • the upper half of the vertical head height V/2 makes an angle 0FRAME V/2 with the camera’s focal point 58 (i.e., the centerline 56 of the camera).
  • FIG. 13 illustrates of an example horizontal dimension determination 406 for a human head 408.
  • a camera 48 is positioned to capture an image of the human head 408 so that a horizonal head width measure H that can be calculated based on an angular extent angle 0FRAME H/2 of one half of the horizontal head width H/2 and the distance d between the camera 48 and head 408.
  • the human head 408 has a head width H which corresponds to the horizonal dimension of a head bounding box. From the vantage of the camera 48, the horizontal head width H makes an angle 0FRAME H extending from the sides of the head 408.
  • FIG. 14 illustrates a camera 48 and a two-dimensional image plane 410.
  • a video conference camera 48 is used to provide an image of a meeting participant located in a first, centered position 412 and a second, panned position 414 that is shifted laterally in the XROOM direction.
  • the fact that the second, panned position 414 is located further away from the camera 48 than the first, centered position 412 results in the angular extent for the second, panned position 414 appearing to be smaller than the angular extent for the first, centered position 412 so that 0FRAME Vl/2 > 0FRAME V2/2.
  • the issue is to find an angular extent for the entire head height HH and then represent it as a percentage of the full frame vertical field of view (VFrame_Percentage) which is then translated into the number of pixels the head will occupy (VHead Pixel Count) at a particular distance and at a pan angle P AN.
  • VHead Pixel Count VFrame Percentage x Vertical FoV in pixels.
  • FIG. 15 illustrates a camera 48 and a two-dimensional image plane 510 to illustrate how to calculate a vertical or depth room distance YROOM (meters) to the meeting participant location from the distance measure XROOM (meters) by calculating a direct distance measure HYP between the camera 48 and the meeting participant location.
  • the two-dimensional image plane 510 includes a plurality of two-dimensional coordinate points 512, 514, 516 that are defined with image plane 510 coordinates ⁇ xi, y as described above.
  • a head bounding box 518 is defined with reference to the starting coordinate point ⁇ xi, yi ⁇ for the head bounding box 518, a Width dimension (measured along the xi axis), and a Height dimension (measured along the yi Tha).
  • HYP V_HEAD/(2 x tan(0/2)), where HYP is the direct distance measure to the meeting participant location at the pan angle OP AN.
  • HYP V_HEAD/(2 x tan(0/2)
  • HYP the direct distance measure to the meeting participant location at the pan angle OP AN.
  • FIG. 16 illustrates of an example camera 548 and microphone array 550, similar to the camera 48 and the microphone array 50, respectively.
  • the camera 548 has a housing 552 with a lens 554 provided in the center to operate with the imager 556.
  • a series of five openings 558 are provided as ports to microphones in the microphone array 550.
  • the microphone openings 558 form a horizontal line 560 to provide the desired angular determination for the SSL process.
  • aspects of the technology can be implemented as a system, method, apparatus, or article of manufacture using standard programming or engineering techniques to produce software, firmware, hardware, or any combination thereof to control a processor device (e.g., a serial or parallel general purpose or specialized processor chip, a single- or multi-core chip, a microprocessor, a field programmable gate array, any variety of combinations of a control unit, arithmetic logic unit, and processor register, and so on), a computer (e.g., a processor device operatively coupled to a memory), or another electronically operated controller to implement aspects detailed herein.
  • a processor device e.g., a serial or parallel general purpose or specialized processor chip, a single- or multi-core chip, a microprocessor, a field programmable gate array, any variety of combinations of a control unit, arithmetic logic unit, and processor register, and so on
  • a computer e.g., a processor device operatively coupled to a memory
  • another electronically operated controller to implement
  • the technology can be implemented as a set of instructions, tangibly embodied on a non- transitory computer-readable media, such that a processor device can implement the instructions based upon reading the instructions from the computer-readable media.
  • a control device such as, e.g., an automation device, a special purpose or general-purpose computer including various computer hardware, software, firmware, and so on, consistent with the discussion below.
  • a control device can include a processor, a microcontroller, a field-programmable gate array, a programmable logic controller, logic gates etc., and other suitable components for implementation of appropriate functionality e.g., memory, communication systems, power sources, user interfaces and other inputs, etc.).
  • FIG. 17 illustrates a method 600 of implementing a two-dimensional acoustic fence in a conference room.
  • step 602 it is determined if sound is present in the output of the microphone. If not, operation proceeds to step 616 to mute the microphone, and operation then returns to step 602.
  • step 604 SSL is used to determine the angle of the sound from the centerline of the microphone array by analyzing the audio output signals of the microphone array (i.e., the microphone output), which centerline is preferably aligned parallel with a line normal to the center of the camera lens.
  • SSL can be performed as disclosed in U.S. Pat No. 6,912,178, which is hereby incorporated herein by reference in its entirety, or by other desired methods.
  • another step (not shown) of determining the FoV angle and centerline angle of the camera are determined. If the camera performs physical pan, tilt and zoom (MTZ), the centerline angle and FoV are determined based on the pan angle and the zoom amount. If the camera performs electronic MZ (EPTZ), the centerline angle is determined by the number of pixels the center of the image or frame is displaced from the center of the camera full image and the number of pixels in the full width of the camera. The FoV is determined by the ratio of the width of the image to the width of the full image and applying that ratio to the FoV of the camera.
  • MZ physical pan, tilt and zoom
  • EPTZ electronic MZ
  • the FoV is determined by the ratio of the width of the image to the width of the full image and applying that ratio to the FoV of the camera.
  • an Al head detection process is used to detect human heads in the conference room.
  • the Al head detection process uses the sound angle and is applied to images captured by the camera in order to identify, for each detected human head, a head bounding box with specified room coordinates and dimension information which are then used to calculate horizontal pan distance and depth dimension distance measured from the camera from a top-down perspective of the room. In this way, a two-dimensional room coordinate location for each detected human head is determined.
  • step 610 the location of each detected human head is compared to a location of an acoustic fence that is created during a calibration phase or inputted by the user/administrator during room or video conferencing set-up step.
  • the acoustic fence can be drawn by a moderator, or the boundaries of the acoustic fence can be manually input to the videoconferencing system using image coordinates.
  • step 612 the room coordinates and dimension information for each detected human head are checked against the boundaries of the acoustic fence, based on the centerline angle, the camera FoV, and the depth dimension measured from the camera. If the centerline of the microphone array and the camera are aligned, this is a simple comparison.
  • the angle between the two centerlines is determined and used to adjust the determined angle of the sound to be based on the camera centerline. If the sound is determined to originate from within the acoustic fence, the microphone is unmuted in step 614. If the sound is determined not to have originated from within the acoustic fence, the microphone is muted in step 616. Operation returns to step 602 so that the muting and unmuting are automatic as detection of sound is performed.
  • the microphone is unmuted if sound is found to have originated from within the acoustic fence.
  • a further step is included to determine if the sound is speech before unmuting the microphone. This keeps the microphone muted for just noise sources, such as fans or other environmental noise, when there is no speech.
  • FIG. 18 illustrates a process to determine image plane coordinates for a detected human head using an Al human head detector process.
  • the Al human head detector process analyzes incoming room-view video frame images 702 of a meeting room scene with a head detector machine learning model 704 to detect and display human heads with corresponding head bounding boxes 706, 708, 710. As depicted, each incoming room-view video frame image 702 may be captured by a camera 48 in the video conferencing system.
  • a first view of the meeting participants is captured by a first camera (not shown) in a first profile image or video frame 702a
  • a second camera (not shown) captures a second profde image or video frame 702b
  • a third camera (not shown) captures a third profde image or video frame 702c.
  • Each incoming room-view video frame image 702 may be processed with an on-device Al human head detector model 704 that may be located at the respective camera which captures the video frame images.
  • the Al human head detector model 704 may be located at a remote or centralized location, or a single camera is used.
  • the Al human head detector model 704 may include a plurality of processing modules 712, 714, 716, 718 which implement a machine learning model which is trained to detect or classify human heads from the incoming video frame images, and to identify, for each detected human head, a head bounding box with specified image plane coordinate and dimension information.
  • the Al human head detector model 704 may include a first preprocessing module 712 that applies image pre-processing (such as color conversion, image scaling, image enhancement, image resizing, etc.) so that the input video frame image is prepared for subsequent Al processing.
  • a second module 714 may include training data parameters or model architecture definitions which may be pre-defined and used to train and define the human head detection model 716 to accurately detect or classify human heads from the incoming video frame images.
  • the human head detection model 716 may be implemented as a model inference software or machine learning model, such as a Convolutional Neural Network (CNN) model that is specially trained for video codec operations to detect heads in an input image by generating pixel-wise locations for each detected head and by generating, for each detected head, a corresponding head bounding box which frames the detected head.
  • the Al human head detector model 704 may include a post-processing module 718 which is applies image postprocessing to the output from the Al human head detector model 704 to make the processed images suitable for human viewing and understanding.
  • the post-processing module 718 may also reduce the size of the data outputs generated by the human head detection model 716, such as by consolidating or grouping a plurality of head bounding boxes or frames which are generated from a single meeting participant so that a single head bounding box or frame is specified.
  • the Al human head detector model 704 may generate output video frame images 702 in which the detected human heads are framed with corresponding head bounding boxes 706, 708, 710.
  • the first output video frame image 702a includes head bounding boxes 706a-c which are superimposed around each detected human head.
  • the second output video frame image 702b includes head bounding boxes 708a-c which are superimposed around each detected human head
  • the third output video frame image 702c includes head bounding boxes 710a, 710b which are superimposed around each detected human head.
  • the Al human head detector model 704 may specify each head bounding box using any suitable pixel-based parameters, such as defining the x and y pixel coordinates of a head bounding box or frame in combination with the height and width dimensions of the head bounding box or frame.
  • the Al human head detector model 704 may specify a distance measure between the camera location and the location of the detected human head using any suitable measurement technique.
  • the Al human head detector model 704 may also compute, for each head bounding box, a corresponding confidence measure or score which quantifies the model’s confidence that a human head is detected.
  • the Al human head detector model 704 may specify all head detections in the data structure that holds the coordinates of each detected human head along with their detection confidence.
  • the human head data structure for a number, n, of human heads may be generated as follows:
  • xi and yi refer to the image plane coordinates of the i 111 detected head
  • Widthi and Height refer to the width and height information for the head bounding box of the i th detected head.
  • Scorei is in the range (0, 100] and reflect confidence in percentage for the i th detected head.
  • This data structure may be used as an input to various applications, such as framing, tracking, composing, recording, switching, reporting, encoding, etc.
  • the first detected head is in the image frame in a head bounding box located at pixel location parameters xi, yi and extending laterally by Widthi and vertically down by Heighti.
  • the second detected head is in the image frame in a head bounding box located at pixel location parameters X2, yi and extending laterally by Width2 and vertically down by Height2, and the n th detected head is in the image frame in a head bounding box located at pixel location parameters xn, y n and extending laterally by Widthn and vertically down by Heightn.
  • This human head data structure may then be used as an input to the distance estimation process that takes the ⁇ Width, Height ⁇ parameters of each head bounding box to pick the best matching distance in terms of meeting room coordinates ⁇ XROOM, yROOM ⁇ from the look-up table by first using one of the Width or Height parameters with a first lookup table, and then using the other parameter as a tie breaking if multiple meeting room coordinates ⁇ XROOM, yROOM) are determined by the one.
  • the human head data structure itself may then be modified to also embed the distance information with each Head, resulting in a modified human head data structure that looks like the following: where ⁇ XROOMI, yROOMi ⁇ , ⁇ XROOM2, yROOM2 ⁇ , ... , ⁇ XROOMU, yROOMn ⁇ specify the distance of Headi, Head2, . . ., Headn, from the camera, respective, in two-dimensional coordinates.
  • FIG. 19 illustrates aspects of a codec according to some examples of the present disclosure.
  • the codec 800 may include loudspeaker ⁇ s) 802, though in many cases the loudspeaker 802 is provided in the monitor 804.
  • the codec 800 may include microphone(s) 806 interfaced via a bus 808.
  • the microphones 806 are connected through an analog to digital (AID) converter 810, and the loudspeaker 802 is connected through a digital to analog (D/A) converter 812.
  • the codec 800 also includes a processing unit 814, a network interface 816, a flash or other non-transitory memory 818, RAM 820, and an input/output (I/O) general interface 822, all coupled by a bus 808.
  • I/O input/output
  • a camera 824 is connected to the I/O interface 822.
  • Microphone(s) 806 are connected to the network interface 816.
  • An HDMI interface 826 is connected to the bus 808 and to the external display or monitor 804.
  • Bus 808 is illustrative and any interconnect between the elements can used, such as Peripheral Compo-nent Interconnect Express (PCie) links and switches, Universal Serial Bus (USB) links and hubs, and combinations thereof.
  • PCie Peripheral Compo-nent Interconnect Express
  • USB Universal Serial Bus
  • the camera 824 and microphones 806, 806 can be contained in housings containing the other components or can be external and removable, connected by wired or wireless connections.
  • the processing unit 814 can include digital signal processors (DSPs), central processing units (CPUs), graphics processing units (GPUs), dedicated hardware elements, such as neural network accelerators and hardware codecs.
  • the flash memory 818 stores modules of varying functionality in the form of software and firmware, generically programs, for controlling the codec 800. Illustrated modules include a video codec 828, camera control 830, framing 832, other video processing 834, audio codec 836, audio processing 838, network operations 840, user interface 842 and operating system, and various other modules 844. In some examples, an Al head detector module is included the modules included in the flash memory 818. Additionally, at least some of the operations of FIG. 17 are performed in the audio processing 838, and the muting and unmuting operations of FIG. 17 are performed in the audio processing 838.
  • the RAM 820 is used for storing any of the modules in the flash memory 818 when the module is executing, storing video images of video streams and audio samples of audio streams and can be used for scratchpad operation of the processing unit 814.
  • the network interface 816 enables communications between the codec 800 and other devices and can be wired, wireless or a combination.
  • the network interface 816 is connected or coupled to the Internet 846 to communicate with remote endpoints 848 in a videoconference Tn
  • the general interface 822 provides data transmission with local devices (not shown) such as a keyboard, mouse, printer, projector, display, exter-nal loudspeakers, additional cameras, and microphone pods, etc.
  • the camera 824 and the microphones 806 capture video and audio, respectively, in the videoconference environment and produce video and audio streams or signals transmitted through the bus 808 to the processing unit 814.
  • the processing unit 814 processes the video and audio using processes in the modules stored in the flash memory 818. Processed audio and video streams can be sent to and received from remote devices coupled to network interface 816 and devices coupled to general interface 822.
  • Microphones in the microphone array used for SSL can be used as the microphones providing speech to the far site, or separate microphones, such as microphone 806, can be used.
  • Certain operations of methods according to the technology, or of systems executing those methods can be represented schematically in the figures or otherwise discussed herein. Unless otherwise specified or limited, representation in the figures of particular operations in particular spatial order can not necessarily require those operations to be executed in a particular sequence corresponding to the particular spatial order.
  • certain operations represented in the figures, or otherwise disclosed herein can be executed in different orders than are expressly illustrated or described, as appropriate for particular examples of the technology. Further, in some examples, certain operations can be executed in parallel, including by dedicated parallel processing devices, or separate computing devices that interoperate as part of a large system.
  • a plurality of hardware and software-based devices, as well as a plurality of different structural components can be used to implement the disclosed technology.
  • examples of the disclosed technology can include hardware, software, and electronic components or modules that, for purposes of discussion, can be illustrated and described as if the majority of the components were implemented solely in hardware.
  • the electronic based aspects of the disclosed technology can be implemented in software (for example, stored on non- transitory computer-readable medium) executable by a processor.
  • certain drawings illustrate hardware and software located within particular devices, these depictions are for illustrative purposes. In some examples, the illustrated components can be combined or divided into separate software, firmware, hardware, or combinations thereof.
  • any suitable non-transitory computer usable or computer readable medium may be utilized.
  • the computer-usable or computer-readable medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device.
  • a computer-usable or computer-readable medium may be any medium that can contain, store, communicate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
  • a component can be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer.
  • a component can be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer.
  • an application running on a computer and the computer can be a component.
  • Components or system, module, and so on

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)

Abstract

Embodiments of the disclosure provide a device, system, and method using a two-dimensional acoustic fence (62). The system includes a microphone (50) that determines if a sound is present in a room (40) and a camera (48) that captures an image of the room. A subject detector model is applied to the image to determine image coordinates and room coordinates in width and depth dimensions from the camera for each subject in the room. An acoustic fence is defined within the image. An artificial intelligence process compares the room coordinates of the subjects with the acoustic fence dimensions to determine if the sound from the subjects has originated within the acoustic fence. The audio output for the videoconference is muted if the sound is determined not to have originated within the acoustic fence, and the audio output is unmuted if the sound is determined to have originated within the acoustic fence.

Description

VIDEO CONFERENCING DEVICE, SYSTEM, AND METHOD USING TWO- DIMENSIONAL ACOUSTIC FENCE
BACKGROUND
[0001] An issue in videoconferencing is the intrusion of external noise sources, be it environmental noise or other individuals. Various techniques provide what is called an acoustic fence around the video-conference area. Tn on variation, microphones are arranged in the form of a perimeter and used to detect background or far field noise which can be subtracted or used to mute or unmute the primary microphone audio. This technique uses multiple microphones located in various places and it is difficult to determine if the sound is inside or outside of the perimeter of the acoustically fenced area.
[0002] In another variation, an acoustic fence is set to be within an angle of the centerline or an angle of the sensing microphone array. If the microphone array is located in the camera body, the centerlines of the camera and the microphone array can be matched. This results in an acoustic fence occurring for areas outside of an angle of the array centerline, which is an angle relating to the camera field of view. The desired capture angle of the sound source localization can be varied manually. This technique operates in the one width (z.e., horizontal) dimension.
BRIEF DESCRIPTION OF THE DRAWINGS
[0003] FIG. 1 is a top view of an example conference room including a camera with a microphone array, according to some aspects of the present disclosure.
[0004] FIG. 2 is a schematic perspective view of another example conference room with three individuals located at different coordinate positions in relation to a video conference camera.
[0005] FIG. 3 is a top view of the conference room of FIG. 2.
[0006] FIG. 4 is a table of pan angle values for a video conference camera corresponding to different coordinate positions in an example conference room.
[0007] FIG. 5 is top view of yet another example conference room, according to some aspects of the present disclosure. [0008] FTG. 6 is a front view of still another example conference room with three individuals located at different coordinate positions, according to some examples of the present disclosure.
[0009] FIG. 7 is a top view of the conference room of FIG. 1 with one individual speaking from within an acoustic fence.
[0010] FIG. 8 is a top view of the conference room of FIGS. 1 and 7 with one individual speaking from outside of an acoustic fence.
[0011] FIG. 9 is a top view of the conference room of FIGS. 1, 7, and 8 with individuals speaking on either side of an acoustic fence.
[0012] FIG. 10 is a top view of the conference room of FIGS. 1 and 7-9 one individual speaking from within another acoustic fence, according to an example of the present disclosure.
[0013] FIG. 11 is a top view of the conference room of FIGS. 1 and 7-10 with one individual speaking from within yet another acoustic fence, according to an example of the present disclosure.
[0014] FIG. 12 is a side view of an example vertical dimension determination for a human head of a subject, according to an example of the present disclosure
[0015] FIG. 13 is a top view of an example horizontal dimension determination for a human head of a subject, according to an example of the present disclosure.
[0016] FIG. 14 is a perspective view of subjects in a conference room with different vertical dimensions based on the determinations of FIGS. 12 and 13.
[0017] FIG. 15 is a schematic diagram of a camera and a two-dimensional image plane with an example determination of room coordinates for a head bounding box, according to an example of the present disclosure.
[0018] FIG. 16 is a front view of a camera according to an example of the present disclosure.
[0019] FIG. 17 is a flowchart of a method of implementing a two-dimensional acoustic fence in a conference room, according to an example of the present disclosure. [0020] FIG. 18 is a flowchart of a method of determining image plane coordinates for a detected subject, according to an example of the present disclosure.
[0021] FIG. 19 is a schematic of an example codec, according to an example of the present disclosure.
DETAILED DESCRIPTION
[0022] Videoconferencing systems typically connect people at a videoconferencing endpoint, such as a videoconference room, with people at other videoconferencing endpoints. In such systems, the framing of specific groups or individuals in the videoconference room can be improved by determining the location of individual participants in the room. For example, if Person A is sitting at 2.5 meters from the camera and Person B is sitting at 4 meters from the camera, the ability to detect this location information can enable various advanced framing and tracking experiences. For example, participant location information can be used to design a framing or bounding box that excludes people located more than a certain distance from the camera from framing and tracking.
[0023] When a microphone array of a videoconference system is used in a public place or a large conference room with two or more participants, background sounds, side conversations, or distracting noise may be present in the audio signal that the microphone array records and outputs to other participants in the videoconference. This is particularly true when the background sounds, side conversations, or distracting noises originate from within a field of view (FoV) of a camera used to record visual data for the videoconference system. When the microphone array is being used to capture a user’s voice as audio for use in a teleconference, another participant or participants in the conference may hear the background sounds, side conversations, or distracting noise on their respective audio devices or speakers. Further, no industry standard or specification has been developed to reduce unwanted sounds in a videoconferencing system based on the distance from which the unwanted sounds are determined to originate from a videoconferencing camera.
[0024] For many applications, it is useful to know the horizontal and vertical location of the humans in the room to provide for a more comprehensive and complete understanding of the videoconference room environments. The ability to determine two-dimensional room distance parameters for each meeting participant can be enabled by using a depth estimation/detection sensor or computationally intensive machine learning-based monocular depth estimation models, but such approaches impose significant hardware and/or processing costs without providing the accuracy for measuring participant locations.
[0025] Accordingly, in some examples, the present disclosure provides a method of implementing a two-dimensional acoustic fence to remove or reduce the background sounds, side conversations, or distracting noises. By utilizing the two-dimensional acoustic fence method, the communication between participants in the teleconference may be clearer, and the overall videoconferencing experience may be more enjoyable for the participants.
[0026] FIG. 1 illustrates a conference room 40 for use in videoconferencing. The conference room 40 includes a conference table 42 and a series of chairs 44. A videoconferencing device 46 includes a camera 48 and a microphone array 50 which are connected to a monitor 52 or television that is provided to display the far end conference site or sites and generally to provide the loudspeaker output. The camera 48 is provided in the conference room 40 to view individuals (as shown in FIG. 7) seated in the various chairs 44, and the camera 48 has a FoV, horizontal and vertical, and an axis or centerline 56 extending in a direction that corresponds to the direction in which the camera 48 is pointing (z.e., the camera’s 48 line of sight that is straight at 90-degrees from its focal point). In some examples, the microphone array 50 is housed on or within a housing for the camera 48, and the microphone array 50 can be used to record and transmit audio data in the videoconference using sound source localization (SSL). In some examples, SSL is used in a way that is similar to the uses described in U.S. Patent App. Pub. No. 2023/0053202, which is incorporated herein by reference in its entirety. In the layout of FIG. I, the centerline 56 of the camera 48 is centered along the conference table 42. In some examples, a central microphone 60 is provided to capture the speaker (i.e., the person speaking) for transmission to a far end of the videoconference. Additionally, the camera 48 and the microphone array 50 are used in combination to provide a two-dimensional acoustic fence 62 or boundary so that audio signals originating from within the acoustic fence 62 are unmuted and transmitted to the far end of the videoconference via the microphone array 50. In this way, audio signals originating outside of the acoustic fence 62 are muted by the microphone array 50 This is accomplished using SSL and subject detection processes as described below.
[0027] Referring now to FIGS. 2 and 3, a conference room 64 is illustrated with three videoconference participants 66, 68, 70 located at different coordinate positions. In the conference room 64, the camera 48 has horizontal and vertical FoV and camera location with respect to 72 of the room is denoted by the three-dimensional (3D) coordinates {0, 0, 0}. Further, the camera 48 captures a view of all three participants 66, 68, 70 having locations that can be characterized in terms of a pan angle <bPAN and a distance measure between the camera 48 and each participant 66, 68, 70. In particular, a first participant 66 has a location defined by a first pan angle 76 and a first distance 78. In addition, a second participant 68 has a location defined by pan angle 80 and a second distance 82, and a third participant 70 has a location defined by pan angle 84 and a third distance measure 86.
[0028] Referring now specifically to FIG. 3, a top view is illustrated of the conference room 64 of FIG. 2. In some examples, the location of each participant 66, 68, 70 may be characterized in terms of the pan angles 76, 80, 84 and distances 78, 82, 86 that are derived from an XROOM dimension or axis 88 and yROOM dimension or axis 90. In particular, the first participant 66 has a location defined by the first pan angle 76 and a first distance measure 78 which is characterized by two-dimensional room distance parameters {-0.5, 1 } to indicate that the participant is located at a vertical distance of 1 meter, measured from the camera 48 to along a ynooM axis 90, and at a horizontal distance of -0.5 meters, measured along an XROOM axis 88 that is perpendicular to the yROOM axis 90. In addition, the second participant 68 has a location defined by the second pan angle 80 and a second distance measure 82 which is characterized by two-dimensional room distance parameters {0, 3} to indicate that the participant is located at a vertical distance of 3 meters (measured along the yROOM axis 90) and at a horizontal distance of 0 meters (measured along the XROOM axis 88) to indicate that the second person is located along the centerline 56 of the camera 48. Finally, the third participant 70 has a location defined by the third pan angle 84 and a third distance measure 86 which is characterized by two-dimensional room distance parameters { 1, 2.5} to indicate that the participant is located at a vertical distance of 2.5 meters (measured along the yROOM axis 90) and at a horizontal distance of 1 meter (measured along the X OOM axis 88). [0029] To demonstrate the relationship between the pan angle values (OP AN) and the two- dimensional room distance parameters {XROOM, YROOM}, reference is now made to FIG. 4 which depicts a reference coordinate table 92 in which pan angle OP AN values 94 for a video conference camera 48 are computed for meeting participants located at different coordinate positions {XROOM, YROOM} in the example conference room 64 of FIGS. 2 and 3. An identical table (not shown) of negative pan angle OP AN values (e.g., -OP AN) would be computed for coordinate positions {- XROOM, YROOM} in the example conference room 64. As depicted, the same pan angle OP AN value ((e. ., OP AN = 0) will be generated for a meeting participant located along the centerline 56 of the camera 48 (e.g., XROOM = 0) at any depth measure (e.g., YROOM = 0.5-8). Similarly, the same pan angle OP AN value (e.g., OP AN = 45) will be generated for a meeting participant located at any coordinate position where XROOM = YROOM. AS illustrated, the pan angle OP AN alone is not sufficient information for determining the two-dimensional ROOM distance parameters {XROOM, YROOM} for the location of an individual.
[0030] This effect is illustrated in FIG. 5 which depicts a top view of another example conference room 96. The conference room 96 has a video conference camera 48 with meeting participants 98, 100, 102, 104, 106, 108 in different locations to illustrate how the perspective projection on the camera image sensor 110 changes to make an object appear smaller to the videoconferencing system as the object moves further from the camera 48. In particular, any geometrical shape or object that is located along the centerline 56 of the camera 48 when the camera 48 is pointed at a far wall 112 of the conference room 96 is represented by straight horizontal and vertical edges (not shown). However, the object appears to decrease in size as it moves further away from the camera 48 in the same perspective or viewing angle. In this way, a depth dimension of the object measured from the camera 48 along the YROOM direction may be determined. For example, a first meeting participant 98 located on the centerline 56 of the camera 48 (e.g., P AN = 0) at a distance of d = 0.5 meters will appear larger to the camera 48 than a second meeting participant 100 located on the centerline 56 of the camera 48 (e.g., OP AN = 0) at a larger distance of d = 1.0 meters due to vanishing points perspective. Similarly, the meeting participants 98, 100, 102, 104 located along the centerline 56 of the camera 48 will appear to decrease in size at larger distances (e.g., d = 1.50 meters and d = 2.0 meters). Thus, as a meeting participant moves further away from the camera 48, the height and width of the participant become smaller to the videoconferencing system, and when projected to camera image sensor 110, meeting participants are represented with a smaller number of pixels compared to participants that are nearer to the camera 48.
[0031] In addition, any meeting participant who is not determined to be located directly along the centerline 56 of the camera 48 (e.g., P AN 0) will create perspective projection on the camera image sensor 110. For example, a fifth meeting participant 106 who is not located along the centerline 56 of sight of camera 48 (e.g., pan angle P AN 0) will appear smaller than a third participant 102 that is located at the same vertical distance as the fifth meeting participant 106 measured along the yRooM axis 90. Likewise, a sixth meeting participant 108 who is not located along the centerline 56 of the camera 48 (e.g., pan angle OP AN 0) may appear to have the same size as the third meeting participant 102 even though they are located at different vertical distances as measured along the ynooM axis 90. As a result, if two heads are seen by the camera 48 has having the same size, they are not necessarily located at the same distance, and their locations in a two-dimensional XR00M-yR00M plane 114 may be different due to the pan angle OP AN and distortion in the height and width.
[0032] To illustrate an example of the challenges posed by perspective projection effects when determining the locations of meeting participants, reference is now made to FIG. 6 which illustrates a picture image 116 of another example conference room 118. Three subj ects or meeting participants 120, 122, 124 are located in the conference room 118 at different coordinate positions and with corresponding head frames or bounding boxes 126, 128, 130 identified in terms of the coordinate positions for each of the participants 120, 122, 124. As depicted, the coordinate positions may be measured with reference to a room width dimension XROOM and a room depth dimension yRooM. The room width dimension XROOM extends across a width of the conference room 118 from the centerline 56 (see FIG. 2) of the camera 48 (see FIG. 2) so that negative values of XROOM are located to the left of the centerline 56 (see FIG. 2) and positive values of XROOM are located to the right of the centerline 56 (see FIG. 2). In addition, the room depth dimension y ooM extends down a length of the room 118 across the centerline 56 of the camera 48 (see FIG. 2). By applying computer vision processing to the picture image 116, a first meeting participant 120 is detected in the back left comer of the room 118, and an interest region around the head of the first meeting participant 120 is framed with a first head bounding box 126, where the first meeting participant 120 is located at the two-dimensional room distance parameters {XROOM = -3, ynooM = 21 } . Tn similar fashion, a second meeting participant 122 seated at a table 136 is detected with the head of the second meeting participant 122 framed with a second head bounding box 128, where the second meeting participant 122 is located at the two-dimensional room distance parameters {XROOM = -1, YROOM = 13}. Finally, a third meeting participant 124 standing to the right is detected with the head of the third meeting participant 124 framed with a third head bounding box 130, where the third meeting participant 124 is located at the two-dimensional room distance parameters {XROOM = 5, yRooM = 14}.
[0033] In particular, the statistical distribution of human head height and width measurements may be used to determine a min-median-max measure for the head size in centimeters. Additionally, by knowing the FoV resolution of the camera 48 in both horizontal and vertical directions with the respective horizonal and vertical pixel counts, the measured angular extent of each head can be used to compute the percentage of the overall frame occupied by the head and the number of pixels for the head height and width measures. Using this information to compute a look-up table for min-median-max head sizes (height and width) at various distances, an artificial intelligence (Al) human head detector model can be applied to detect the location of each head in a two-dimensional viewing plane with specified image plane coordinates and associated width and height measures for a head frame or bounding box (e.g., {xbox, ybox, width, height}). By using the reverse look-up table operation, the distance can be determined between the camera 48 and each head that is located on the centerline 56 the camera 48.
[0034] With this understanding of the relationship between image projection size at the camera 48 and distance to an object, the present disclosure provides a method, device, system, and computer readable medium to accurately determining if a source of a sound originates within an acoustic fence 62 or acoustic boundary. The location of meeting subject is determined using an Al human head detector model using room distance parameters, as discussed above. The room distance parameters of human heads are then compared to room parameters that correspond to a two-dimensional acoustic fence 62. In this way, it becomes possible to determine if a particular sound recorded by the microphone array 50 has originated from within an area delimited by the acoustic fence 62 or from outside of the area delimited by the acoustic fence 62. If the sound is determined to have originated from within the acoustic fence 62, the microphone array 50 unmutes the sound and transmits the sound to the videoconference. However, if the sound is determined to have originated from outside of the acoustic fence 62, the microphone array 50 mutes the sound and does not transmit the sound to other participants in the videoconference.
[0035] FIG. 7 illustrates a top view of the conference room 40 of FIG. 1 with one person speaking within the acoustic fence 62. In the example of FIG. 7, four subjects or individuals 166, 168, 170, 172 are seated in the chairs 44 and two individuals 174, 176 are standing outside of a door 178 of the conference room 40. In this example, individual 166 is speaking, as indicated by the shading of individual 166 and an SSL line 180 directed towards individual 166. In some examples, any individual detected within the FoV of the camera 48 who is speaking or otherwise creating sound will also be detected using SSL, and subsequent SSL lines will be used to detect the other individuals who are creating sound. In some aspects, the FoV of the camera 48 is between 10 degrees and 120 degrees, or the FoV of the camera 48 is dependent on the camera 48. Additionally, the pan angle P AN (i.e., the angle defined between the centerline 56 of the microphone array 50 and the SSL line) is used in combination with an Al human head detector process as described above in order to define spatial locations for each of the individuals 166, 168, 170, 172 in the FoV of the camera 48. The Al human head detector model applies head bounding boxes 182, 184, 186, 188 to each detected individual 166, 168, 170, 172, and sizes of the bounding boxes 182, 184, 186, 188 are dependent upon an individual’s distance (e.g., distance 190) from the camera 48 and pan angle OP AN from the centerline 56 of the camera 48. Individuals 174, 176 are not within the FoV of the camera 48 and are therefore not captured by the camera 48 or provided with bounding boxes. In the illustrated non-limiting example, the FoV of the camera 48 is wider than dimensions of the conference room 40, so portions of walls of the conference room 40 are also captured by the camera 48. However, in some examples, it is contemplated that the camera 48 may have a wider or narrower FoV than illustrated in FIG. 7.
[0036] After the individuals 166, 168, 170, 172 are detected by the Al head detection process and locations of the detected individuals 166, 168, 170, 172 are determined using a combination of the Al head detection process and SSL, the location of each of the detected individuals 166, 168, 170, 172 are compared to two-dimensional boundaries 192, 194, 196 of the acoustic fence 62. In some examples, the acoustic fence 62 is created during a calibration phase in which a moderator or participant (not shown) draws the acoustic fence 62 or sets the boundaries of the acoustic fence 62 manually. In some examples, the videoconferencing system automatically detects dimensions for a room 40 and locations of meeting participants, and the videoconferencing system sets the acoustic fence 62 boundaries based on the detected room 40 dimensions and the locations of the meeting participants. In the example, the acoustic fence 62 is arranged in a generally triangular configuration defined by a first side 192 extending from the camera 48 to a first wall 198 of the conference room 40, a second side 194 extending from the camera 48 to a second wall 200 of the conference room 40, and a base side 196 that extends between the first side 192 and the second side 194 in a direction that is perpendicular with respect to the centerline 56 of the camera 48. In some examples, the acoustic fence 62 defines a first distance 202 from the videoconferencing device 46 and a second distance 204 from the videoconferencing device 46 that is greater than the first distance 202. Sounds that are determined to have originated between the first distance 202 and the second distance 204 are unmuted. In this way, sounds that originate at distances that are less than the first distance 202 are muted, and sounds that originate at distances that are greater than the second distance 204 are muted. In some examples, the acoustic fence 62 defines a first pan angle 206 from the centerline 56 of the camera 48 and a second pan angle 208 from the centerline 56 of the camera 48. Sounds that are determined to have originated between the first pan angle 206 and the second pan angle 208 are unmuted. In other words, an origin point of the sound is determined based on the pan angle, the centerline, the room coordinates, and dimensions of the acoustic fence.
[0037] In some examples, the acoustic fence 62 may be arranged in any suitable shape or configuration that is particularly desirable to videoconference meeting participants, such as configurations based on an environment of a room 40, or the acoustic fence 62 can be manually set by a user using room coordinates. In some examples, a moderator can change the boundaries of the acoustic fence 62 during a videoconference. Examples of different configurations in which the acoustic fence 62 can be arranged are described below with respect to FIGS. 10 and 11.
[0038] The comparison of the locations of the individuals with the boundaries of the acoustic fence 62 determines if a sound has originated within an area 210 delimited by the acoustic fence 62 or if a sound has originated from outside the area 210 delimited by the acoustic fence 62. In the example of FIG. 7, individuals 166, 170 are within the acoustic fence 62 while individuals 168, 172, 174, and 176 are outside of the acoustic fence 62. Since individual 170 (z.e., the active speaker as indicated by the SSL line) is within the acoustic fence 62, the microphone array 50 is unmuted and the sound created by individual 170 is transmitted to the videoconference. If individual 168 begins speaking, the microphone array 50 will also record and transmit the sound created by individual 168 since individual 168 is also within the acoustic fence 62.
[0039] FIG. 8 illustrates the conference room 40 of FIGS. 1 and 7 with one person speaking from outside of the acoustic fence 62. In this example, individual 168 is speaking, and the videoconferencing device 46 has used SSL to locate the sound originating from individual 168 as indicated by the SSL line. Since individual 168 is outside of the area 210 delimited by the acoustic fence 62, the microphone array 50 mutes sound originating from individual 168 and does not transmit the sound originating from individual 168 to the videoconference. However, if individual 168 moves within the acoustic fence 62, the microphone array 50 will unmute sound that individual 168 creates while within the acoustic fence 62.
[0040] FIG. 9 illustrates the conference room 40 of FIGS. 1, 7, and 8 with multiple individuals speaking on either side of the acoustic fence 62. In this example, individuals 168, 170, 174 are speaking, and the videoconferencing device 46 has used SSL to locate the sound originating from individuals 168, 170 as indicated by the SSL lines, 180a, 180b, respectively. Since individual 174 is located outside of the FoV of the camera 48, sound that is created by individual 174 is automatically muted and is not transmitted to the far end of the videoconference. Since individual 170 is within the area 210 delimited by the acoustic fence 62, the microphone array 50 is unmuted and the sound created by individual 170 is transmitted to the videoconference. Since individual 168 is outside of the area 210 delimited by the acoustic fence 62, the microphone array 50 mutes sound originating from individual 168 and does not transmit the sound originating from individual 168 to the videoconference. The videoconferencing device 46 is capable of identifying any number of individuals that are speaking within the FoV of the camera 48, but the microphone array 50 unmutes sound and transmits the sound to the videoconference if the sound is determined to have originated within the acoustic fence 62.
[0041] FIG. 10 illustrates the conference room 40 of FIGS. 1 and 7-9 with another acoustic fence 262 that is arranged as a rectangle. In this example, individuals 168, 170 are located within the acoustic fence 262, and individuals 166, 172 are located outside of the acoustic fence 262. The microphone array 50 unmutes any speech or noise that is determined to have originated from individuals 168, 170, and the microphone array 50 mutes any speech or noise that is determined to have originated from individuals 166, 172. In some examples, muting is particularly useful for sounds that are closer to the videoconferencing device 46 than active participants of the videoconference , such as situations in which a prompter or moderator is located toward a front 264 of the room 40 while another participant is actively presenting or speaking. Any noise created by the moderator is muted and thus is not transmitted to the videoconference.
[0042] FIG. 11 illustrates the conference room 40 of FIGS. 1 and 7-10 with one individual speaking from within another acoustic fence 362. As discussed above, the acoustic fence 362 may be arranged in any suitable shape or configuration that is particularly desirable to videoconference meeting participants. In some examples, more than one acoustic fence 362 is used in order to delimit particular areas of the videoconferencing room 40. In this example, a first acoustic fence 364 is arranged as a rectangle around individual 170, and a second acoustic fence 366 is arranged as a rectangle around individual 172. Since individual 170 is within the first acoustic fence 364, and individual 172 is within the second acoustic fence 366, the microphone array 50 unmutes noise created by individual 170 and/or individual 172 and the noise is transmitted to the videoconference. In some examples, this arrangement is useful in other environments in addition to conference rooms or enclosed rooms, such as public places or open concept workspaces, and hybrid workspaces. Specifically, the two-dimensional acoustic fence 362 can be used to isolate individual participants in a videoconference to prevent unwanted background noise from being recorded by the microphone array 50 and being transmitted to other participants in the videoconference.
[0043] The Al subject detection process will now be described in greater detail. In some examples, the subject detection process is similar to the Al head detection process as disclosed in U.S. Patent Application No. 17/971,564 filed on October 22, 2022, which is incorporated by reference herein in its entirety. Referring specifically to FIG. 12, a side view is illustrated of an example vertical dimension determination 400 for a human head 402. In some examples, a camera 48 is positioned to capture an image of the human head 402 so that a vertical head height measure V that can be calculated based on an angular extent angle 0FRAME V/2 of the upper half of a vertical head height V/2 and a distance d between the camera 48 and head 402. As illustrated, the human head 402 has a head height V which corresponds to the vertical dimension of a head bounding box (not shown). From the vantage of the camera 48, the vertical head height V makes an angle 0FRAME V extending from the bottom to the top of the head 402. Upon bisecting the angle 0FRAME V, the upper half of the vertical head height V/2 makes an angle 0FRAME V/2 with the camera’s focal point 58 (i.e., the centerline 56 of the camera). As a result, the vertical head height measure V can be calculated based on an angular extent 0FRAME V/2 and the distance d between the camera 48 and head 402 using the equation tan(0FRAME_V/2) = (V/2)/d. Solving for V, the vertical head height measure V may be computed as V = 2d x tan(9FRAME_V/2).
[0044] FIG. 13 illustrates of an example horizontal dimension determination 406 for a human head 408. In some examples, a camera 48 is positioned to capture an image of the human head 408 so that a horizonal head width measure H that can be calculated based on an angular extent angle 0FRAME H/2 of one half of the horizontal head width H/2 and the distance d between the camera 48 and head 408. As depicted, the human head 408 has a head width H which corresponds to the horizonal dimension of a head bounding box. From the vantage of the camera 48, the horizontal head width H makes an angle 0FRAME H extending from the sides of the head 408. Upon bisecting the angle 0FRAME H, the upper half of the horizontal head width H/2 makes an angle 0FRAME H/2 with the camera’s focal point 58 (or the line of sight that is straight at 90-degrees from the camera focal point 58). As a result, the horizontal head width measure H can be calculated based on an angular extent 0FRAME H/2 and the distance d between the camera 48 and head 408 using the equation tan(0 FRAME H/2) = (H/2)/d. Solving for H, the horizontal head width measure H may be computed as H = 2d x tan(FRAME_H/2). As the human head is moved laterally or sideways from the centerline 56 of the camera focal point 58 at a pan angle P AN, the perspective projection cause the head to appear to be smaller than its original vertical and horizontal head measures V, H.
[0045] FIG. 14 illustrates a camera 48 and a two-dimensional image plane 410. In some examples, a video conference camera 48 is used to provide an image of a meeting participant located in a first, centered position 412 and a second, panned position 414 that is shifted laterally in the XROOM direction. In the first, centered position, the meeting participant is located on the centerline 56 the camera (e.g., OP AN = 0) at dO = Y meters, so the two-dimensional room distance parameters for the first, centered position 412 are {XROOM = 0, yRooM = Y). In the second, panned position, the meeting participant is shifted laterally in the XROOM direction by a panned angle OP AN and is located at dl > dO meters, so the two-dimensional room distance parameters for the second, panned position 414 are {XROOM = P, YROOM = Y}. Further, the same vertical head height measure V/2 for the meeting participant locations 412, 414 will result in an angular extent 0FRAME V I/2 for the first meeting participant location 412 that is larger than the angular extent 0FRAME V2/2 for the second meeting participant location 414. In effect, the fact that the second, panned position 414 is located further away from the camera 48 than the first, centered position 412 (dl > dO) results in the angular extent for the second, panned position 414 appearing to be smaller than the angular extent for the first, centered position 412 so that 0FRAME Vl/2 > 0FRAME V2/2.
[0046] From the foregoing, the issue is to find an angular extent for the entire head height HH and then represent it as a percentage of the full frame vertical field of view (VFrame_Percentage) which is then translated into the number of pixels the head will occupy (VHead Pixel Count) at a particular distance and at a pan angle P AN. To this end, the angular extent for the entire head height HHI for the first meeting participant location 412 may be calculated by starting with the equation, tan(0HHi/2) = (V/2)/d0. Solving for the angular extent 01, the angular extent for the entire head height HHI may be calculated as HHI = 2 arctan((V/2)/dO). In similar fashion, the angular extent for the entire head height HH2 for the second meeting participant location 414 located at the pan angle P AN may be calculated by starting with the equation, tan(0HH2/2) = (V/2)/dl, where dl = VdO2 + P2. Solving for the angular extent HH2, the angular extent for the entire head height HH2 may be calculated as HH2 = 2 X arctan((V/2)/dl) = 2 x arctan((V/2 dO2 + P2)). Based on this computation, the percentage of the frame occupied by the head height for the second meeting participant location 414 can be computed as VFrame Percentage = 0HH2/Vertical FoV. In addition, the corresponding number of pixels for the head height for the second meeting participant location 414 can be computed as VHead Pixel Count = VFrame Percentage x Vertical FoV in pixels. Based on the foregoing calculations, the angular extent for the entire head height HH = 0FRAME V may be calculated at discrete distances of 0.5 meters in each of the XROOM and yRooM directions that are equivalent to the various angular pan angles OP AN listed in FIG 3.
[0047] FIG. 15 illustrates a camera 48 and a two-dimensional image plane 510 to illustrate how to calculate a vertical or depth room distance YROOM (meters) to the meeting participant location from the distance measure XROOM (meters) by calculating a direct distance measure HYP between the camera 48 and the meeting participant location. The two-dimensional image plane 510 includes a plurality of two-dimensional coordinate points 512, 514, 516 that are defined with image plane 510 coordinates {xi, y as described above. In addition, a head bounding box 518 is defined with reference to the starting coordinate point {xi, yi} for the head bounding box 518, a Width dimension (measured along the xi axis), and a Height dimension (measured along the yi Tha). To locate the vertical or depth room distance YROOM (meters) from the camera 48, a vertical angular extent (0) for the head bounding box 518 is computed as 0 = Height * V FoV / V PIXELS, where Height is the height of the head bounding box in pixels, where V FoV is the Vertical FoV in degrees, and where V_PIXELS is the Vertical FoV in Pixels. Next, a vertical angular extent for the upper half of the head bounding box is computed (0/2) and used to derive the direct distance measure HYP between the camera 48 and the meeting participant location, HYP = V_HEAD/(2 x tan(0/2)), where HYP is the direct distance measure to the meeting participant location at the pan angle OP AN. Finally, the vertical or depth room distance YROOM (meters) is derived from the direct distance measure HYP and the distance measure XROOM (meters) using Pythagorean’s Theorem, YROOM = HYP2 — XR00M 2.
[0048] FIG. 16 illustrates of an example camera 548 and microphone array 550, similar to the camera 48 and the microphone array 50, respectively. The camera 548 has a housing 552 with a lens 554 provided in the center to operate with the imager 556. A series of five openings 558 are provided as ports to microphones in the microphone array 550. In some examples, the microphone openings 558 form a horizontal line 560 to provide the desired angular determination for the SSL process. This is an example illustration of a camera 548 and numerous other configurations are possible, with varying lens and microphone configurations. In some examples, aspects of the technology, including computerized implementations of methods according to the technology, can be implemented as a system, method, apparatus, or article of manufacture using standard programming or engineering techniques to produce software, firmware, hardware, or any combination thereof to control a processor device (e.g., a serial or parallel general purpose or specialized processor chip, a single- or multi-core chip, a microprocessor, a field programmable gate array, any variety of combinations of a control unit, arithmetic logic unit, and processor register, and so on), a computer (e.g., a processor device operatively coupled to a memory), or another electronically operated controller to implement aspects detailed herein. Accordingly, for example, the technology can be implemented as a set of instructions, tangibly embodied on a non- transitory computer-readable media, such that a processor device can implement the instructions based upon reading the instructions from the computer-readable media. Some examples of the technology can include (or utilize) a control device such as, e.g., an automation device, a special purpose or general-purpose computer including various computer hardware, software, firmware, and so on, consistent with the discussion below. As specific examples, a control device can include a processor, a microcontroller, a field-programmable gate array, a programmable logic controller, logic gates etc., and other suitable components for implementation of appropriate functionality e.g., memory, communication systems, power sources, user interfaces and other inputs, etc.).
[0049] The above description assumes that the axes of camera 48 and the microphone array 50 were collocated. If the axes are displaced, the displacement is used in translating the determined sound angle from the micro-phone array to the camera frames of reference.
[0050] FIG. 17 illustrates a method 600 of implementing a two-dimensional acoustic fence in a conference room. In step 602, it is determined if sound is present in the output of the microphone. If not, operation proceeds to step 616 to mute the microphone, and operation then returns to step 602. If sound is present, in step 604, SSL is used to determine the angle of the sound from the centerline of the microphone array by analyzing the audio output signals of the microphone array (i.e., the microphone output), which centerline is preferably aligned parallel with a line normal to the center of the camera lens. SSL can be performed as disclosed in U.S. Pat No. 6,912,178, which is hereby incorporated herein by reference in its entirety, or by other desired methods. In some examples, another step (not shown) of determining the FoV angle and centerline angle of the camera are determined. If the camera performs physical pan, tilt and zoom (MTZ), the centerline angle and FoV are determined based on the pan angle and the zoom amount. If the camera performs electronic MZ (EPTZ), the centerline angle is determined by the number of pixels the center of the image or frame is displaced from the center of the camera full image and the number of pixels in the full width of the camera. The FoV is determined by the ratio of the width of the image to the width of the full image and applying that ratio to the FoV of the camera.
[0051] In step 606, an Al head detection process is used to detect human heads in the conference room. In step 608, the Al head detection process uses the sound angle and is applied to images captured by the camera in order to identify, for each detected human head, a head bounding box with specified room coordinates and dimension information which are then used to calculate horizontal pan distance and depth dimension distance measured from the camera from a top-down perspective of the room. In this way, a two-dimensional room coordinate location for each detected human head is determined.
[0052] In step 610, the location of each detected human head is compared to a location of an acoustic fence that is created during a calibration phase or inputted by the user/administrator during room or video conferencing set-up step. In some examples, the acoustic fence can be drawn by a moderator, or the boundaries of the acoustic fence can be manually input to the videoconferencing system using image coordinates. In step 612, the room coordinates and dimension information for each detected human head are checked against the boundaries of the acoustic fence, based on the centerline angle, the camera FoV, and the depth dimension measured from the camera. If the centerline of the microphone array and the camera are aligned, this is a simple comparison. If the two centerlines are not aligned, as would be the case in an EPTZ camera, the angle between the two centerlines is determined and used to adjust the determined angle of the sound to be based on the camera centerline. If the sound is determined to originate from within the acoustic fence, the microphone is unmuted in step 614. If the sound is determined not to have originated from within the acoustic fence, the microphone is muted in step 616. Operation returns to step 602 so that the muting and unmuting are automatic as detection of sound is performed.
[0053] In the above description the microphone is unmuted if sound is found to have originated from within the acoustic fence. In some examples, a further step is included to determine if the sound is speech before unmuting the microphone. This keeps the microphone muted for just noise sources, such as fans or other environmental noise, when there is no speech.
[0054] FIG. 18 illustrates a process to determine image plane coordinates for a detected human head using an Al human head detector process. The Al human head detector process analyzes incoming room-view video frame images 702 of a meeting room scene with a head detector machine learning model 704 to detect and display human heads with corresponding head bounding boxes 706, 708, 710. As depicted, each incoming room-view video frame image 702 may be captured by a camera 48 in the video conferencing system. For example, a first view of the meeting participants is captured by a first camera (not shown) in a first profile image or video frame 702a, a second camera (not shown) captures a second profde image or video frame 702b, and a third camera (not shown) captures a third profde image or video frame 702c. Each incoming room-view video frame image 702 may be processed with an on-device Al human head detector model 704 that may be located at the respective camera which captures the video frame images. However, in other examples, the Al human head detector model 704 may be located at a remote or centralized location, or a single camera is used. Wherever located, the Al human head detector model 704 may include a plurality of processing modules 712, 714, 716, 718 which implement a machine learning model which is trained to detect or classify human heads from the incoming video frame images, and to identify, for each detected human head, a head bounding box with specified image plane coordinate and dimension information.
[0055] In this example, the Al human head detector model 704 may include a first preprocessing module 712 that applies image pre-processing (such as color conversion, image scaling, image enhancement, image resizing, etc.) so that the input video frame image is prepared for subsequent Al processing. In addition, a second module 714 may include training data parameters or model architecture definitions which may be pre-defined and used to train and define the human head detection model 716 to accurately detect or classify human heads from the incoming video frame images. In selected examples, the human head detection model 716 may be implemented as a model inference software or machine learning model, such as a Convolutional Neural Network (CNN) model that is specially trained for video codec operations to detect heads in an input image by generating pixel-wise locations for each detected head and by generating, for each detected head, a corresponding head bounding box which frames the detected head. Finally, the Al human head detector model 704 may include a post-processing module 718 which is applies image postprocessing to the output from the Al human head detector model 704 to make the processed images suitable for human viewing and understanding. In addition, the post-processing module 718 may also reduce the size of the data outputs generated by the human head detection model 716, such as by consolidating or grouping a plurality of head bounding boxes or frames which are generated from a single meeting participant so that a single head bounding box or frame is specified.
[0056] Based on the results of the processing modules 712, 714, 716, 718, the Al human head detector model 704 may generate output video frame images 702 in which the detected human heads are framed with corresponding head bounding boxes 706, 708, 710. As depicted, the first output video frame image 702a includes head bounding boxes 706a-c which are superimposed around each detected human head. In addition, the second output video frame image 702b includes head bounding boxes 708a-c which are superimposed around each detected human head, and the third output video frame image 702c includes head bounding boxes 710a, 710b which are superimposed around each detected human head. The Al human head detector model 704 may specify each head bounding box using any suitable pixel-based parameters, such as defining the x and y pixel coordinates of a head bounding box or frame in combination with the height and width dimensions of the head bounding box or frame. In addition, the Al human head detector model 704 may specify a distance measure between the camera location and the location of the detected human head using any suitable measurement technique. The Al human head detector model 704 may also compute, for each head bounding box, a corresponding confidence measure or score which quantifies the model’s confidence that a human head is detected.
[0057] In some examples of the present disclosure, the Al human head detector model 704 may specify all head detections in the data structure that holds the coordinates of each detected human head along with their detection confidence. Specifically, the human head data structure for a number, n, of human heads may be generated as follows:
In this example, xi and yi refer to the image plane coordinates of the i111 detected head, and where Widthi and Height refer to the width and height information for the head bounding box of the ith detected head. In addition, Scorei is in the range (0, 100] and reflect confidence in percentage for the ith detected head. This data structure may be used as an input to various applications, such as framing, tracking, composing, recording, switching, reporting, encoding, etc. In this example data structure, the first detected head is in the image frame in a head bounding box located at pixel location parameters xi, yi and extending laterally by Widthi and vertically down by Heighti. In addition, the second detected head is in the image frame in a head bounding box located at pixel location parameters X2, yi and extending laterally by Width2 and vertically down by Height2, and the nth detected head is in the image frame in a head bounding box located at pixel location parameters xn, yn and extending laterally by Widthn and vertically down by Heightn.
[0058] This human head data structure may then be used as an input to the distance estimation process that takes the {Width, Height} parameters of each head bounding box to pick the best matching distance in terms of meeting room coordinates {XROOM, yROOM} from the look-up table by first using one of the Width or Height parameters with a first lookup table, and then using the other parameter as a tie breaking if multiple meeting room coordinates {XROOM, yROOM) are determined by the one. The human head data structure itself may then be modified to also embed the distance information with each Head, resulting in a modified human head data structure that looks like the following: where {XROOMI, yROOMi}, {XROOM2, yROOM2}, ... , {XROOMU, yROOMn} specify the distance of Headi, Head2, . . ., Headn, from the camera, respective, in two-dimensional coordinates.
[0059] FIG. 19 illustrates aspects of a codec according to some examples of the present disclosure. The codec 800 may include loudspeaker} s) 802, though in many cases the loudspeaker 802 is provided in the monitor 804. The codec 800 may include microphone(s) 806 interfaced via a bus 808. The microphones 806 are connected through an analog to digital (AID) converter 810, and the loudspeaker 802 is connected through a digital to analog (D/A) converter 812. The codec 800 also includes a processing unit 814, a network interface 816, a flash or other non-transitory memory 818, RAM 820, and an input/output (I/O) general interface 822, all coupled by a bus 808. A camera 824 is connected to the I/O interface 822. Microphone(s) 806 are connected to the network interface 816. An HDMI interface 826 is connected to the bus 808 and to the external display or monitor 804. Bus 808 is illustrative and any interconnect between the elements can used, such as Peripheral Compo-nent Interconnect Express (PCie) links and switches, Universal Serial Bus (USB) links and hubs, and combinations thereof. The camera 824 and microphones 806, 806 can be contained in housings containing the other components or can be external and removable, connected by wired or wireless connections. [0060] The processing unit 814 can include digital signal processors (DSPs), central processing units (CPUs), graphics processing units (GPUs), dedicated hardware elements, such as neural network accelerators and hardware codecs.
[0061] The flash memory 818 stores modules of varying functionality in the form of software and firmware, generically programs, for controlling the codec 800. Illustrated modules include a video codec 828, camera control 830, framing 832, other video processing 834, audio codec 836, audio processing 838, network operations 840, user interface 842 and operating system, and various other modules 844. In some examples, an Al head detector module is included the modules included in the flash memory 818. Additionally, at least some of the operations of FIG. 17 are performed in the audio processing 838, and the muting and unmuting operations of FIG. 17 are performed in the audio processing 838. The RAM 820 is used for storing any of the modules in the flash memory 818 when the module is executing, storing video images of video streams and audio samples of audio streams and can be used for scratchpad operation of the processing unit 814.
[0062] The network interface 816 enables communications between the codec 800 and other devices and can be wired, wireless or a combination. In one example, the network interface 816 is connected or coupled to the Internet 846 to communicate with remote endpoints 848 in a videoconference Tn one example, the general interface 822 provides data transmission with local devices (not shown) such as a keyboard, mouse, printer, projector, display, exter-nal loudspeakers, additional cameras, and microphone pods, etc.
[0063] In one example, the camera 824 and the microphones 806 capture video and audio, respectively, in the videoconference environment and produce video and audio streams or signals transmitted through the bus 808 to the processing unit 814. In one example of this disclosure, the processing unit 814 processes the video and audio using processes in the modules stored in the flash memory 818. Processed audio and video streams can be sent to and received from remote devices coupled to network interface 816 and devices coupled to general interface 822.
[0064] Microphones in the microphone array used for SSL can be used as the microphones providing speech to the far site, or separate microphones, such as microphone 806, can be used. [0065] Certain operations of methods according to the technology, or of systems executing those methods, can be represented schematically in the figures or otherwise discussed herein. Unless otherwise specified or limited, representation in the figures of particular operations in particular spatial order can not necessarily require those operations to be executed in a particular sequence corresponding to the particular spatial order. Correspondingly, certain operations represented in the figures, or otherwise disclosed herein, can be executed in different orders than are expressly illustrated or described, as appropriate for particular examples of the technology. Further, in some examples, certain operations can be executed in parallel, including by dedicated parallel processing devices, or separate computing devices that interoperate as part of a large system.
[0066] The disclosed technology is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the following drawings. Other examples of the disclosed technology are possible and examples described and/or illustrated here are capable of being practiced or of being carried out in various ways.
[0067] A plurality of hardware and software-based devices, as well as a plurality of different structural components can be used to implement the disclosed technology. In addition, examples of the disclosed technology can include hardware, software, and electronic components or modules that, for purposes of discussion, can be illustrated and described as if the majority of the components were implemented solely in hardware. However, in one example, the electronic based aspects of the disclosed technology can be implemented in software (for example, stored on non- transitory computer-readable medium) executable by a processor. Although certain drawings illustrate hardware and software located within particular devices, these depictions are for illustrative purposes. In some examples, the illustrated components can be combined or divided into separate software, firmware, hardware, or combinations thereof. As one example, instead of being located within and performed by a single electronic processor, logic and processing can be distributed among multiple electronic processors. Regardless of how they are combined or divided, hardware and software components can be located on the same computing device or can be distributed among different computing devices connected by a network or other suitable communication links. [0068] Any suitable non-transitory computer usable or computer readable medium may be utilized. The computer-usable or computer-readable medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium would include the following: a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, or a magnetic storage device. In the context of this disclosure, a computer-usable or computer-readable medium may be any medium that can contain, store, communicate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
[0069] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms “component,” “system,” “module,” “block,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component can be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. Components (or system, module, and so on) can reside within a process or thread of execution, can be localized on one computer, can be distributed between two or more computers or other processor devices, or can be included within another component (or system, module, and so on).

Claims

1. A method of detecting sound in a location and selecting an audio output to be transmitted to a videoconference, the method comprising: analyzing a microphone output to determine a pan angle of the sound in the location; capturing an image of the location; applying a subject detector model to the image; identifying room coordinates with reference to a camera for each subject detected in the image; defining an acoustic boundary for the image; determining if the sound originates within the acoustic boundary; muting the audio output if the sound does not originate within the acoustic boundary; and unmuting the audio output if the sound does originate within the acoustic boundary.
2. The method of claim 1, wherein analyzing the microphone output includes using a sound source localization process.
3. The method of claim 1, wherein capturing the image of the location includes using the camera having a field of view, and wherein defining the acoustic boundary includes defining the acoustic boundary within the field of view.
4. The method of claim 1 wherein capturing the image of the location includes capturing the image of at least a portion of an enclosed room or a portion of an open concept workspace.
5. The method of claim 1, wherein the acoustic boundary is delimited by a first side, a second side, and a base side, and wherein the base side extends between the first side and the second side.
6. The method of claim 1, wherein a machine leaning human head detector model designs bounding boxes for each human head of each subject that is detected in the image.
7. The method of claim 1 , the method further comprising: measuring, from the camera, a first distance for the sound, wherein the pan angle is measured from a centerline of the camera.
8. The method of claim 7, wherein the acoustic boundary defines a second distance and a third distance.
9. The method of claim 8, the method further comprising: muting the audio output if the first distance is greater than or equal to the second distance and less than or equal to the third distance of the acoustic boundary; and unmuting the audio output if the first distance is less than the second distance or greater than the third distance of the acoustic boundary.
10. A videoconferencing device to select audio to be transmitted in a videoconference, the device comprising: a camera that captures an image; a microphone that receives sound; a processor connected to the camera and the microphone, the processor executing programs to perform videoconferencing operations; and a memory coupled to the processor, the memory storing programs executed by the processor, the programs performing the operations of: determining if a sound is present in an audio output from the microphone; identifying room coordinates for each subject that is detected in the image; defining an acoustic boundary for the image; determining if the sound originates within the acoustic boundary; muting the audio output if the sound does not originate within the acoustic boundary; and unmuting the audio output of the sound does originate within the acoustic boundary.
11. The device of claim 10, wherein determining if the sound is present in the audio output includes using a sound source localization process.
12. The device of claim 10, wherein the camera has a field of view, and wherein defining the acoustic boundary includes defining the acoustic boundary within the field of view.
13. The device of claim 10, wherein the acoustic boundary is delimited by a first side, a second, and a base side, and wherein the base side extends between the first side and the second side.
14. The device of claim 10, wherein the acoustic boundary in includes a first acoustic boundary and a second acoustic boundary, and wherein each of the first acoustic boundary and the second acoustic boundary isolate individual participants in the videoconference.
15. The device of claim 10, wherein identifying the room coordinates for each subject that is detected in the image includes designing bounding boxes for each human head of each subject that is detected in the image.
16. The device of claim 10, and further comprising: measuring a first distance and a pan angle for each human head that is detected in the image, wherein the pan angle is measured from a centerline of the camera.
17. The device of claim 16, wherein the acoustic boundary defines a second distance and a third distance.
18. The device of claim 17, further being performing the operations of: muting the audio output if the first distance is greater than or equal to the second distance and less than or equal to the third distance of the acoustic boundary; and unmuting the audio output if the first distance is less than the second distance or greater than the third distance of the acoustic boundary.
19. A non-transitory computer-readable medium containing instructions that when executed cause a processor to: analyze a microphone output to determine a pan angle of a sound in a location; capture an image of the location; define an acoustic boundary for the image; determine if the sound originates within the acoustic boundary; mute an audio output if the sound does not originate within the acoustic boundary; and unmute the audio output if the sound does originate within the acoustic boundary.
20. The non-transitory computer-readable medium of claim 19, wherein capturing the image of the location includes capturing the image using a camera having a centerline, wherein a microphone array and sound source localization (SSL) determine the pan angle of the sound; wherein room coordinates are determined for each subject that is detected in the image; and wherein an origin point of the sound is determined based on the pan angle, the centerline, the room coordinates, and boundaries of the acoustic boundary.
EP23719164.8A 2023-03-29 2023-03-29 Video conferencing device, system, and method using two-dimensional acoustic fence Pending EP4690769A1 (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/US2023/016764 WO2024205583A1 (en) 2023-03-29 2023-03-29 Video conferencing device, system, and method using two-dimensional acoustic fence

Publications (1)

Publication Number Publication Date
EP4690769A1 true EP4690769A1 (en) 2026-02-11

Family

ID=86142860

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23719164.8A Pending EP4690769A1 (en) 2023-03-29 2023-03-29 Video conferencing device, system, and method using two-dimensional acoustic fence

Country Status (3)

Country Link
EP (1) EP4690769A1 (en)
CN (1) CN121176001A (en)
WO (1) WO2024205583A1 (en)

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6912178B2 (en) 2002-04-15 2005-06-28 Polycom, Inc. System and method for computing a location of an acoustic source
US9215543B2 (en) * 2013-12-03 2015-12-15 Cisco Technology, Inc. Microphone mute/unmute notification
US9530426B1 (en) * 2015-06-24 2016-12-27 Microsoft Technology Licensing, Llc Filtering sounds for conferencing applications
US10939202B2 (en) * 2018-04-05 2021-03-02 Holger Stoltze Controlling the direction of a microphone array beam in a video conferencing system
US11778407B2 (en) 2021-08-10 2023-10-03 Plantronics, Inc. Camera-view acoustic fence

Also Published As

Publication number Publication date
WO2024205583A1 (en) 2024-10-03
CN121176001A (en) 2025-12-19

Similar Documents

Publication Publication Date Title
US8749607B2 (en) Face equalization in video conferencing
WO2017215295A1 (en) Camera parameter adjusting method, robotic camera, and system
US12315189B2 (en) Estimation of human locations in two-dimensional coordinates using machine learning
US11778407B2 (en) Camera-view acoustic fence
US11501578B2 (en) Differentiating a rendered conference participant from a genuine conference participant
CN107005677A (en) Adjust the Space Consistency in video conferencing system
WO2020103078A1 (en) Joint use of face, motion, and upper-body detection in group framing
EP4106327A1 (en) Intelligent multi-camera switching with machine learning
CN107005678A (en) Adjust the Space Consistency in video conferencing system
WO2020103068A1 (en) Joint upper-body and face detection using multi-task cascaded convolutional networks
EP4207751A1 (en) System and method for speaker reidentification in a multiple camera setting conference room
JP5793975B2 (en) Image processing apparatus, image processing method, program, and recording medium
WO2024119902A1 (en) Image stitching method and apparatus
US12154287B2 (en) Framing in a video system using depth information
US20230306698A1 (en) System and method to enhance distant people representation
EP4187898A2 (en) Securing image data from unintended disclosure at a videoconferencing endpoint
US11937057B2 (en) Face detection guided sound source localization pan angle post processing for smart camera talker tracking and framing
EP4614430A1 (en) Systems and methods for image correction in camera systems using adaptive image warping
EP4690769A1 (en) Video conferencing device, system, and method using two-dimensional acoustic fence
US20240338924A1 (en) System and Method for Fewer or No Non-Participant Framing and Tracking
US20250139968A1 (en) Using inclusion zones in videoconferencing
WO2025095949A1 (en) Participant reidentification in multi-camera videoconferencing including central camera
WO2025075641A1 (en) Systems and methods for participant reidentification in multi-camera videoconferencing
US20240303918A1 (en) Generating representation of user based on depth map
HK40045336B (en) Detecting deceptive speakers in video conference

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250919

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR