WO2025259908A1 - Methods and systems for identifying potentially harmful misdirection by urls to thwart malicious phishing attacks - Google Patents
Methods and systems for identifying potentially harmful misdirection by urls to thwart malicious phishing attacksInfo
- Publication number
- WO2025259908A1 WO2025259908A1 PCT/US2025/033403 US2025033403W WO2025259908A1 WO 2025259908 A1 WO2025259908 A1 WO 2025259908A1 US 2025033403 W US2025033403 W US 2025033403W WO 2025259908 A1 WO2025259908 A1 WO 2025259908A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- user
- url
- target url
- warning
- risk
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L9/00—Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols
- H04L9/32—Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols including means for verifying the identity or authority of a user of the system or for message authentication, e.g. authorization, entity authentication, data integrity or data verification, non-repudiation, key authentication or verification of credentials
- H04L9/3263—Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols including means for verifying the identity or authority of a user of the system or for message authentication, e.g. authorization, entity authentication, data integrity or data verification, non-repudiation, key authentication or verification of credentials involving certificates, e.g. public key certificate [PKC] or attribute certificate [AC]; Public key infrastructure [PKI] arrangements
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L63/00—Network architectures or network communication protocols for network security
- H04L63/14—Network architectures or network communication protocols for network security for detecting or protecting against malicious traffic
- H04L63/1408—Network architectures or network communication protocols for network security for detecting or protecting against malicious traffic by monitoring network traffic
- H04L63/1416—Event detection, e.g. attack signature detection
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L63/00—Network architectures or network communication protocols for network security
- H04L63/14—Network architectures or network communication protocols for network security for detecting or protecting against malicious traffic
- H04L63/1441—Countermeasures against malicious traffic
- H04L63/1483—Countermeasures against malicious traffic service impersonation, e.g. phishing, pharming or web spoofing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V2201/00—Indexing scheme relating to image or video recognition or understanding
- G06V2201/02—Recognising information on displays, dials, clocks
Definitions
- the invention generally relates to methods and systems for identifying potentially harmful misdirections to malicious uniform resource locators (URLs) to help prevent phishing attacks and other similar mechanisms for misdirecting a user to a malicious URL.
- URLs uniform resource locators
- Phishing is a common method used by malicious actors to obtain sensitive information from Internet users by misdirecting a user through subterfuge to a URL (address) associated with a nefarious Internet website. Phishing typically involves tricking individuals into disclosing sensitive information through deceptive computer-based means, as nonlimiting examples, by sending an electronic communication (e g., emails, Internet websites, SMS texts, instant messaging texts, etc.) to the user that appears to be from a trustworthy source, but in fact tricks the user into accessing a fake website and divulging sensitive personal information to the malicious actor.
- an electronic communication e g., emails, Internet websites, SMS texts, instant messaging texts, etc.
- a website, email, text, instant message, or other electronic communication directed to the user is provided that is intended to appear to be from a desired, safe, or otherwise trustworthy website (e.g., of a reputable company) with which the user might want to interact; however, the website is in fact controlled by a malicious actor and disguised to persuade the user to freely divulge sensitive information (e.g., passwords, financial account information, secret information, etc.) to the website, load malicious software code onto the user's computer, or otherwise take nefarious actions against the user's interests based on tricking the user into providing sensitive information to the website.
- sensitive information e.g., passwords, financial account information, secret information, etc.
- the present invention provides, but is not limited to, computer-implemented methods of identifying potentially harmful misdirections to malicious URLs to thwart a malicious phishing attack, non-transitory computer-readable media comprising instructions for implementing such computer-implemented methods, and computerized systems for implementing such computer- implemented methods.
- a computer-implemented method for identifying potentially harmful misdirection of a malicious URL to thwart a malicious phishing attack includes capturing an image of a target URL of a high-risk website.
- the captured image is analyzed using a computer image processing module to identify an apparent meaning of the captured image of the target URL.
- the apparent meaning of the captured image is compared to a set of known URLs. Any difference between the captured image and the known URLs is calculated based on a zone of familiarity for an Internet user calculated on prior computer activity of the user and a departure of the target URL from the zone of familiarity. Based on the calculated difference, whether the target URL is a potentially malicious URL is identified. Based on the calculated difference, a warning is generated if the target URL is identified as a potentially malicious URL.
- one or more non-transitory computer-readable media include instructions which, when executed by one or more processors, cause the one or more processors to perform operations of a computer-implemented method as described above.
- a computerized system to identify potentially harmful misdirection to a malicious URL to thwart a malicious phishing attack includes one or more processors configured to implement a computer-implemented method as described above.
- FIG. l is a schematic logic flow diagram of a computer implemented system and method that implements certain nonlimiting aspects of the invention to assist Internet users to more readily identify potential phishing attacks and make better-informed decisions on whether to accept a potential risk associated with following a suspicious link to a target URL.
- FIGS. 2A, 2B, and 2C are screen shots of example actions that may be taken by the system and method of FIG. 1 to warn an Internet user of a phishing risk, including personalized warnings and user interaction functionalities, according to nonlimiting aspects of the invention.
- FIG. 3 schematically illustrates some geographic information that may be associated with a given URL and then utilized by the system and method of FIG. 1 to warn a user of a phishing risk according to a nonlimiting aspect of the invention.
- FIG. 4 is an interactive screen that may be provided with warnings of a phishing risk associated with following a suspicious link to a target URL, wherein the screen includes a risk level designator with which an Internet user can interact with the system and method to accept or reject the risk.
- a data communications network such as the Internet, telecommunications networks, and electromagnetic transmissions, including but not limited to emails, SMS text messages, instant messenger messages (IMs) (e.g., Microsoft Teams® or Slack®), in which may be embedded a hyperlink to a target destination, such as a URL for a website, embedded or remote computer software code, a certificate of authority, web pages (e g., HTTP/HTTPS), file transfer protocols (FTP), emails (mailto), database access (JDBC), and/or other types of computer readable files.
- IMs instant messenger messages
- FTP file transfer protocols
- emails emailto
- JDBC database access
- the terms “a” and “an” to introduce a feature are used as open-ended, inclusive terms to refer to at least one, or one or more of the features, and are not limited to only one such feature unless otherwise expressly indicated.
- use of the term “the” in reference to a feature previously introduced using the term “a” or “an” does not thereafter limit the feature to only a single instance of such feature unless otherwise expressly indicated.
- Various aspects of the present invention are directed to technologies associated with computers and their use, including data science and usable security for computers and data through the Internet.
- the method in some nonlimiting aspects includes generation of local personalized models for Internet users using geographical, temporal, and traditional technical data. These personalized models may be integrated with holistic global models.
- the methods and systems preferably provide a human-centered integrated approach to warning and blocking phishing attacks by integrating cognitive science, behavioral economics, and/or warning science to provide more effective systems and methods of helping users avoid or thwart phishing attacks.
- the methods and systems disclosed herein are preferably capable of generating holistic risk measures for individual users, individual sites, and/or components within such sites. This may include, for example, local identification of phishing websites, malicious scripts, unfamiliar domains, and questionable certificate authorities by identifying anomalies for a single machine in the browsing history of an individual Internet user (identifying locally trusted and familiar actors), as well as by anomaly detection with global data to provide information about universally untrustworthy actors.
- the methods and systems enable the ability to thwart phishing attacks by identifying a zone of familiarity for a specific Internet user and then providing signals to the user about target URLs that depart from that zone of familiarity.
- FIG. 1 diagrammatically represents a system and method 100 (hereinafter optionally referred to as the system 100 or the method 100 as a matter of convenience) of thwarting phishing attacks through electronic communications received by an individual (user).
- the electronic communication may be an email, text, IM, website, or other received electronic communication.
- Such electronic communications are received at a computing device (also referred to simply as a "computer") with one or more processors, hardware, electronic memory, and software configured to receive electronic communications, such as a server, desktop computer, laptop computer, mobile telephone, tablet computer, smart watch, Internet of Things (loT) device, or any other generally similar electronic computing device that is capable of or can be made capable of communications with one or more remote electronic computing devices.
- the electronic communications are typically transmitted via a data communications network, such as the Internet, telecommunications networks, electromagnetic transmissions, etc., from one or more host servers or computers to a receiving computing device.
- the electronic communications could be transmitted physically, for example with a removable memory device, such as a flash drive, compact disc, floppy disc, magnetic tape, punchcards, organic computer memory, or by any other possible mechanism of transmitting the electronic communication to the receiving computer.
- the receiving computer may be one or more computing devices that receive the electronic communication, and typically, although not necessarily, from which the user accesses the electronic communication and interacts with the electronic communication.
- the system and method 100 are preferably implemented by one or more computing devices executing computer readable instructions (e.g., software code and/or hardware configurations) that cause the computing device(s) to execute some or all of the steps of the method 100.
- the system and method 100 may be provided in the form of software instructions encoded in and/or on a physical memory storage device, such as a flash drive, compact disc, floppy disc, magnetic tape, punchcards, organic computer memory, or other non-transitory computer-readable media which, when read and executed by one or more processors, cause the one or more processors to perform the various operational steps of the method 100.
- the system and method 100 are implemented by the receiving computer; however, the system and method 100 may be implemented at least partly with other computing devices remote from the receiving computer.
- the software instructions may be resident on the receiving computer and/or resident on other computing devices operatively coupled with the receiving computer.
- the receiving computer will include some type of electronic display, such as a computer screen, and human machine interface (HMI), such as a mouse and/or keyboard, with which the user can view and interact with the electronic communication to activate a hyperlink in the received electronic communication.
- HMI human machine interface
- an individual e.g., the computer user encounters a URL from any source, such as in email, through Microsoft Teams®, as a text, or loading directly on a browser.
- the user may receive an electronic communication that includes a hyperlink to a target URL.
- the target URL is typically, although not necessarily, embedded within a text string or image that is visible to the user on the computer screen.
- the actual target URL e.g., "http//exampleurladdres.com/" may be hidden from view on the display in that it is not directly displayed on the computer screen, although often the target URL can be seen, for example, if the cursor is hovered over the hyperlink.
- the hyperlink may be embedded within a text string (e.g., "click here” or "XYZ Company Name") and/or within a graphic image, such as a picture and/or company logo.
- a text string e.g., "click here” or "XYZ Company Name”
- a graphic image such as a picture and/or company logo.
- the user is provided with a potentially malicious website with a target URL.
- the system and method 100 leverage information about a group of frequently or frequently targeted websites to train models for developing a list of websites for which access is to be automatically prevented (a "block list"). Phishing risks vary for an individual and an organization. Different organizations are targeted by phishing more or less often and in different campaigns. For example, financial institutions (e.g., banks, electronic payment processors, and stock brokerages) are commonly targeted. Different jurisdictions, sectors, or individual institutions may be targeted based on out-of-band factors (e.g., social or political reasons).
- the system and method 100 identify and/or evaluate risk of subversion based on global variables. In some configurations, this identification may be implemented using an artificial intelligence (“Al”) model, such as convolution neural networks (CNN) that provide for deep learning, machine learning, large language models, and/or generative Al.
- Al artificial intelligence
- CNN convolution neural networks
- the system and method 100 create a personalized group of targeted websites that are known to be legitimate and acceptable to use and from which the system and method 100 are able to generate a list of allowable websites to which access is to be automatically granted (an "allow list").
- an “allow list” For a single person (e.g., the user), individual logins can be defined as more risky or less risky based on various factors, including but not limited to, the role of the individual, the function of the website, and the controls on the network. For example, if the user runs an organization social network account for a company, then the security of that web page and login are of concern to the company.
- a personal social network account may be accessed by an employee in a different role, for example, a developer, with a different level of institutional risk.
- domain name risk is dominated by local reputations in the service of anomaly detection. This can be implemented by Al learning individualized for each participant (individual). New domain names and changes in domain names may be designated as anomalous by definition, especially where the resources they depend on and connect to do not match baseline legitimate behavior.
- Either or both of Blocks 2a and 2b of FIG. 1 may include the creation of data and models of the zone of familiarity associated with individual websites ("site-specific") and/or associated with an individual user ("user-specific” or “customer-specific”).
- This zone of familiarity may be used to identify relative specific websites that are known as safe or harmful as well as other factors that may influence whether a specific URL or website is likely to be (but not yet known) safe or harmful.
- Factors contributing to defining the zone of familiarity may include, for example, geographic information related to a URL, past use behavior of the user relative to what URLs, organizations, and/or areas of interest the user has accessed previously, previous risk level designations selected by the user for various URLs, and known block lists and/or allow lists. Other factors could also be used to help define the zone of familiarity for a given user.
- the system and method 100 may create custom Al models (Block 3a) based on PKI (public key infrastructure), domain name, IP address, BGP (Border Gateway Protocol), ASN (autonomous system number), and/or other information to further define additional aspects of the zone of familiarity that can be used to help evaluate the potential risk of newly encountered hyperlinks or new URLs, websites, computer code, or other potentially malicious links.
- the system and method 100 may create known block lists (Block 3b) of, for example, PKIs, domain names, routing, and/or ASN information.
- the system and method 100 may also obtain information (Block 3c) including geographic (e.g., map) locations and/or political jurisdictions (e.g., country) of any of the PKI, domain name, IP address, BGP, routing, and/or ASN information for one or more of the websites on the lists in either or both of the allow lists and block lists from Blocks 2a and/or 2b of FIG. 1.
- Other information may also be used to help build the zone of familiarity.
- previous work has identified other risk measures that can be used to inform the determination of risk level of an individual user interacting with a specific website. These measures, including changes in location and hosting of previously trusted sites, may be integrated into the risk measure.
- the captured image(s) may include the entire website and URL address bar.
- the image(s) may be captured, for example, with a screen capture module, JavaScript, or other image capture mechanism that is capable of capturing an image of the URL as displayed and seen by the user on the display.
- the captured image(s) of the URL may optionally include an image of any link text or images associated with the URL.
- the captured image(s) is then processed by an image processing and/or natural language processing application that recognizes and meaningfully interprets the apparent or implied meaning of the visible text strings and/or graphic images associated with the URL in the received electronic communication.
- the image processing and NLP applications may incorporate one or more Al models that are configured for language recognition and image recognition capabilities.
- the image processing and Al model may process the visible URL and/or image logo of a legitimate bank and/or a text string name of the legitimate bank that is embedded in the electronic communication to analyze that the image is intended to convey the implication (intended meaning) that the hyperlink is associated with that legitimate bank, and thus intended to make the user believe that the URL and/or website is from or controlled or authorized by that legitimate bank.
- the system and method 100 calculate any differences between the implied or apparent meaning of the captured image and various known URLs from Blocks 2a, 2b, and 3 of FIG. 1 .
- the system and method 100 calculate the difference between the new URL (e.g., web page and/or domain name) presented in the received electronic communication (Block 1 of FIG. 1) as seen by the user and the high risk (block list) websites and/or high value (allow list) websites identified in Blocks 2a, 2b, and 3 of FIG. 1. High levels of apparent similarity to any of these known websites may indicate high levels of risk.
- the present system and method 100 compare the manner in which the text is visually presented to the user on the display.
- the standard conventional method for identification of domains as being suspicious is when a domain has identifiable words or has words that are similar to the target, or includes sensitive words such as "reset. " In a classic example of a Punycode attack, the domain "xn— pple-43d" can be registered, which appears to the human eye as the word “apple” when read on a typical electronic display.
- the URL which will read as “apple” uses the Cyrillic "a” (Unicode U+0430) rather than the ASCII "a” (Unicode U+0061). These are indistinguishable to a human reading the URL on a typical display device.
- the attack is detected if multiple languages are used; however, detection fails when only a single foreign alphabet is used.
- the present system and method 100 capture (convert) the URL as an image as seen by the user on the display and compares what is seen by the human user with possible target domains.
- homograph attacks can include simple manipulation (such as duplication of letters), inclusion of the targeted domain as a subdomain, or inclusion of the targeted domain as a filename to the right of the domain name.
- Subdomains currently dominate.
- the URL "google.resetaccount.com” uses the well- known name “google” as a subdomain, but is probably not actually associated with the actual company by the same name.
- the system and method 100 can identify one or more differences between the apparent suggest source or target of the URL and the actual source or target of the URL.
- the system and method 100 may optionally take one or more risk mitigating actions as deemed appropriate, such as, but not limited to, generating a warning, blocking site elements (as shown in the examples shown in FIGS. 2A, 2B, and 2C), or blocking an entire site.
- the system and method 100 may generate a warning to the user relative to the difference or differences between the apparent or implied legitimate URL and the actual target URL.
- the system and method 100 can present a risk assessment 30 and/or warning indicator 32, as well as an optional risk level designator 34 that indicates what risk level, such as low, medium, or high risk, the user has chosen to accept relative to the target URL or website.
- the risk assessment 30 indicates that "this web page is safe” in green with a green light indicator to show that the web page is safe.
- the warning indicator 32 explains a potential risk consideration, in this case that "the local system firewall is not enabled.”
- the system and method 100 generate user-specific and site-specific warnings using geography information about the target website and individual user browsing and/or other computer usage history.
- the system and method 100 generate a site-specific warning using geographical information about the target website and global data regarding known safe ("allow lists") and unsafe (“block lists”) URLs. These warnings are generated based on the information assembled at Block 3 of FIG. 1 relative to the custom site Al models (Block 3a), known block lists (Block 3b), and map locations and jurisdiction information (Block 3c) relative to various URLs.
- the geographic information, as well as other data, can be used to explain the warnings. In the example shown in FIG.
- a URL 50 of a suspect or high-risk website has an unusual domain name as well as being from a site located in China.
- the target URL 50 includes the term "myaccount.google.com” in the URL, that is followed by securitysettingpage” and hosted by top level domain at "tk.” This may be analyzed as potentially suspicious because the use of the term "google” in the URL suggests that the web page is associated with Google, but the actual address does not appear to be associated with a known Google URL.
- the geographic information associated with the target URL 50 shows that the target URL is from Tokelau (the ".tk” TLD) and that the website is hosted in Beijing.
- the user will bring that information to the system and method 100 and can override or request an override, for example, by using the risk level designator 34 (FIGS. 2A-2C and 4).
- the risk level designator 34 FIGS. 2A-2C and 4
- the user would know that this is a new or unusual domain based on the warnings 30 and 32 provided and thus have a reduced risk of succumbing to a phishing attack.
- the leveraged information may include, for example, risk perception, a mental model, personal history of the user, social network information, expertise of the user, and accessibility, among other factors.
- highly detailed warnings may be provided based on and tailored to the literacy of the particular user.
- the optional acceptable risk level designator 34 shows three different levels 36, 38, 40 of possible risk level that the user can choose to accept for the website (low, medium, and high).
- the acceptable risk level indicator 34 may include a control 42 that allows the user to choose to take a high level of risk.
- a radio button is associated with each of the three levels of risk, so that the user can select what level of risk the user has chosen to take for the particular target website.
- the user has selected to accept a high risk by selecting the radio button 42 associated with the high-risk option 40.
- more or fewer levels of risk could be presented to the user as options.
- FIG. 2A the user has selected to accept only a low level of risk (the piggy in the brick house) with the website.
- the user has selected to accept only a medium level of risk (the piggy in the house made of sticks) with the website.
- FIG. 2C the user has selected to accept only a high level of risk (the piggy in the house made of straw) with the website.
- FIGS. 2A-2C and 4 illustrate an example of an integrated control and warning suitable for an Internet user with relatively low literacy or low technical literacy. These examples have been used for non-technical participants in tests and found to be highly effective; however, other more sophisticated warnings would typically be used for users with higher technical literacy, such as for example, a security operations person.
- the system and method 100 combine the assessments, warnings, and information from Blocks 6, 7a, 7b, and 8 of FIG. 1 to generate customized waming(s) and interaction with the user.
- the system and method 100 may provide one or more warnings and interactions that require the user to click multiple times to choose to take a risk.
- the system and method 100 may allow the user to simply override a given warning or receive a final custom warning.
- the final custom warning may be in the form of a text, video, and/or audio to ensure efficacy and accessibility.
- the system and method 100 can be configured to generate a final custom warning to the user that requires the user to again consider the risk and again choose to accept the risk of following the hyperlink before allowing the user to connect to the site in accordance with the user's over-ride of the system's blocking recommendation.
- the risk communication provided at Block 10 of FIG. 1 can include risk mitigation so that the extension of trust is limited, as illustrated in FIG. 4 and explained in conjunction with Block 8 of FIG. 1. This risk mitigation integrates the warnings 30 and 32 and multiple levels of risk mitigation 36, 38, and 40 from none to high while still allowing the site to connect.
- the system and method 100 provide an advantage over current conventional phishing protection systems that do not provide such warnings and interactions.
- the warnings and actions integrated into the system and method 100 as described above function to support human decision-making by helping to transition a user from fast thinking ("system 1 thinking") to slow thinking (“system 2 thinking”) through the provision of simple to evaluate graphics that encourage the transition from intuitive to rational thinking.
- system and method 100 integrate an understanding of why people are vulnerable to phishing.
- the system and method 100 can be implemented in the loT (Internet of Things) and on general purpose machines, such as servers, desktop computers, laptop computers, mobile telephones, tablet computers, smart watches, Internet of Things (loT) devices, or any other generally similar electronic computing device that is capable of or can be made capable of communications with one or more remote electronic computing devices.
- loT Internet of Things
- general purpose machines such as servers, desktop computers, laptop computers, mobile telephones, tablet computers, smart watches, Internet of Things (loT) devices, or any other generally similar electronic computing device that is capable of or can be made capable of communications with one or more remote electronic computing devices.
- the system and method 100 are well configured for use by individuals and particularly by large enterprises that have experience with ineffective anti-phishing technology. Employees, particularly high-level corporate officers, who have experienced burdensome and disruptive antiphishing training has especially recognize the benefits provided by the system and method 100.
- the system and method 100 may be provided is a single system or implemented as two or more separate products.
- an “awareness” product can be offered that notifies users that a site is new, unfamiliar, and potentially untrustworthy and requires explicit user decision decisions as to whether or not proceed to a website
- a separate “risk modeling” product can be offered that performs risk modeling without requiring any explicit user decisions and can be utilized by the awareness product to generate user warnings or over-ride user decisions.
- Each of the awareness and risk modeling products can be offered or further divided into different product offerings for individuals and commercial settings.
- the awareness product is independently beneficial as it is capable of filling the gaps in phishing resilience and providing insights to individuals and organizations about the extent and driving factors of their vulnerability to phishing.
- Human behavior is a critical component of social engineering attacks exploiting human decision-making, and notifying users that a site is new, unfamiliar, and potentially untrustworthy helps to embed a deep understanding of phishing resilience, integrating behavioral factors, demographical factors, susceptibility, and resilience, as well as technical aspects of phishing site identification.
- the risk modeling product can utilize Al models built upon ground truth about high-risk websites and other malicious entities, including global, public information such as IP address of subverted machines, knowledge of bullet-proof hosting, and public reports of malicious machines, without any explicit user decisions.
- Local distance measure between newly encountered domains with banks and financial institutions can be calculated for domain names.
- routing can be used to map hosting, history, and resolver including identification of malicious and bullet-proof hosts and resolvers.
- Such models can include a comparison of historical patterns of public key certificates with newly encountered or changed certificates. This not only blocks rogues (as does Certificate Transparency (CT), an Internet security standard for monitoring and auditing the issuance of digital certificates) but also refuses engagement with remote jurisdictions without requiring any explicit user decision.
- CT Certificate Transparency
- Certificate Transparency cannot implement this, as it is a global system.
- AS autonomous system
- the relationships between every autonomous system (AS) can be mapped, and the difference documented between the local and global. All of these together provide estimates of the familiarity of a site. If a site is not evaluated as safe and familiar, warnings can be generated to advise as to the reasons that the Al model classified the connection as different. That communication can identify unfamiliar resources as newly visited or as sites that are different from prior interactions.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Security & Cryptography (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Computer Hardware Design (AREA)
- Computing Systems (AREA)
- General Engineering & Computer Science (AREA)
- Information Transfer Between Computers (AREA)
Abstract
Methods and systems for identifying potentially harmful misdirection to a malicious URL to thwart a malicious phishing attack. The methods and systems capture an image of a target URL of a high-risk website and analyze the captured image to compare the apparent meaning of the captured image to a set of known URLs. Differences between the captured image and the known URLs are calculated based on a zone of familiarity for an Internet user calculated on prior computer activity of the user and a departure of the target URL from the zone of familiarity. The method and system identify whether the target URL is a potentially malicious URL based on the calculated difference. A warning is generated if the target URL is identified as a potentially malicious URL.
Description
METHODS AND SYSTEMS FOR IDENTIFYING POTENTIALLY HARMFUL
MISDIRECTION BY URLS TO THWART MALICIOUS PHISHING ATTACKS
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63/658,974 filed June 12, 2024, the contents of which are incorporated herein by reference.
STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under Contract No. 1535375 awarded by the National Science Foundation. The Government has certain rights in the invention.
BACKGROUND OF THE INVENTION
[0003] The invention generally relates to methods and systems for identifying potentially harmful misdirections to malicious uniform resource locators (URLs) to help prevent phishing attacks and other similar mechanisms for misdirecting a user to a malicious URL.
[0004] Phishing is a common method used by malicious actors to obtain sensitive information from Internet users by misdirecting a user through subterfuge to a URL (address) associated with a nefarious Internet website. Phishing typically involves tricking individuals into disclosing sensitive information through deceptive computer-based means, as nonlimiting examples, by sending an electronic communication (e g., emails, Internet websites, SMS texts, instant messaging texts, etc.) to the user that appears to be from a trustworthy source, but in fact tricks the user into accessing a fake website and divulging sensitive personal information to the malicious actor.
[0005] In a typical phishing attack, a website, email, text, instant message, or other electronic communication directed to the user is provided that is intended to appear to be from a desired,
safe, or otherwise trustworthy website (e.g., of a reputable company) with which the user might want to interact; however, the website is in fact controlled by a malicious actor and disguised to persuade the user to freely divulge sensitive information (e.g., passwords, financial account information, secret information, etc.) to the website, load malicious software code onto the user's computer, or otherwise take nefarious actions against the user's interests based on tricking the user into providing sensitive information to the website.
[0006] The success of phishing attacks is fundamentally based on taking advantage of a digital form of social engineering that relies on the user making one or more rapid decisions based on the overall appearance of legitimacy of an electronic communication combined with the user's familiarity with the purported source of the electronic communication without stopping to investigate whether the communication is actually from the purported source. Phishing is effective because the targeted user is engaged in implicit reasoning, using heuristics and familiarity in decision-making. Such intuitive thinking has been called "System 1," or fast thinking, in contrast to "System 2" thinking that requires deliberate, effortful, and orderly reasoning, applying rational deliberate rules. Kahneman, “Thinking, Fast and Slow,” Macmillan Press (2011). Using this categorization of thinking, phishing attacks work when victims respond with System 1 thinking.
[0007] Various systems and methods of trying to combat phishing attacks have been developed. However, typical conventional systems simply show the target URL with a generic warning that there may be something unusual about the target domain, and then provide the user with the opportunity to either follow or not follow the link. However, these conventional systems typically do not provide much useful information to help the user understand what is unusual or otherwise why the target URL has been identified as suspicious. Rather, these conventional systems require the user to carefully read the target URL and have a relatively high-base level of technical understanding to decide what may be suspicious and whether to interact with the website based solely on reading the target URL. However, many users either do not take the time to carefully read the URL and identify what the potential threat is, or they simply do not have the level of technical understanding necessary to decide what may be suspicious about the URL and
whether or not to interact with the website based solely on reading the target URL.
[0008] Therefore, it would be desirable if methods and systems existed that were able to help Internet users to more readily identify potential phishing attacks and make better-informed decisions on whether to accept the potential risk of following a suspicious link to a target URL. Such methods and systems would also preferably provide means for local identification of risks by focusing not only on anomaly detection with global data (enabling the identification of universally untrustworthy actors), but also by identifying anomalies for a single computer in the browsing history of one user (enabling the identification of locally trusted and familiar actors).
BRIEF SUMMARY OF THE INVENTION
[0009] The intent of this section of the specification is to briefly indicate the nature and substance of the invention, as opposed to an exhaustive statement of all subject matter and aspects of the invention. Therefore, while this section identifies subject matter recited in the claims, additional subject matter and aspects relating to the invention are set forth in other sections of the specification, particularly the detailed description, as well as any drawings.
[0010] The present invention provides, but is not limited to, computer-implemented methods of identifying potentially harmful misdirections to malicious URLs to thwart a malicious phishing attack, non-transitory computer-readable media comprising instructions for implementing such computer-implemented methods, and computerized systems for implementing such computer- implemented methods.
[0011] According to a nonlimiting aspect, a computer-implemented method for identifying potentially harmful misdirection of a malicious URL to thwart a malicious phishing attack includes capturing an image of a target URL of a high-risk website. The captured image is analyzed using a computer image processing module to identify an apparent meaning of the captured image of the target URL. The apparent meaning of the captured image is compared to a set of known URLs. Any difference between the captured image and the known URLs is calculated based on a zone of
familiarity for an Internet user calculated on prior computer activity of the user and a departure of the target URL from the zone of familiarity. Based on the calculated difference, whether the target URL is a potentially malicious URL is identified. Based on the calculated difference, a warning is generated if the target URL is identified as a potentially malicious URL.
[0012] According to another nonlimiting aspect, one or more non-transitory computer-readable media include instructions which, when executed by one or more processors, cause the one or more processors to perform operations of a computer-implemented method as described above.
[0013] According to yet another nonlimiting aspect, a computerized system to identify potentially harmful misdirection to a malicious URL to thwart a malicious phishing attack includes one or more processors configured to implement a computer-implemented method as described above.
[0014] Technical aspects of methods, computer-readable media, and computerized systems as described above preferably include the ability to promote the identification and thwarting of phishing attacks.
[0015] These and other aspects, arrangements, features, and/or technical effects will become apparent upon detailed inspection of the figures and the following description.
BRIEF DESCRIPTION OF THE DRAWINGS
[0016] FIG. l is a schematic logic flow diagram of a computer implemented system and method that implements certain nonlimiting aspects of the invention to assist Internet users to more readily identify potential phishing attacks and make better-informed decisions on whether to accept a potential risk associated with following a suspicious link to a target URL.
[0017] FIGS. 2A, 2B, and 2C are screen shots of example actions that may be taken by the system and method of FIG. 1 to warn an Internet user of a phishing risk, including personalized warnings and user interaction functionalities, according to nonlimiting aspects of the invention.
[0018] FIG. 3 schematically illustrates some geographic information that may be associated with a given URL and then utilized by the system and method of FIG. 1 to warn a user of a phishing risk according to a nonlimiting aspect of the invention.
[0019] FIG. 4 is an interactive screen that may be provided with warnings of a phishing risk associated with following a suspicious link to a target URL, wherein the screen includes a risk level designator with which an Internet user can interact with the system and method to accept or reject the risk.
DETAILED DESCRIPTION OF THE INVENTION
[0020] The intended purpose of the following detailed description of the invention and the phraseology and terminology employed therein is to describe one or more nonlimiting embodiments of the invention, and to describe certain but not all aspects of the embodiment s) to which the drawings relate. The following detailed description also identifies certain but not all alternatives of the embodiment(s). As nonlimiting examples, the invention encompasses additional or alternative embodiments in which one or more features or aspects shown and/or described as part of a particular embodiment could be eliminated, and also encompasses additional or alternative embodiments that combine two or more features or aspects shown and/or described as part of different embodiments. Therefore, the appended claims, and not the detailed description, are intended to particularly point out subject matter regarded as the invention, including certain but not necessarily all of the aspects and alternatives described in the detailed description.
[0021] Although the invention will be described hereinafter in reference to systems and methods shown in the drawings that are directed to thwarting phishing attacks from Internet websites, it will be appreciated that the teachings of the invention are more generally applicable to a variety of types of electronic communications that may be transmitted via a data communications network, such as the Internet, telecommunications networks, and electromagnetic transmissions, including but not limited to emails, SMS text messages, instant messenger messages (IMs) (e.g., Microsoft
Teams® or Slack®), in which may be embedded a hyperlink to a target destination, such as a URL for a website, embedded or remote computer software code, a certificate of authority, web pages (e g., HTTP/HTTPS), file transfer protocols (FTP), emails (mailto), database access (JDBC), and/or other types of computer readable files.
[0022] As used herein, the terms "a" and "an" to introduce a feature are used as open-ended, inclusive terms to refer to at least one, or one or more of the features, and are not limited to only one such feature unless otherwise expressly indicated. Similarly, use of the term "the" in reference to a feature previously introduced using the term "a" or "an" does not thereafter limit the feature to only a single instance of such feature unless otherwise expressly indicated.
[0023] Various aspects of the present invention are directed to technologies associated with computers and their use, including data science and usable security for computers and data through the Internet. The method in some nonlimiting aspects includes generation of local personalized models for Internet users using geographical, temporal, and traditional technical data. These personalized models may be integrated with holistic global models. The methods and systems preferably provide a human-centered integrated approach to warning and blocking phishing attacks by integrating cognitive science, behavioral economics, and/or warning science to provide more effective systems and methods of helping users avoid or thwart phishing attacks.
[0024] The methods and systems disclosed herein are preferably capable of generating holistic risk measures for individual users, individual sites, and/or components within such sites. This may include, for example, local identification of phishing websites, malicious scripts, unfamiliar domains, and questionable certificate authorities by identifying anomalies for a single machine in the browsing history of an individual Internet user (identifying locally trusted and familiar actors), as well as by anomaly detection with global data to provide information about universally untrustworthy actors. In some aspects, the methods and systems enable the ability to thwart phishing attacks by identifying a zone of familiarity for a specific Internet user and then providing signals to the user about target URLs that depart from that zone of familiarity.
[0025] Turning now to the nonlimiting embodiments represented in the drawings, FIG. 1
diagrammatically represents a system and method 100 (hereinafter optionally referred to as the system 100 or the method 100 as a matter of convenience) of thwarting phishing attacks through electronic communications received by an individual (user). The electronic communication may be an email, text, IM, website, or other received electronic communication. Typically, such electronic communications are received at a computing device (also referred to simply as a "computer") with one or more processors, hardware, electronic memory, and software configured to receive electronic communications, such as a server, desktop computer, laptop computer, mobile telephone, tablet computer, smart watch, Internet of Things (loT) device, or any other generally similar electronic computing device that is capable of or can be made capable of communications with one or more remote electronic computing devices. The electronic communications are typically transmitted via a data communications network, such as the Internet, telecommunications networks, electromagnetic transmissions, etc., from one or more host servers or computers to a receiving computing device. However, the electronic communications could be transmitted physically, for example with a removable memory device, such as a flash drive, compact disc, floppy disc, magnetic tape, punchcards, organic computer memory, or by any other possible mechanism of transmitting the electronic communication to the receiving computer. The receiving computer may be one or more computing devices that receive the electronic communication, and typically, although not necessarily, from which the user accesses the electronic communication and interacts with the electronic communication.
[0026] The system and method 100 are preferably implemented by one or more computing devices executing computer readable instructions (e.g., software code and/or hardware configurations) that cause the computing device(s) to execute some or all of the steps of the method 100. The system and method 100 may be provided in the form of software instructions encoded in and/or on a physical memory storage device, such as a flash drive, compact disc, floppy disc, magnetic tape, punchcards, organic computer memory, or other non-transitory computer-readable media which, when read and executed by one or more processors, cause the one or more processors to perform the various operational steps of the method 100. Preferably, the system and method 100 are implemented by the receiving computer; however, the system and method 100 may be
implemented at least partly with other computing devices remote from the receiving computer. The software instructions may be resident on the receiving computer and/or resident on other computing devices operatively coupled with the receiving computer. Typically, the receiving computer will include some type of electronic display, such as a computer screen, and human machine interface (HMI), such as a mouse and/or keyboard, with which the user can view and interact with the electronic communication to activate a hyperlink in the received electronic communication.
[0027] At Block 1 in FIG. 1 , an individual (e.g., the computer user) encounters a URL from any source, such as in email, through Microsoft Teams®, as a text, or loading directly on a browser. For example, the user may receive an electronic communication that includes a hyperlink to a target URL. The target URL is typically, although not necessarily, embedded within a text string or image that is visible to the user on the computer screen. The actual target URL (e.g., "http//exampleurladdres.com/") may be hidden from view on the display in that it is not directly displayed on the computer screen, although often the target URL can be seen, for example, if the cursor is hovered over the hyperlink. For example, the hyperlink may be embedded within a text string (e.g., "click here" or "XYZ Company Name") and/or within a graphic image, such as a picture and/or company logo. In other scenarios, the user is provided with a potentially malicious website with a target URL.
[0028] At Block 2a of FIG. 1, the system and method 100 leverage information about a group of frequently or frequently targeted websites to train models for developing a list of websites for which access is to be automatically prevented (a "block list"). Phishing risks vary for an individual and an organization. Different organizations are targeted by phishing more or less often and in different campaigns. For example, financial institutions (e.g., banks, electronic payment processors, and stock brokerages) are commonly targeted. Different jurisdictions, sectors, or individual institutions may be targeted based on out-of-band factors (e.g., social or political reasons). The system and method 100 identify and/or evaluate risk of subversion based on global variables. In some configurations, this identification may be implemented using an artificial
intelligence ("Al") model, such as convolution neural networks (CNN) that provide for deep learning, machine learning, large language models, and/or generative Al.
[0029] At Block 2b of FIG. 1, the system and method 100 create a personalized group of targeted websites that are known to be legitimate and acceptable to use and from which the system and method 100 are able to generate a list of allowable websites to which access is to be automatically granted (an "allow list"). For a single person (e.g., the user), individual logins can be defined as more risky or less risky based on various factors, including but not limited to, the role of the individual, the function of the website, and the controls on the network. For example, if the user runs an organization social network account for a company, then the security of that web page and login are of concern to the company. A personal social network account may be accessed by an employee in a different role, for example, a developer, with a different level of institutional risk. For each institution/group/collaborative or individual, domain name risk is dominated by local reputations in the service of anomaly detection. This can be implemented by Al learning individualized for each participant (individual). New domain names and changes in domain names may be designated as anomalous by definition, especially where the resources they depend on and connect to do not match baseline legitimate behavior.
[0030] Either or both of Blocks 2a and 2b of FIG. 1 may include the creation of data and models of the zone of familiarity associated with individual websites ("site-specific") and/or associated with an individual user ("user-specific" or "customer-specific"). This zone of familiarity may be used to identify relative specific websites that are known as safe or harmful as well as other factors that may influence whether a specific URL or website is likely to be (but not yet known) safe or harmful. Factors contributing to defining the zone of familiarity may include, for example, geographic information related to a URL, past use behavior of the user relative to what URLs, organizations, and/or areas of interest the user has accessed previously, previous risk level designations selected by the user for various URLs, and known block lists and/or allow lists. Other factors could also be used to help define the zone of familiarity for a given user.
[0031] At Block 3 of FIG. 1, the system and method 100 may create custom Al models (Block
3a) based on PKI (public key infrastructure), domain name, IP address, BGP (Border Gateway Protocol), ASN (autonomous system number), and/or other information to further define additional aspects of the zone of familiarity that can be used to help evaluate the potential risk of newly encountered hyperlinks or new URLs, websites, computer code, or other potentially malicious links. The system and method 100 may create known block lists (Block 3b) of, for example, PKIs, domain names, routing, and/or ASN information. The system and method 100 may also obtain information (Block 3c) including geographic (e.g., map) locations and/or political jurisdictions (e.g., country) of any of the PKI, domain name, IP address, BGP, routing, and/or ASN information for one or more of the websites on the lists in either or both of the allow lists and block lists from Blocks 2a and/or 2b of FIG. 1. Other information may also be used to help build the zone of familiarity. Thus, for example, previous work has identified other risk measures that can be used to inform the determination of risk level of an individual user interacting with a specific website. These measures, including changes in location and hosting of previously trusted sites, may be integrated into the risk measure.
[0032] At Block 4 of FIG. 1, for each high-risk website or other URL in the received electronic communication that is above some selected threshold on the block list of Block 2a of FIG. 1, at least one image of the URL of the high-risk website is captured. As such, the system and method 100 do not rely simply on text within a URL, but on an image of that text as it appears to the human eye. The captured image(s) may include the entire website and URL address bar. The image(s) may be captured, for example, with a screen capture module, JavaScript, or other image capture mechanism that is capable of capturing an image of the URL as displayed and seen by the user on the display. The captured image(s) of the URL may optionally include an image of any link text or images associated with the URL. The captured image(s) is then processed by an image processing and/or natural language processing application that recognizes and meaningfully interprets the apparent or implied meaning of the visible text strings and/or graphic images associated with the URL in the received electronic communication. The image processing and NLP applications may incorporate one or more Al models that are configured for language recognition and image recognition capabilities. For example, the image processing and Al model
may process the visible URL and/or image logo of a legitimate bank and/or a text string name of the legitimate bank that is embedded in the electronic communication to analyze that the image is intended to convey the implication (intended meaning) that the hyperlink is associated with that legitimate bank, and thus intended to make the user believe that the URL and/or website is from or controlled or authorized by that legitimate bank.
[0033] At Block 5 of FIG. 1, the system and method 100 calculate any differences between the implied or apparent meaning of the captured image and various known URLs from Blocks 2a, 2b, and 3 of FIG. 1 . For example, using computerized image processing, for example, CNN zero shot, few shot, or other Al training models, the system and method 100 calculate the difference between the new URL (e.g., web page and/or domain name) presented in the received electronic communication (Block 1 of FIG. 1) as seen by the user and the high risk (block list) websites and/or high value (allow list) websites identified in Blocks 2a, 2b, and 3 of FIG. 1. High levels of apparent similarity to any of these known websites may indicate high levels of risk. Thus, unlike previously known methods, in addition to comparing the text as coded within the electronic communication, the present system and method 100 compare the manner in which the text is visually presented to the user on the display. As an example, the standard conventional method for identification of domains as being suspicious is when a domain has identifiable words or has words that are similar to the target, or includes sensitive words such as "reset. " In a classic example of a Punycode attack, the domain "xn— pple-43d" can be registered, which appears to the human eye as the word "apple" when read on a typical electronic display. Thus, it may not be obvious to the person that the URL, which will read as "apple" uses the Cyrillic "a" (Unicode U+0430) rather than the ASCII "a" (Unicode U+0061). These are indistinguishable to a human reading the URL on a typical display device. Conventionally, for current browsers, the attack is detected if multiple languages are used; however, detection fails when only a single foreign alphabet is used. In contrast to other conventional systems that attempt to read only the URL, the present system and method 100 capture (convert) the URL as an image as seen by the user on the display and compares what is seen by the human user with possible target domains. In addition, homograph attacks can include simple manipulation (such as duplication of letters), inclusion of the targeted domain as a
subdomain, or inclusion of the targeted domain as a filename to the right of the domain name. Subdomains currently dominate. For example, the URL "google.resetaccount.com" uses the well- known name "google" as a subdomain, but is probably not actually associated with the actual company by the same name. By analyzing the image of the hyperlink and the URL to determine what legitimate target with which the image appears to most likely be associated, and then comparing that to what the actual URL appears to link to, the system and method 100 can identify one or more differences between the apparent suggest source or target of the URL and the actual source or target of the URL.
[0034] At Block 6 of FIG. 1, the system and method 100 may optionally take one or more risk mitigating actions as deemed appropriate, such as, but not limited to, generating a warning, blocking site elements (as shown in the examples shown in FIGS. 2A, 2B, and 2C), or blocking an entire site. For example, the system and method 100 may generate a warning to the user relative to the difference or differences between the apparent or implied legitimate URL and the actual target URL. As illustrated in FIGS. 2A-2C, the system and method 100 can present a risk assessment 30 and/or warning indicator 32, as well as an optional risk level designator 34 that indicates what risk level, such as low, medium, or high risk, the user has chosen to accept relative to the target URL or website. In FIGS. 2A-2C, the risk assessment 30 indicates that "this web page is safe" in green with a green light indicator to show that the web page is safe. The warning indicator 32 explains a potential risk consideration, in this case that "the local system firewall is not enabled."
[0035] At Block 7a of FIG. 1, the system and method 100 generate user-specific and site-specific warnings using geography information about the target website and individual user browsing and/or other computer usage history. At Block 7b of FIG. 1, the system and method 100 generate a site-specific warning using geographical information about the target website and global data regarding known safe ("allow lists") and unsafe ("block lists”) URLs. These warnings are generated based on the information assembled at Block 3 of FIG. 1 relative to the custom site Al models (Block 3a), known block lists (Block 3b), and map locations and jurisdiction information
(Block 3c) relative to various URLs. The geographic information, as well as other data, can be used to explain the warnings. In the example shown in FIG. 3, a URL 50 of a suspect or high-risk website has an unusual domain name as well as being from a site located in China. Specifically, the target URL 50 includes the term "myaccount.google.com" in the URL, that is followed by securitysettingpage" and hosted by top level domain at "tk." This may be analyzed as potentially suspicious because the use of the term "google" in the URL suggests that the web page is associated with Google, but the actual address does not appear to be associated with a known Google URL. The geographic information associated with the target URL 50 shows that the target URL is from Tokelau (the ".tk" TLD) and that the website is hosted in Beijing. If the user is engaged in seeking new contacts and business opportunities in Beijing or generally in China, the user will bring that information to the system and method 100 and can override or request an override, for example, by using the risk level designator 34 (FIGS. 2A-2C and 4). In this case, the user would know that this is a new or unusual domain based on the warnings 30 and 32 provided and thus have a reduced risk of succumbing to a phishing attack.
[0036] At Block 8 of FIG. 1, information specific to the user is leveraged to help personalize the warnings provided to the user. The leveraged information may include, for example, risk perception, a mental model, personal history of the user, social network information, expertise of the user, and accessibility, among other factors. In this way, highly detailed warnings may be provided based on and tailored to the literacy of the particular user. For example, as best seen in FIG. 4, the optional acceptable risk level designator 34 shows three different levels 36, 38, 40 of possible risk level that the user can choose to accept for the website (low, medium, and high). Optionally, the acceptable risk level indicator 34 may include a control 42 that allows the user to choose to take a high level of risk. In this example, a radio button is associated with each of the three levels of risk, so that the user can select what level of risk the user has chosen to take for the particular target website. In this example, the user has selected to accept a high risk by selecting the radio button 42 associated with the high-risk option 40. In other embodiments, more or fewer levels of risk could be presented to the user as options. In FIG. 2A, the user has selected to accept only a low level of risk (the piggy in the brick house) with the website. In FIG. 2B, the user has
selected to accept only a medium level of risk (the piggy in the house made of sticks) with the website. In FIG. 2C, the user has selected to accept only a high level of risk (the piggy in the house made of straw) with the website. The warning figures illustrated in FIGS. 2A-2C and 4 illustrate an example of an integrated control and warning suitable for an Internet user with relatively low literacy or low technical literacy. These examples have been used for non-technical participants in tests and found to be highly effective; however, other more sophisticated warnings would typically be used for users with higher technical literacy, such as for example, a security operations person.
[0037] At Block 9 of FIG. 1, the system and method 100 combine the assessments, warnings, and information from Blocks 6, 7a, 7b, and 8 of FIG. 1 to generate customized waming(s) and interaction with the user. For example, the system and method 100 may provide one or more warnings and interactions that require the user to click multiple times to choose to take a risk. Based on the organization and the role, the system and method 100 may allow the user to simply override a given warning or receive a final custom warning. The final custom warning may be in the form of a text, video, and/or audio to ensure efficacy and accessibility.
[0038] If the user opts or requests to unblock the URL, at Block 10 of FIG. 1, the system and method 100 can be configured to generate a final custom warning to the user that requires the user to again consider the risk and again choose to accept the risk of following the hyperlink before allowing the user to connect to the site in accordance with the user's over-ride of the system's blocking recommendation. In contrast to current conventional warnings, the risk communication provided at Block 10 of FIG. 1 (and/or optionally Block 9 of FIG. 1) can include risk mitigation so that the extension of trust is limited, as illustrated in FIG. 4 and explained in conjunction with Block 8 of FIG. 1. This risk mitigation integrates the warnings 30 and 32 and multiple levels of risk mitigation 36, 38, and 40 from none to high while still allowing the site to connect. By providing automated risk connection and/or limited connectivity, the system and method 100 provide an advantage over current conventional phishing protection systems that do not provide such warnings and interactions.
[0039] Advantageously, the warnings and actions integrated into the system and method 100 as described above function to support human decision-making by helping to transition a user from fast thinking ("system 1 thinking") to slow thinking ("system 2 thinking") through the provision of simple to evaluate graphics that encourage the transition from intuitive to rational thinking. To do this, the system and method 100 integrate an understanding of why people are vulnerable to phishing.
[0040] The system and method 100 can be implemented in the loT (Internet of Things) and on general purpose machines, such as servers, desktop computers, laptop computers, mobile telephones, tablet computers, smart watches, Internet of Things (loT) devices, or any other generally similar electronic computing device that is capable of or can be made capable of communications with one or more remote electronic computing devices.
[0041] The system and method 100 are well configured for use by individuals and particularly by large enterprises that have experience with ineffective anti-phishing technology. Employees, particularly high-level corporate officers, who have experienced burdensome and disruptive antiphishing training has especially recognize the benefits provided by the system and method 100.
[0042] From the foregoing, it can be appreciated that the system and method 100 may be provided is a single system or implemented as two or more separate products. For example, an “awareness” product can be offered that notifies users that a site is new, unfamiliar, and potentially untrustworthy and requires explicit user decision decisions as to whether or not proceed to a website, whereas a separate “risk modeling” product can be offered that performs risk modeling without requiring any explicit user decisions and can be utilized by the awareness product to generate user warnings or over-ride user decisions. Each of the awareness and risk modeling products can be offered or further divided into different product offerings for individuals and commercial settings.
[0043] The awareness product is independently beneficial as it is capable of filling the gaps in phishing resilience and providing insights to individuals and organizations about the extent and driving factors of their vulnerability to phishing. Human behavior is a critical component of social
engineering attacks exploiting human decision-making, and notifying users that a site is new, unfamiliar, and potentially untrustworthy helps to embed a deep understanding of phishing resilience, integrating behavioral factors, demographical factors, susceptibility, and resilience, as well as technical aspects of phishing site identification.
[0044] The risk modeling product can utilize Al models built upon ground truth about high-risk websites and other malicious entities, including global, public information such as IP address of subverted machines, knowledge of bullet-proof hosting, and public reports of malicious machines, without any explicit user decisions. Local distance measure between newly encountered domains with banks and financial institutions can be calculated for domain names. When connected through a network, routing can be used to map hosting, history, and resolver including identification of malicious and bullet-proof hosts and resolvers. Such models can include a comparison of historical patterns of public key certificates with newly encountered or changed certificates. This not only blocks rogues (as does Certificate Transparency (CT), an Internet security standard for monitoring and auditing the issuance of digital certificates) but also refuses engagement with remote jurisdictions without requiring any explicit user decision. Certificate Transparency cannot implement this, as it is a global system. Similarly, in routing, the relationships between every autonomous system (AS) can be mapped, and the difference documented between the local and global. All of these together provide estimates of the familiarity of a site. If a site is not evaluated as safe and familiar, warnings can be generated to advise as to the reasons that the Al model classified the connection as different. That communication can identify unfamiliar resources as newly visited or as sites that are different from prior interactions.
[0045] As previously noted above, though the foregoing detailed description describes certain aspects of one or more particular embodiments of the invention, alternatives could be adopted by one skilled in the art. As such, and again as was previously noted, it should be understood that the invention is not necessarily limited to any particular embodiment described herein or illustrated in the drawings.
Claims
1. A computer-implemented method of identifying potentially harmful misdirection of a malicious URL to thwart a malicious phishing attack, the method comprising: capturing an image of a target URL of a high-risk website; analyzing the captured image using a computer image processing module to identify an apparent meaning of the captured image of the target URL; comparing the apparent meaning of the captured image to a set of known URLs; calculating any difference between the captured image and the known URLs based on a zone of familiarity for an Internet user calculated on prior computer activity of the user and a departure of the target URL from the zone of familiarity; identifying whether the target URL is a potentially malicious URL based on the calculated difference; and generating, based on the calculated difference, a warning if the target URL is identified as a potentially malicious URL.
2. The computer-implemented method of claim 1, wherein the warning includes an explanation of why the target URL is identified as potentially malicious, wherein the explanation provides context relative to what would not be considered potentially malicious.
3. The computer-implemented method of claim 1, further comprising blocking access to the target URL based on a choice that is presented with the warning and selected by the user.
4. The method of claim 1, wherein the step of generating the warning includes applying to the calculated difference a multi-tiered, non-binary risk mitigation scheme that includes at least three levels of acceptable risk threshold.
5. The method of claim 1 , wherein the step of calculating the difference comprises identifying a legitimate website from the known URLs to which the captured image appears to be
directed and identifying a difference between the legitimate website and the target URL, and wherein the step of generating the warning includes indicating the legitimate website and the identified difference between the legitimate website and the target URL.
6. The method of claim 1, wherein the known URLs include a list of allowed URLs that have been generated by past website use behavior of the user.
7. The method of claim 1, wherein the known URLs include a list of allowed URLs that correspond to known legitimate websites.
8. The method of claim 1, wherein the step of generating the warning includes providing geographic information about the target URL to the user, wherein the geographic information includes at least one of geographic origin of the target URL and geographic location of a hosting site of the target URL.
9. The method of claim 8, wherein the step of calculating the difference includes comparing prior usage activity of the user to the geographic information about the target URL, and wherein the step of generating the warning comprises generating a user-specific warning based on differences between the prior activity of the user and the geographic information about the target URL.
10. The method of claim 9, wherein the step of calculating the difference includes comparing the geographic information for the target URL with geographic information for the known URLs, and the step of generating the warning comprises generating a site-specific warning based on the comparison between the geographic information for the target URL and the geographic information for the known URLs.
11. The method of claim 10, wherein the step of generating the warning comprises creating
a customized warning that combines the user-specific warning and the site-specific warning.
12. The method of claim 1, wherein the step of generating the warning comprises: providing the user with a choice of at least three levels of risk to associate with the target
URL; and having the user select one of the levels of risk to associate with target URL.
13. The method of claim 1, further comprising if the user selects to unblock the target URL, providing a second warning customized for the user and requiring the user to make a second selection to override the warning before allowing the user to connect to the site.
14. The method of any previous claim, wherein the method is performed on or by one or more of a server, desktop computer, laptop computer, mobile telephone, tablet computer, smart watch, and Internet of Things device.
15. The method of any previous claim, wherein the method is implemented with an awareness product, the method further comprising risk modeling implemented with a separate risk modeling product that utilizes artificial intelligence models built upon ground truth about high- risk websites without requiring any explicit user decisions.
16. One or more non-transitory computer-readable media comprising instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising the computer-implemented method of any one of the previous claims.
17. A computerized system to identify potentially harmful misdirection to a malicious URL to thwart a malicious phishing attack, the computerized system comprising one or more processors configured to implement the computer-implemented method of any one of claims 1-15.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463658974P | 2024-06-12 | 2024-06-12 | |
| US63/658,974 | 2024-06-12 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025259908A1 true WO2025259908A1 (en) | 2025-12-18 |
Family
ID=98051646
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/033403 Pending WO2025259908A1 (en) | 2024-06-12 | 2025-06-12 | Methods and systems for identifying potentially harmful misdirection by urls to thwart malicious phishing attacks |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025259908A1 (en) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090064330A1 (en) * | 2004-05-02 | 2009-03-05 | Markmonitor Inc. | Methods and systems for analyzing data related to possible online fraud |
| KR101761513B1 (en) * | 2016-06-13 | 2017-07-26 | 한남대학교 산학협력단 | Method and system for detecting counterfeit and falsification using image |
| US20220182410A1 (en) * | 2020-09-21 | 2022-06-09 | Tata Consultancy Services Limited | Method and system for layered detection of phishing websites |
| US20230403298A1 (en) * | 2022-05-23 | 2023-12-14 | Gen Digital Inc. | Systems and methods for utilizing user profile data to protect against phishing attacks |
| US20240171610A1 (en) * | 2019-03-26 | 2024-05-23 | Proofpoint, Inc. | Uniform Resource Locator Classifier and Visual Comparison Platform for Malicious Site Detection |
-
2025
- 2025-06-12 WO PCT/US2025/033403 patent/WO2025259908A1/en active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090064330A1 (en) * | 2004-05-02 | 2009-03-05 | Markmonitor Inc. | Methods and systems for analyzing data related to possible online fraud |
| KR101761513B1 (en) * | 2016-06-13 | 2017-07-26 | 한남대학교 산학협력단 | Method and system for detecting counterfeit and falsification using image |
| US20240171610A1 (en) * | 2019-03-26 | 2024-05-23 | Proofpoint, Inc. | Uniform Resource Locator Classifier and Visual Comparison Platform for Malicious Site Detection |
| US20220182410A1 (en) * | 2020-09-21 | 2022-06-09 | Tata Consultancy Services Limited | Method and system for layered detection of phishing websites |
| US20230403298A1 (en) * | 2022-05-23 | 2023-12-14 | Gen Digital Inc. | Systems and methods for utilizing user profile data to protect against phishing attacks |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Hijji et al. | A multivocal literature review on growing social engineering based cyber-attacks/threats during the COVID-19 pandemic: challenges and prospective solutions | |
| US11621953B2 (en) | Dynamic risk detection and mitigation of compromised customer log-in credentials | |
| Huang et al. | Security and privacy concerns in ChatGPT | |
| Aswathy et al. | Privacy breaches through cyber vulnerabilities: Critical issues, open challenges, and possible countermeasures for the future | |
| Adu-Manu et al. | Phishing attacks in social engineering: a review | |
| Wedman et al. | An analytical study of web application session management mechanisms and HTTP session hijacking attacks | |
| Lebed et al. | Large Language Models in Cyberattacks | |
| Keerthana et al. | Critical Strategies for Phishing Defense and Digital Asset Protection: A Case Study Approach | |
| Blancaflor et al. | Risk assessments of social engineering attacks and set controls in an online education environment | |
| US20130312115A1 (en) | Human-authorized trust service | |
| WO2025259908A1 (en) | Methods and systems for identifying potentially harmful misdirection by urls to thwart malicious phishing attacks | |
| Pal et al. | Attacks on social media networks and prevention measures | |
| Zainabi et al. | Security risks of chatbots in customer service: a comprehensive literature review | |
| Calderon et al. | Toward a Protocol for Tax Data Security | |
| Kant et al. | Dark Net and Deep Web: Legal Issues and Regulations | |
| Manikanta et al. | AI and Automation in Cybersecurity: Future Skilling for Efficient Defense | |
| Behare et al. | Securing Sign-Up and Sign-In in the Indian Market: Fraud Prevention in CIAM for CRM-A Comprehensive Review and Future Directions | |
| Al-Hyasat et al. | A comprehensive review on digital security and privacy on social networks: The role of users’ awareness | |
| Khadilkar | Securing Internet Banking Against Data Phishing Using Cryptography | |
| Ranjani et al. | Phishing attack detector and awareness generator | |
| Almotiri | Security & Privacy Awareness & Concerns of Computer Users Posed by Web Cookies and Trackers | |
| Alghamdi | Understanding Social Engineering in Cybersecurity and Mitigating Human-Centric Threats | |
| Wayne | Social Engineering: The Effects of Cybercriminals on the Human Mind | |
| Saha et al. | Clickjacking in the Modern Era: Analyzing Emerging Attack Vectors and Advanced Defense Strategies | |
| Vishwanath | The Beginning of the End of Security Awareness Training: How AI Marks a New Era in Cybersecurity |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25822755 Country of ref document: EP Kind code of ref document: A1 |