EP4689919A1 - Computer-implemented methods for searching over the internet, databases, methods of producing such databases, and related systems and computer program products - Google Patents

Computer-implemented methods for searching over the internet, databases, methods of producing such databases, and related systems and computer program products

Info

Publication number
EP4689919A1
EP4689919A1 EP24827102.5A EP24827102A EP4689919A1 EP 4689919 A1 EP4689919 A1 EP 4689919A1 EP 24827102 A EP24827102 A EP 24827102A EP 4689919 A1 EP4689919 A1 EP 4689919A1
Authority
EP
European Patent Office
Prior art keywords
search
travel
travel search
responses
parameters
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24827102.5A
Other languages
German (de)
French (fr)
Inventor
Andreea PASCU
Adam Berk
Nathan COMISKEY
Steve Morley
Cristobal BERGER
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Skyscanner Ltd
Original Assignee
Skyscanner Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from GBGB2408265.3A external-priority patent/GB202408265D0/en
Priority claimed from GBGB2413724.2A external-priority patent/GB202413724D0/en
Application filed by Skyscanner Ltd filed Critical Skyscanner Ltd
Publication of EP4689919A1 publication Critical patent/EP4689919A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/24Querying
    • G06F16/242Query formulation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/24Querying
    • G06F16/245Query processing
    • G06F16/2455Query execution
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/95Retrieval from the web
    • G06F16/951Indexing; Web crawling techniques
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/02Reservations, e.g. for tickets, services or events
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q50/00Information and communication technology [ICT] specially adapted for implementation of business processes of specific business sectors, e.g. utilities or tourism
    • G06Q50/10Services
    • G06Q50/14Travel agencies

Definitions

  • the field of the invention relates to computer-implemented methods for searching over the internet, especially to reducing energy usage in such methods, to related machine learning models, to related training databases for training such machine learning models, and related methods of producing such databases, to related databases for providing input to such trained machine learning models, and related methods of producing such databases, and to related systems, servers, computers and computer program products.
  • EP3557437A1 and EP3557437B1 disclose a case management system which is configured to generate search templates based on selection of a search type and one or more data sources. As configured, the case management system enables execution of searches using the generated search template on synchronous and asynchronous data sources and provides periodic polling of the asynchronous data sources to generate consolidated search results.
  • a computer- implemented method of predicting when to stop polling early in a travel search including the steps of
  • the machine learning model generating an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by applying the learned parameters to the inputs using the machine learning model;
  • step (ix) proceeding to step (x) if the predetermined time is reached, or if the output indicator indicates that the polling should be stopped; otherwise repeating steps (iii) to (viii); (x) storing the distinct responses received;
  • An advantage is that because polling is stopped early, energy consumption is reduced.
  • An advantage is that because polling is stopped early, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
  • the method may be one in which steps (vi) to (viii) take less than 100ms.
  • the method may be one wherein the predetermined time is in the range of 30s to 120s, or wherein the predetermined time is in the range of 45s to 90s, or wherein the predetermined time is in the range of 55s to 65s, or wherein the predetermined time is 60s.
  • the method may be one wherein even when polling (e.g. from a front end) has been stopped, the search is continued (e.g. in a back end), up to a final predetermined time (e.g. 60s) to obtain and to store quotes search data which can be used in future training of the model.
  • a final predetermined time e.g. 60s
  • the method may be one wherein the searches which continue (e.g. in the back end) are a sample of all the searches that are performed, e.g. in the range of 1% to 20% of all the searches that are performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed, so that not all the searches are continued (e.g. in the back end).
  • the searches which continue are a sample of all the searches that are performed, e.g. in the range of 1% to 20% of all the searches that are performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed, so that not all the searches are continued (e.g. in the back end).
  • the method may be one in which the travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
  • the method may be one in which the travel search request is a flight search request.
  • the method may be one in which the flight search request is for a return flight, in which start and destination are specified, outbound and return travel dates are specified, the number of passengers is specified, and cabin class is specified.
  • the method may be one in which when processing the distinct responses received to provide a set of processed distinct responses received, the distinct responses received are processed into itinerary results, wherein each itinerary result includes respective quotes.
  • the method may be one in which for each respective itinerary result, a respective selectable option is provided on a display screen of the user terminal which is selectable to present the quotes associated with the respective itinerary.
  • the method may be one in which when the respective selectable option is selected for the respective itinerary, the user terminal requests the quotes for the respective itinerary from the search service server and the search service server returns the quotes for display on the display screen of the user terminal.
  • the method may be one in which the set of processed distinct responses received are presented in a booking panel.
  • the method may be one in which the search is resumed if the user indicates a potential willingness to make a booking, such as by clicking on a search result to view more detail.
  • the method may be one in which the method is used when contacting an external classifier microservice.
  • the method may be one in which the user terminal is a computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
  • a computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
  • the method may be one in which the trained machine learning model uses Python or Java.
  • the method may be one in which the trained machine learning is deployed in all geographic regions in which the search service (e.g. FPS) is present.
  • the search service e.g. FPS
  • the method may be one in which the search service uses a Java client, e.g. one which is already integrated into Itinerary Filtering/Construction.
  • the method may be one wherein the model travel search parameters derived from the travel search parameters include a query market derived from a destination (e.g. airport).
  • a destination e.g. airport
  • the method may be one wherein the model travel search parameters derived from the travel search parameters include a booking horizon, which is derived from the departure date.
  • the method may be one wherein the model travel search parameters derived from the travel search parameters include a trip duration, which is derived from a departure date and a return date.
  • the method may be one wherein the machine learning model has been trained using a method of any aspect of the second aspect of the invention.
  • the method may be one wherein the trained machine learning model has been trained using a training database produced by a method of any aspect of the fourth aspect of the invention.
  • a computer- implemented method for training a machine learning model, to predict when to stop polling early in a travel search including the steps of:
  • each dataset comprising a respective model travel search request including respective model travel search parameters, and a respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
  • step (vii) storing the parameters which correspond to the trained machine learning model which met the predefined convergence criterion in step (vi).
  • An advantage is that because polling may be stopped early using the trained machine learning model, energy consumption is reduced.
  • An advantage is that because polling may be stopped early using the trained machine learning model, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
  • the method may be one in which the model travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
  • the method may be one in which the stored datasets in the database are a sample of all searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed.
  • the method may be one wherein the database is updated after a predefined time interval, e.g. once per week, or once per day, and the machine learning model is correspondingly retrained using the updated database.
  • the method may be one in which a moving time window is used for the training.
  • the method may be one in which the number of datasets is at least one thousand, or at least ten thousand, at least one hundred thousand, or at least one million, or at least two million.
  • the method may be one wherein the trained machine learning model is trained using a training database produced by a method of any of aspect of the fourth aspect of the invention.
  • a stored trained machine learning model produced by a method of any aspect of the second aspect of the invention.
  • An advantage is that because polling may be stopped early using the trained machine learning model, energy consumption is reduced.
  • An advantage is that because polling may be stopped early using the trained machine learning model, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
  • the method may be one wherein the predetermined time is in the range of 30s to 120s, or wherein the predetermined time is in the range of 45s to 90s, or wherein the predetermined time is in the range of 55s to 65s, or wherein the predetermined time is 60s.
  • the method may be one wherein the threshold is a predetermined threshold.
  • the method may be one wherein the predetermined threshold is in the range 0.40 to 0.95, or wherein the predetermined threshold is in the range 0.60 to 0.90, or wherein the predetermined threshold is in the range 0.70 to 0.85, or wherein the predetermined threshold is 0.8.
  • the method may be one wherein the threshold is a function of the respective model travel search parameters (e.g. in a simple example it is 0.6 when the target market is UK, or it is 0.8 when the target market is USA).
  • the threshold is a function of the respective model travel search parameters (e.g. in a simple example it is 0.6 when the target market is UK, or it is 0.8 when the target market is USA).
  • the method may be one including in step (vi) storing in the first data set the received responses, and in step (x) the respective processed first dataset includes processed received responses derived from the received responses; in which the threshold is a function of the processed received responses.
  • the processed received responses may be processed such that the threshold is set so as to include a cheapest result of the processed received responses in order for the second indicator to be stored.
  • the processed received responses may be processed to rank the processed received responses, and wherein the threshold is set so as to include a most highly ranked processed received response of the processed received responses in order for the second indicator to be stored.
  • the method may be one in which the polling intervals are in the range of 0.2s to 50s, or in which the polling intervals are in the range of 0.5s to 45s, or in which the polling intervals are in the range of 1.0s to 30s.
  • the method may be one in which the stored processed first datasets in the database are a sample of all searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed.
  • the method may be one wherein the database is updated after a predefined time interval, e.g. once per week, or once per day.
  • the method may be one in which the number of search requests is at least one thousand, or at least ten thousand, at least one hundred thousand, or at least one million, or at least two million.
  • the method may be one in which the travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
  • the method may be one in which the travel search request is a flight search request.
  • the method may be one in which the flight search request is for a return flight, in which start and destination are specified, outbound and return travel dates are specified, the number of passengers is specified, and cabin class is specified.
  • the method may be one wherein the respective model travel search parameters derived from the respective travel search parameters include a query market derived from a destination (e.g. airport).
  • a destination e.g. airport
  • the method may be one wherein the respective model travel search parameters derived from the respective travel search parameters include a booking horizon, which is derived from the departure date.
  • the method may be one wherein the respective model travel search parameters derived from the respective travel search parameters include a trip duration, which is derived from a departure date and a return date.
  • the method may be one wherein step (vi) includes deriving and storing an itinerary count from the stored received responses.
  • the method may be one wherein the travel search request is received from a user terminal computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
  • An advantage is that because polling may be stopped early using the trained machine learning model, energy consumption is reduced.
  • An advantage is that because polling may be stopped early using the trained machine learning model, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
  • a stored database for training a machine learning model to predict when to stop polling early in a travel search, the database having been produced and stored using a method of any aspect of the fourth aspect of the invention.
  • Advantages includes those of the fourth aspect of the invention.
  • a computer- implemented method of producing and storing a database for predicting total search results as a function of an input travel search request including input travel search parameters including the steps of:
  • step (ix) processing the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including the respective travel search request including the respective travel search parameters and the respective total number of search results, and the respective time associated with step (vi);
  • step (xi) processing the processed first datasets to produce a database of total number of search results as a function of travel search request including respective travel search parameters, wherein the database entry for a given travel search request including respective travel search parameters is a most recent request for the travel search including respective travel search parameters, wherein the most recent request for the travel search including respective travel search parameters is determined using respective times associated with step (vi).
  • An advantage is that because polling may be stopped early using the trained machine learning model which uses data from the database as input, energy consumption is reduced.
  • An advantage is that because polling may be stopped early using the trained machine learning model which uses data from the database as input, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
  • the method may be one in which the travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
  • the method may be one wherein the predetermined time is in the range of 30s to 120s, or wherein the predetermined time is in the range of 45s to 90s, or wherein the predetermined time is in the range of 55s to 65s, or wherein the predetermined time is 60s.
  • the method may be one in which the polling intervals are in the range of 0.2s to 50s, or in which the polling intervals are in the range of 0.5s to 45s, or in which the polling intervals are in the range of 1.0s to 30s.
  • the method may be one in which the processed first datasets in the database are a sample of all searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed.
  • the method may be one wherein the database is updated after a predefined time interval, e.g. once per week, or once per day.
  • the method may be one in which the number of processed first datasets is at least one thousand, or at least ten thousand, at least one hundred thousand, or at least one million, or at least two million.
  • the method may be one wherein the travel search request is received from a user terminal computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
  • a user terminal computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
  • a stored database for predicting total search results as a function of an input travel search request including input travel search parameters the database having been produced and stored using a method of any aspect of the sixth aspect of the invention. Advantages include those of the sixth aspect of the invention.
  • a system including a server and a computer, the computer configured to execute a trained machine learning model to predict when to stop polling early in a travel search, wherein:
  • the server is configured to receive a travel search request from a user terminal, the travel search request including travel search parameters;
  • the server is configured to consult a database, wherein the database returns an expected total number of search results in response to receiving the travel search request including the travel search parameters from the server;
  • the server is configured to poll a plurality of search service servers using the travel search request including the travel search parameters;
  • the server is configured to receive responses from at least one of the plurality of search service servers, and to store the received responses;
  • the server is configured to process the stored received responses to determine distinct responses received, and then to determine a cumulative number of distinct responses received;
  • the server is configured to convert the travel search request including the travel search parameters into a model travel search request including model travel search parameters, and to input the model travel search request including the model travel search parameters and the determined cumulative number of distinct responses received divided by the expected total number of search results, into the trained machine learning model;
  • the computer is configured to execute the trained machine learning model to process the model travel search request including the model travel search parameters, and the determined cumulative number of distinct responses received divided by the expected total number of search results, through the machine learning model, wherein the machine learning model has been previously trained and comprises a set of learned parameters;
  • the computer is configured to execute the machine learning model to generate an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by applying the learned parameters to the inputs using the machine learning model;
  • (x) the server is configured to store the distinct responses received
  • the server is configured to process the distinct responses received to provide a set of processed distinct responses received
  • the server is configured to send the set of processed distinct responses received to the user terminal.
  • Advantages include those of the first aspect of the invention.
  • the system may be configured to perform a method of any aspect of the first aspect of the invention.
  • a computer configured to train a machine learning model, to predict when to stop polling early in a travel search, wherein the computer is configured to:
  • each dataset comprising a respective model travel search request including respective model travel search parameters, and a respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
  • the computer may be configured to perform a method of any aspect of the second aspect of the invention.
  • a server system configured to produce and to store a training database for training a machine learning model to predict when to stop polling early in a travel search, the server system configured to:
  • (viii) define a first indicator (e.g. a first Boolean value) which indicates that a search should be continued, and define a second indicator (e.g. a second Boolean value different to the first Boolean value) different to the first indicator which indicates that a search should be stopped;
  • a first indicator e.g. a first Boolean value
  • a second indicator e.g. a second Boolean value different to the first Boolean value
  • (x) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including a respective model travel search request including respective model travel search parameters derived from the respective travel search parameters; the respective processed first dataset further including the determined cumulative number of distinct received responses, divided by the respective total number of search results, as a function of successive polling interval; wherein for each stored determined cumulative number of distinct received responses divided by the respective total number of search results, if the value of the cumulative number of distinct received responses divided by the respective total number of search results is less than a threshold, then the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the first indicator, otherwise the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the second indicator;
  • the server system may be configured to perform a method of any aspect of the fourth aspect of the invention.
  • a server system configured to produce and to store a database for predicting total search results as a function of an input travel search request including input travel search parameters, the server system configured to:
  • (xi) process the processed first datasets to produce a database of total number of search results as a function of travel search request including respective travel search parameters, wherein the database entry for a given travel search request including respective travel search parameters is a most recent request for the travel search including respective travel search parameters, wherein the most recent request for the travel search including respective travel search parameters is determined using respective times associated with part (vi).
  • the server system may be configured to perform a method of any aspect of the sixth aspect of the invention.
  • a computer program product executable on a server to:
  • the computer program product may be executable on the server to perform a method of any aspect of the first aspect of the invention.
  • a computer program product executable on a computer to train a machine learning model, to predict when to stop polling early in a travel search, wherein the computer program product is executable to:
  • each dataset comprising a respective model travel search request including respective model travel search parameters, and a respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
  • the computer program product may be executable to perform a method of any aspect of the second aspect of the invention.
  • a fourteenth aspect of the invention there is provided a computer program product executable on a server system to produce and to store a training database for training a machine learning model to predict when to stop polling early in a travel search, the computer program product executable to:
  • (viii) define a first indicator (e.g. a first Boolean value) which indicates that a search should be continued, and define a second indicator (e.g. a second Boolean value different to the first Boolean value) different to the first indicator which indicates that a search should be stopped;
  • a first indicator e.g. a first Boolean value
  • a second indicator e.g. a second Boolean value different to the first Boolean value
  • (x) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including a respective model travel search request including respective model travel search parameters derived from the respective travel search parameters; the respective processed first dataset further including the determined cumulative number of distinct received responses, divided by the respective total number of search results, as a function of successive polling interval; wherein for each stored determined cumulative number of distinct received responses divided by the respective total number of search results, if the value of the cumulative number of distinct received responses divided by the respective total number of search results is less than a threshold, then the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the first indicator, otherwise the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the second indicator;
  • the computer program product may be executable to perform a method of any aspect of the fourth aspect of the invention.
  • a computer program product executable on a server system to produce and to store a database for predicting total search results as a function of an input travel search request including input travel search parameters, the computer program product executable to:
  • (xi) process the processed first datasets to produce a database of total number of search results as a function of travel search request including respective travel search parameters, wherein the database entry for a given travel search request including respective travel search parameters is a most recent request for the travel search including respective travel search parameters, wherein the most recent request for the travel search including respective travel search parameters is determined using respective times associated with part (vi).
  • the computer program product may be executable to perform a method of any aspect of the sixth aspect of the invention.
  • a computer- implemented method of stopping an internet search early the internet search being in a class of internet searches, the method using a trained machine learning model, in which the trained machine learning model has been trained on earlier internet searches in the class, and the successive results produced by those searches, to recognize when sufficient search results have been received to make continuing searching not worthwhile, and to output whether or not to continue searching; the method including inputting the internet search into the trained machine learning model, and the search’s present results, and stopping the search in response to the trained machine learning model outputting that the search should be stopped.
  • An advantage is that because searching is stopped early, energy consumption is reduced.
  • An advantage is that because searching is stopped early, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
  • the method may be one including any aspect of any other aspect of the invention.
  • a seventeenth aspect of the invention there is provided a system configured to perform a method of any aspect of the sixteenth aspect of the invention. Advantages include those of the sixteenth aspect of the invention.
  • computer program products may be embodied on a non-transitory storage medium.
  • Figure 1 shows an example within a travel search service, e.g. within a (e.g flight) pricing service (FPS) in which a quote request is stopped early, using a fixed timeout of 15s.
  • a travel search service e.g. within a (e.g flight) pricing service (FPS) in which a quote request is stopped early, using a fixed timeout of 15s.
  • FPS flight pricing service
  • Figure 2 shows an example, within a travel search service, e.g. within a (e.g flight) pricing service (FPS), of a logic within this service for contacting an external classifier microservice, and if that classifier determines that we should stop the session, only then do we stop the session, within a time interval. If the search time is greater than a lower time limit of 12.8s, but less than an upper time limit of 14.5s, then a classifier is used to determine if the search session should be stopped; if the classifier determines that the search should be stopped, then the search is stopped; if the classifier determines that the search should not be stopped, then the search is not stopped.
  • a travel search service e.g. within a (e.g flight) pricing service (FPS)
  • FPS flight pricing service
  • Figure 3 shows an example of experimentally obtained SHAP values for TTLR, as a function of TTLR.
  • P40 denotes the 40 th percentile of the distribution of SHAP values.
  • P80 denotes the 80 th percentile of the distribution of SHAP values.
  • Figure 4 shows an example of first and second itinerary search results displayed at a user terminal.
  • Figure 5 shows an example of quotes associated with the first itinerary search result of Figure 4 displayed at the user terminal.
  • Figure 6 shows an example of quotes associated with the second itinerary search result of Figure 4 displayed at the user terminal.
  • Figure 7 shows an example of a system including a client computer device e.g. including a browser or an application program executing on the device, a search home website, a front end in the Cloud, a search back-end for front-end (BFF) server, and a search service server, which are in connection with the internet.
  • a client computer device e.g. including a browser or an application program executing on the device
  • a search home website e.g. including a browser or an application program executing on the device
  • a search home website e.g. including a browser or an application program executing on the device
  • a search home website e.g. including a search home website, a front end in the Cloud, a search back-end for front-end (BFF) server, and a search service server, which are in connection with the internet.
  • BFF search back-end for front-end
  • Figure 8 shows data from Ayala et al, in which for a Galaxy Nexus mobile phone, energy consumption in Joules is reduced as a function of increasing polling interval in milliseconds, for messages of 3000 bytes (This derives from Fig. 3(c) in Ayala et al).
  • TTLR time to last result
  • Reducing TTLR is considered to be strongly related to reducing the rate at which users give up on performing a travel search using a travel search website, or in an app, and reducing TTLR is considered to provide a better traveler experience. If we can reduce the time waiting for results that the user will not wait to see, or that the user is unlikely to choose, then we can improve the overall experience for the user. By reducing the TTLR, we may reduce the total energy consumed in a search process.
  • a pricing service (e.g flight) pricing service (FPS) we may stop a quote request early; it is possible to resume the search when the user indicates a potential willingness to make a booking, such as by clicking on a search result to view more detail.
  • a fixed timeout e.g. of 15s
  • all corresponding search sessions will be curtailed at the fixed timeout (e.g. of 15s) point.
  • An example is shown in Figure 1.
  • StopPollingEarlyService e.g. within a (e.g flight) pricing service (FPS)
  • a StopPollingEarlyService which has the responsibility of identifying if this poll should be stopped early.
  • a travel search service e.g. within a (e.g flight) pricing service (FPS)
  • a logic within this service to contact an external classifier microservice, and if that classifier determines that we should stop the session, e.g. during a time interval, only then do we stop the session.
  • the search time is greater than a lower time limit (e.g. 12.8s), but less than an upper time limit (e.g.
  • a classifier is used to determine if the search session should be stopped; if the classifier determines that the search should be stopped, then the search is stopped; if the classifier determines that the search should not be stopped, then the search is not stopped.
  • An example is shown in Figure 2.
  • a classifier is used to determine if the search session should be stopped. Why might we use an upper time limit (e.g. 14.5s) that is less than a target cut off time (e.g. 15s)? Our objective is to increase the number of searches that take less than the target cut off time. So if we consider a poll that will be returned after the target cut off time, we've missed our objective target, and in this approach we should continue to allow polling and to complete the search as usual. By only choosing the searches between a lower time limit (e.g. 12.8s) and an upper time limit (e.g. 14.5s) this gives the poll enough time to complete within a search (e.g flight pricing service (FPS)) within the target cut off time (e.g. 15s).
  • a search e.g flight pricing service (FPS)
  • search results times are obtained in advance, and stored, and a cache is generated which determines when polling should be stopped early, as a function of search parameters. Then when searches are performed, the cache is consulted based on the search parameters of the search, to determine at what time polling should be stopped early, if the search has not yet completed. If the search parameters of the search do not have a good enough match within the cache, then the cache is not used, and polling is not stopped early, or polling is stopped when a target time is reached.
  • the cache file is fetched periodically and is stored within the memory of the search service, e.g. FPS memory.
  • the classifier uses heuristics.
  • the classifier uses a machine learning model.
  • a machine learning model example we may use a language more suited to hosting machine learning models e.g. Python, rather than being limited to Java.
  • a machine learning model example we use Java.
  • the classifier is consulted only once per search session.
  • the machine learning model classifier is consulted only once per search session.
  • the classifier is provided as a service, e.g. as a micro-service, e.g. in python.
  • the service is deployed in all geographic regions that the search service (e.g. FPS) is present; the service receives the current search service (e.g. FPS) state and returns a boolean, e.g. TRUE to stop, FALSE to continue polling.
  • the service may
  • Java client e.g. one which is already integrated into Itinerary Filtering/Construction.
  • Search Service e.g FPS interaction with the Relevance Service
  • the Relevance Service may be integrated into itinerary-filtering.
  • the Relevance Service may be integrated into itinerary-construction.
  • the model the Relevance Service hosts requires access to the filtered priced itineraries.
  • the model the Relevance Service hosts requires access to statistics derived from the filtered priced itineraries.
  • the call to the Relevance Service is made after the call to construction/filtering.
  • the call to the Relevance Service is made before the state has been written to the State store.
  • the call to the Relevance Service is a blocking request, so the classifier has tight constraints around latency as it will directly impact the latency of search service (e.g. FPS).
  • FPS latency of search service
  • the call to the Relevance Service may be a single poll, there is no impact on time to first result (TTFR).
  • the classifier may require two new sources of data to be generated.
  • the first is a source of data that can be used for training, generated during the polling process.
  • the second is a source of data emitted a significant time (e.g. 60s) after we stop a poll: this source of data needs to contain the details of what would have been returned had we not stopped early.
  • the two new sources of data can then be used for monitoring our impact on coverage and price accuracy, as well as feedback for the classifier.
  • an end result is that when the search service (e.g. FPS) stops a session early we send an event to a worker process, which after a significant time (e.g. 60s) polls the search that was completed early.
  • a significant time e.g. 60s
  • Monitoring the Relevance Service Information display may be provided showing service metrics, e.g. memory usage, central processing unit (CPU) usage etc. Errors may be monitored from the service itself, but also from the search service’s (e.g. FPS’s) perspective. Circuit breaker status and/or percentage of errors/timeouts may be monitored.
  • service metrics e.g. memory usage, central processing unit (CPU) usage etc.
  • Errors may be monitored from the service itself, but also from the search service’s (e.g. FPS’s) perspective. Circuit breaker status and/or percentage of errors/timeouts may be monitored.
  • times for a response to be returned from a poll are as follows.
  • the 50 th percentile for the poll response time was 130ms.
  • the 95 th percentile for the poll response time was 1.2s.
  • the 99 th percentile for the poll response time was 4.6s.
  • the polling interval is 1.0 s.
  • flights search for a return flight, the start and destination are specified, outbound and return travel dates are specified, the number of passengers is specified, cabin class is specified, etc.
  • the search website polls sources of data for relevant data, and the search website waits for data.
  • the search website can poll for results at regular time intervals, e.g. every second, for example from a system which receives data and constructs itineraries for the user. Then for example the constructed itineraries are sorted and returned to the front end of the website. Or for example the constructed itineraries are returned to the front end of the website where they are sorted and presented.
  • the algorithm interrupts the process of polling when there is sufficient confidence for this particular search, taking into account the search parameters, that there is insufficient value in continuing polling further.
  • a search progress bar is stopped from being presented on a screen to a user who requested the search.
  • the model has been trained on training data, in which the training data has been selected based on a search time cut off, in which the search time cut off has been determined to be the length of time beyond which users are likely to give up on the search.
  • An example length of time beyond which users are likely to give up on the search is 15s.
  • the polling request search time cut off is 14.8s, which leaves 0.2s for final results to be returned, before the search process is terminated at 15.0s.
  • the classifier may be consulted for the previous e.g. one or two pollings (made at 1.0 s intervals), e.g. hence at 12.8s and 13.8s, to determine if polling should be stopped early.
  • a flights search is requested by a user terminal in communication with a search service server, for flights for one adult in economy class, with the outward flight from London Luton to Tenerife (any airport) on 5 Jan 2025, and the return flight from Tenerife (any airport) to London Luton on 9 Jan 2025, with the flights selected to be restricted to direct flights only.
  • the results are assembled by the search service server and are sent to the user terminal for display on a display screen of the user terminal.
  • two itineraries are returned.
  • the first itinerary is Easy Jet departing Luton at 0700 and arriving at TFS at 1135 on 5 Jan 2025, and returning Easy Jet from TFS at 1255 to Luton at 1725 on 9 Jan 2025.
  • a respective selectable option is provided on the display screen which is selectable to present the quotes, or “deals” associated with the respective itinerary.
  • the user terminal requests the quotes for the first itinerary from the search service server and the search service server returns the quotes for display on a display screen of the user terminal.
  • Figure 5 An example is shown in Figure 5, in which the nine quotes or “deals” are listed by provider, namely easyJet, Mytrip, Gotogate, lastminute.com, Booking.com, eDreams, Expedia, Kiwi.com and BudgetAir.
  • the user terminal when the respective selectable option is selected for the second itinerary, the user terminal requests the quotes for the second itinerary from the search service server and the search service server returns the quotes for display on a display screen of the user terminal.
  • An example is shown in Figure 6, in which the two quotes or “deals” are listed by provider, namely Ryanair +easyJet, and Kiwi.com.
  • a classifier is consulted to decide when to stop updating displayed search results with new itineraries and options. If we wait for every single source of search results to return results from their Application Programming Interface (API) call, we can potentially wait a long time, e.g. for up to 60 seconds, before completing the displayed search results. Often waiting for this long can result in very little advantage for the traveller so in an example a model, a classifier, is consulted during a specific time window, to decide whether to continue waiting for more results or to stop updating the displayed search results.
  • API Application Programming Interface
  • a model has been trained on data from searches, e.g. from small screen devices (e.g. smartphones) based on the state of the search service (e.g. Flights Pricing Service) in a time interval, e.g. 5 - 50 seconds, after the session starts compared to the state of the session, e.g. the search results, when the search completes.
  • search service e.g. Flights Pricing Service
  • flight search sessions are stopped early, if the model is called and the model predicts that the session should be stopped early, and if zero, one, or more, or all, of the following set of constraints are met:
  • the searching device’s platform is android, or ios, or banana or acorn
  • the time of the poll is in between a lower time and an upper time, e.g. between 5 sec and 50 sec
  • the model is closed, specific to a single use case.
  • a travel search request including travel search parameters is converted into a model travel search request including model travel search parameters, for use in querying a trained machine learning model, or for use in generating a training data set (e.g. a database) for training a machine learning model.
  • travel search parameters including a destination airport in Thailand may be converted into model travel search parameters including for example a query market, which is a market code, which here is TH for Thailand, because the destination airport is in Thailand.
  • travel search parameters including a departure date may be converted into model travel search parameters including for example a booking horizon, which is the period of time from the present day to the departure date.
  • travel search parameters including a departure date and a return date may be converted into model travel search parameters including for example a trip duration, which is the period of time from the departure date to the return date.
  • An example input variable is a candidate id, which is a candidate id, which is a string, which may or may not be used in the model.
  • An example input variable is query market, which is a market code eg. TH for Thailand, which is a string, which may be used in the model.
  • An example input variable is query trip type, which is a search kind eg. ONE WAY, which is a string, which may be used in the model.
  • An example input variable is query cabin class, which is a cabin class eg ECONOMY, which is a string, which may be used in the model.
  • An example input variable is query booking horizon, which is a booking horizon e.g. in days eg. 18, which is a number, which may be used in the model.
  • the term "booking horizon" refers to the period of time in advance during which bookings or reservations are being searched for, such as for a flight, hotel room, event, train journey or vehicle rental. This period starts from the current date and extends into the future.
  • An example input variable is query trip duration, which is a trip duration in days, which is a number, eg. zero for a one-way trip, which may be used in the model.
  • An example input variable is query adults, which is the number of adults in the search query eg. 1, which is a number, which may be used in the model.
  • An example input variable is stats itineraries itinerary count, which is a count of itineraries in the poll eg. 15, which is a number, which may be used in the model.
  • An example input variable is stats quote requests current count, which is current quotes eg. 300, which is a number, which may be used in the model.
  • An example input variable is stats quote requests total count, which is total quotes requested eg. 330, which is a number, which may be used in the model.
  • a maximum cutoff time e.g. 60s
  • the input variable stats quote requests total count is a prediction from a database of what the total number of quotes is expected to be, based for example on a search key (e.g. defined as a trip kind (return or one-way trips), route (origin and destination) and inbound and outbound dates).
  • This database can be populated in advance based on stored search results for searches which have been allowed to run for a long time (e.g. 60s) until probably all possible search results have been received.
  • the stored searches in the database which have been allowed to run for a long time may be a sample of all the searches performed, e.g.
  • the number of quotes may be the number of search results, for a given search key.
  • the number of quotes may be the number of instances in which providers or partners have provided search results for a given search key.
  • the database may be updated as new stored search results become available, e.g. the database may be updated once per week, or once per day.
  • the input variable stats quote requests total count is the total number of quotes from providers/partners available for a search key (e.g. defined as a trip kind (return or oneway trips), route (origin and destination) and inbound and outbound dates).
  • a callable service e.g. called “WhoToAsk” maps each search key to the set of partners providing quotes for the search key, e.g. using a database which has been populated in advance based on stored search results for searches which have been allowed to run for a long time (e.g. 60s) until probably all possible search results have been received.
  • the stored searches in the database which have been allowed to run for a long time e.g.
  • 60s may be a sample of all the searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed.
  • the input variable stats quote requests total count doesn’t change during the whole polling process.
  • the database may be updated as new stored search results become available, e.g. the database may be updated once per week, or once per day.
  • stats quote requests current count is the number of quotes from providers/partners returned as the polling process progresses for a given search. For example, let’s assume a polling process lasts for 20 seconds with a 5-second sampling interval: At 5 seconds: Polling returns a Ryanair quote. Then stats quote requests current count will be 1.
  • stats quote requests current count counts the number of unique providers/partners quotes returned which is the cumulative count of distinct quotes from partners as the polling process progresses. That’s why on the first two pollings described in the above example stats quote requests current count is 1.
  • the model internally builds stats quote requests current count divided by stats quote requests total count to construct a quotient e.g. called frac_partners (which is a fraction of partners).
  • the quotient e.g. fraction of partners, is a relevant feature (e.g. the most relevant feature) that helps the model decide when to stop a polling process.
  • the model would be expected to stop polling once a large enough fraction of the possible quotes from partners serving that search have been returned.
  • the total number of providers/partners quotes available for a search key (e.g. defined as a trip kind (return or one-way trips), route (origin and destination) and inbound and outbound dates) is one hundred.
  • the model may be trained to favour stopping polling when at least 80 percent of the expected providers/partners quotes have been returned as search results, so here when 80 of the 100 providers/partners quotes have been returned as search results it is favoured to stop polling early, and if less than 80 of the 100 providers/partners quotes have been returned as search results then stopping polling early is not favoured.
  • the number of unique providers/partners quotes returned which is the cumulative count of distinct quotes by partners as the polling process progresses, is counted.
  • a database is queried to provide the total number of quotes by providers/partners available for the particular search key.
  • the front end queries the trained model which includes as input parameters the cumulative count of distinct quotes by partners as the polling process progresses, the total number of quotes by providers/partners available for the particular search key, and the search key, to obtain an output which indicates whether or not the polling should be stopped early.
  • the trained model tends to stop the polling early when the ratio of the cumulative count of distinct partners which have returned search results quotes for the particular search key, to the total number of quotes by providers/partners available for the particular search key, reaches the target ratio, which may be 80%.
  • the search may continue in the back end, e.g. up to a final time (e.g. 60s) to obtain and to store quotes search data which can be used in future training of the model.
  • the searches which continue in the back end are a sample of all the searches that are performed, e.g. 1% to 20% of all the searches that are performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed, so that not all the searches are continued in the back end.
  • An example output variable is a candidate id, which is a candidate id, which is a string, which may or may not be used in the model.
  • An example output variable is prediction, which is a Boolean, in which for example the prediction value can only be 1 (stop session) or 0 (don't stop).
  • a search is requested for return flights between London, UK and Barcelona, ES on particular dates, and the polling is stopped early after 15s.
  • a search is requested for return flights between London, UK and Fresno, CA, USA on particular dates, and the polling is stopped early after 35s.
  • Response code 412 has meaning Deadline exceeded, which indicates the request to the model has taken too long and a deadline has been exceeded.
  • the model is retrained weekly, typically using a new data set compared to the week before.
  • a moving time window is used for the training e.g. from one week to the next week.
  • the model is retrained after a predetermined time interval.
  • the number of searches on which the model has been trained is at least one thousand. In an example, the number of searches on which the model has been trained is at least ten thousand. In an example, the number of searches on which the model has been trained is at least one hundred thousand. In an example, the number of searches on which the model has been trained is at least one million. In an example, the number of searches on which the model has been trained is at least two million.
  • the number of searches included in the training set on which a model can be trained is at least one thousand. In an example, the number of searches included in the training set on which a model can be trained is at least ten thousand. In an example, the number of searches included in the training set on which a model can be trained is at least one hundred thousand. In an example, the number of searches included in the training set on which a model can be trained is at least one million. In an example, the number of searches included in the training set on which a model can be trained is at least two million.
  • model training procedure follows common approaches, as would be known to those skilled in the art.
  • the framework can be extended to other applications, such as for car hire, for vehicle hire, for accommodation, for events, for train journeys or for hotels.
  • frontend or sometimes referred to as front end or front-end
  • backend or sometimes referred to as back end or back-end
  • front end and backend may be used in respect of software.
  • front end and backend may be used in respect of hardware.
  • Example method of producing and storing a training database for training a machine learning model to predict when to stop polling early in a search e.g. travel search
  • a computer-implemented method of producing and storing a training database for training a machine learning model to predict when to stop polling early in a search e.g. travel search
  • the method including the steps of:
  • search e.g. travel search
  • search e.g. travel search
  • search e.g. travel search
  • search parameters search parameters
  • search e.g. travel search
  • search e.g. travel search
  • steps (i) to (vi) for a plurality of search (e.g. travel search) requests, each respective search (e.g. travel search) request including respective search (e.g. travel search) parameters;
  • search e.g. travel search
  • a respective model search e.g. travel search
  • respective model search e.g. travel search
  • the predetermined time may be in the range of 30s to 120s, or the predetermined time may be in the range of 45s to 90s, or the predetermined time may be in the range of 55s to 65s, or the predetermined time may be 60s.
  • the threshold may be a predetermined threshold.
  • the predetermined threshold may be in the range 0.40 to 0.95, or the predetermined threshold may be in the range 0.60 to 0.90, or the predetermined threshold may be in the range 0.70 to 0.85, or the predetermined threshold may be 0.8.
  • Example method for training a machine learning model, to predict when to stop polling early in a search e.g. travel search
  • a computer-implemented method for training a machine learning model, to predict when to stop polling early in a search including the steps of:
  • each dataset comprising a respective model search (e.g. travel search) request including respective model search (e.g. travel search) parameters, and respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
  • a respective model search e.g. travel search
  • model search e.g. travel search
  • step (vii) storing the parameters which correspond to the trained machine learning model which met the predefined convergence criterion in step (vi).
  • Example method of predicting when to stop polling early in a search e.g. travel search
  • a computer implemented method of predicting when to stop polling early in a search including the steps of:
  • a server receiving a search (e.g. travel search) request from a user terminal, the search (e.g. travel search) request including search (e.g. travel search) parameters;
  • the server consulting a database which returns an expected total number of search results in response to receiving the search (e.g. travel search) request including the search (e.g. travel search) parameters;
  • model search e.g. travel search
  • model search e.g. travel search
  • model search e.g. travel search
  • determined cumulative number of distinct received responses divided by the expected total number of search results through the machine learning model, wherein the machine learning model has been previously trained and comprises a set of learned parameters
  • the machine learning model generating an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by applying the learned parameters to the inputs using the machine learning model;
  • step (ix) proceeding to step (x) if the predetermined time is reached, or if the output indicator indicates that the polling should be stopped; otherwise repeating steps (iii) to (viii);
  • the predetermined time may be in the range of 30s to 120s, or the predetermined time may be in the range of 45s to 90s, or the predetermined time may be in the range of 55s to 65s, or the predetermined time may be 60s.
  • Example method of producing and storing a database for predicting total search results as a function of an input search (e.g. travel search) request
  • a computer-implemented method of producing and storing a database for predicting total search results as a function of an input search (e.g. travel search) request including input search (e.g. travel search) parameters the method including the steps of:
  • search e.g. travel search
  • search e.g. travel search
  • search e.g. travel search
  • search parameters search parameters
  • search e.g. travel search
  • search e.g. travel search
  • the search e.g. travel search
  • the determined cumulative number of distinct received responses as a function of successive polling interval, and a time associated with completing this step
  • steps (i) to (vi) for a plurality of search (e.g. travel search) requests, each respective search (e.g. travel search) request including respective search (e.g. travel search) parameters;
  • search e.g. travel search
  • step (ix) processing the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including the respective search (e.g. travel search) request including the respective search (e.g. travel search) parameters and the respective total number of search results, and the respective time associated with step (vi);
  • the respective search e.g. travel search
  • the respective search e.g. travel search
  • the respective search e.g. travel search
  • step (xi) processing the processed first datasets to produce a database of total number of search results as a function of search (e.g. travel search) request including respective search (e.g. travel search) parameters, wherein the database entry for a given search (e.g. travel search) request including respective search (e.g. travel search) parameters is a most recent request for the search (e.g. travel search) including respective search (e.g. travel search) parameters, wherein the most recent request for the search (e.g. travel search) including respective search (e.g. travel search) parameters is determined using respective times associated with step (vi).
  • search e.g. travel search
  • respective search e.g. travel search
  • the predetermined time may be in the range of 30s to 120s, or the predetermined time may be in the range of 45s to 90s, or the predetermined time may be in the range of 55s to 65s, or the predetermined time may be 60s.
  • SHAP Shapley Additive explanations
  • the Shapley value provides a principled way to explain the predictions of nonlinear models in the field of machine learning.
  • Shapley values provide a natural way to compute which features contribute to a prediction or contribute to the uncertainty of a prediction.
  • SHAP values provide a unified approach to explaining the output of any machine learning model.
  • the collective SHAP values can show how much each predictor contributes, either positively or negatively, to the target variable.
  • a dependence plot shows the marginal effect one or two features have on the predicted outcome by plotting a feature's SHAP values across its domain. It shows whether the relationship between the target and a feature is linear, monotonic or more complex.
  • the chart of SHAP values for TTLR shows the TTLR effect on the traveler’s propensity to redirect from a travel search.
  • the distribution of SHAP values follows a negative trend with increasing TTLR, demonstrating an increasing tendency for travelers to redirect away from a travel search with increasing TTLR, although the functonal relationship is largely non-linear with the exception of values located above the 80 th percentile of the distribution of TTLR values.
  • the first percentile of the distribution of TTLR values corresponds to the shortest TTLR values
  • the 100 th percentile of the distribution of TTLR values corresponds to the longest TTLR values.
  • TTLR values show predominantly positive SHAP values, in which the proportion of negative SHAP values gets smaller as TTLR times get shorter. For TTLR times below four seconds, negative SHAP values are practically non-existent.
  • the fraction of positive SHAP values varies as the TTLR increases. Between TTLR values of 5.3 seconds (the 20 th percentile of the distribution) and 10.1 seconds (the 40 th percentile of the distribution) maximum SHAP values are near to +0.2. Between TTLR values of 1.5 and 5.3 seconds, maximum SHAP values are near to +0.3. For TTLR values lower than 1 second, SHAP values reach all the way up to values of +1 as TTLR tends towards zero seconds.
  • SHAP values are slightly negative on average and are distributed reasonably uniformly.
  • SHAP values are -0.3 and +0.1, respectively.
  • a transition region exists between 10.1 seconds (the 40 th percentile of the distribution) and 14.7 seconds (the 80 th percentile of the distribution) as SHAP values average near to zero, or -0.1 in more precise terms.
  • This transition region comprises forty per cent of observations, from the 40 th percentile of the distribution to the 80 th percentile of the distribution.
  • TTLR times in the transition region are expected to provide a beneficial effect if the TTLR times are reduced. Although the TTLR times in the transition region have slightly negative average SHAP values, it is expected that reducing these TTLR values will often change these SHAP values to positive SHAP values. Therefore for TTLR values in the transition region, the goal is to reduce the TTLR values to as close to 10.1 seconds as possible.
  • TTLR values Twenty per cent of the TTLR values are greater than 14.7 seconds (the 80 th percentile of the distribution) and these TTLR values are associated with larger negative SHAP values.
  • SHAP values for TTLR values greater than 14.7 seconds is particularly uniform, we expect improvements in SHAP values will take place if TTLR values greater than, but close to, 14.7 seconds are reduced.
  • values at the tail of the distribution e.g. as the percentile increases from the 90 th percentile of the distribution, we don't expect to achieve a noticeable increase in the SHAP values, if the TTLR is reduced by a small amount.
  • a client computer device e.g. including a browser or an application program executing on the device, is in connection with the internet.
  • a search home website e.g. an HTML server
  • a front end in the cloud e.g. a content delivery network (CDN), e.g. a static files server
  • a back-end for front-end (BFF) application programming interface (API) server is in connection with the internet.
  • a search service server is in connection with the internet.
  • the client computer device e.g. smartphone
  • the client computer device e.g. including a browser or an application program executing on the device, may be in wireless (e.g. cellular or WiFi) connection with the internet.
  • the client computer device e.g. desktop computer
  • the search home website and the BFF server may be in the same server.
  • search services for hotel reservations search services for rental vehicles (e.g. cars, automobiles, vans, trucks, boats, ships, aircraft, motorbikes), search services for home loans, search services for insurance (e.g. car insurance, home insurance, travel insurance, life insurance, pet insurance) search services for money loans, search services for mortgage loans, search services for mobile phone purchase, search services for broadband deals, search services for real estate, search services for rental accommodation, and search services for dating.
  • search services for hotel reservations search services for rental vehicles (e.g. cars, automobiles, vans, trucks, boats, ships, aircraft, motorbikes)
  • search services for home loans search services for insurance (e.g. car insurance, home insurance, travel insurance, life insurance, pet insurance) search services for money loans
  • search services for mortgage loans search services for mobile phone purchase
  • search services for broadband deals search services for real estate, search services for rental accommodation, and search services for dating.
  • server this may refer to a single server, to a group of servers, or to servers in the cloud, as would be clear to the person skilled in the art.
  • machine learning e.g. neural networks
  • all the machine learning (e.g. neural network) parameters can be randomized with standard methods (such as Xavier Initialization). Typically, we find that satisfactory results are obtained with sufficiently small learning rates.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Business, Economics & Management (AREA)
  • General Physics & Mathematics (AREA)
  • Tourism & Hospitality (AREA)
  • Databases & Information Systems (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Mathematical Physics (AREA)
  • Economics (AREA)
  • General Business, Economics & Management (AREA)
  • Software Systems (AREA)
  • Strategic Management (AREA)
  • Marketing (AREA)
  • Human Resources & Organizations (AREA)
  • Computational Linguistics (AREA)
  • Artificial Intelligence (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Evolutionary Computation (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Primary Health Care (AREA)
  • Computing Systems (AREA)
  • Medical Informatics (AREA)
  • Development Economics (AREA)
  • Entrepreneurship & Innovation (AREA)
  • Operations Research (AREA)
  • Quality & Reliability (AREA)
  • Information Transfer Between Computers (AREA)

Abstract

There is provided a computer-implemented method of stopping an internet search early, the internet search being in a class of internet searches, the method using a trained machine learning model, in which the trained machine learning model has been trained on earlier internet searches in the class, and the successive results produced by those searches, to recognize when sufficient search results have been received to make continuing searching not worthwhile, and to output whether or not to continue searching; the method including inputting the internet search into the trained machine learning model, and the search's present results, and stopping the search in response to the trained machine learning model outputting that the search should be stopped. A corresponding system configured to perform the method is also provided.

Description

COMPUTER-IMPLEMENTED METHODS FOR SEARCHING OVER THE INTERNET, DATABASES, METHODS OF PRODUCING SUCH DATABASES, AND RELATED SYSTEMS AND COMPUTER PROGRAM PRODUCTS
BACKGROUND OF THE INVENTION
1. Field of the Invention
The field of the invention relates to computer-implemented methods for searching over the internet, especially to reducing energy usage in such methods, to related machine learning models, to related training databases for training such machine learning models, and related methods of producing such databases, to related databases for providing input to such trained machine learning models, and related methods of producing such databases, and to related systems, servers, computers and computer program products.
2. Technical Background
With more and more search activities being performed by users of the internet, ever more energy is being used to perform search activities on the internet. It is desirable to reduce, or to limit, the energy usage in respect of performing search activities on the internet. It is also desirable to provide pertinent results more quickly if possible.
A portion of the disclosure of this patent document contains material, which is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the Patent and Trademark Office patent file or records, but otherwise reserves all copyright rights whatsoever.
3. Discussion of Related Art
EP3557437A1 and EP3557437B1 disclose a case management system which is configured to generate search templates based on selection of a search type and one or more data sources. As configured, the case management system enables execution of searches using the generated search template on synchronous and asynchronous data sources and provides periodic polling of the asynchronous data sources to generate consolidated search results.
We refer to the publication by Ayala et al, 2017 IEEE 15th Inti Conf on Dependable, Autonomic and Secure Computing, 15th Inti Conf on Pervasive Intelligence and Computing, 3rd Inti Conf on Big Data Intelligence and Computing and Cyber Science and Technology Congress pages 861 to 866 (DOI: 10.1109/DASC-PICom-DataCom- CyberSciTec.2017.144). In this publication, Ayala et al report substantial and significant energy reductions when polling is reduced. Prior art Figure 8 shows data from Ayala et al, in which for a Galaxy Nexus mobile phone, energy consumption in Joules is significantly and substantially reduced as a function of increasing polling interval in milliseconds, for messages of 3000 bytes. We reason that there is expected to be a corresponding reduction in energy usage by network infrastructure, when energy consumption by a mobile phone in connection with the network infrastructure is reduced.
SUMMARY OF THE INVENTION
According to a first aspect of the invention, there is provided a computer- implemented method of predicting when to stop polling early in a travel search, the method including the steps of
(i) a server receiving a travel search request from a user terminal, the travel search request including travel search parameters;
(ii) the server consulting a database, wherein the database returns an expected total number of search results in response to receiving the travel search request including the travel search parameters from the server;
(iii) polling a plurality of search service servers using the travel search request including the travel search parameters;
(iv) receiving responses from at least one of the plurality of search service servers, and storing the received responses;
(v) processing the stored received responses to determine distinct responses received, and then determining a cumulative number of distinct responses received;
(vi) converting the travel search request including the travel search parameters into a model travel search request including model travel search parameters, and inputting the model travel search request including the model travel search parameters and the determined cumulative number of distinct responses received divided by the expected total number of search results, into a trained machine learning model;
(vii) processing the model travel search request including the model travel search parameters, and the determined cumulative number of distinct responses received divided by the expected total number of search results, through the machine learning model, wherein the machine learning model has been previously trained and comprises a set of learned parameters;
(viii) the machine learning model generating an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by applying the learned parameters to the inputs using the machine learning model;
(ix) proceeding to step (x) if the predetermined time is reached, or if the output indicator indicates that the polling should be stopped; otherwise repeating steps (iii) to (viii); (x) storing the distinct responses received;
(xi) processing the distinct responses received to provide a set of processed distinct responses received;
(xii) the server sending the set of processed distinct responses received to the user terminal.
An advantage is that because polling is stopped early, energy consumption is reduced. An advantage is that because polling is stopped early, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
The method may be one in which steps (vi) to (viii) take less than 100ms.
The method may be one wherein the predetermined time is in the range of 30s to 120s, or wherein the predetermined time is in the range of 45s to 90s, or wherein the predetermined time is in the range of 55s to 65s, or wherein the predetermined time is 60s.
The method may be one wherein even when polling (e.g. from a front end) has been stopped, the search is continued (e.g. in a back end), up to a final predetermined time (e.g. 60s) to obtain and to store quotes search data which can be used in future training of the model.
The method may be one wherein the searches which continue (e.g. in the back end) are a sample of all the searches that are performed, e.g. in the range of 1% to 20% of all the searches that are performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed, so that not all the searches are continued (e.g. in the back end).
The method may be one in which the travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
The method may be one in which the travel search request is a flight search request.
The method may be one in which the flight search request is for a return flight, in which start and destination are specified, outbound and return travel dates are specified, the number of passengers is specified, and cabin class is specified. The method may be one in which when processing the distinct responses received to provide a set of processed distinct responses received, the distinct responses received are processed into itinerary results, wherein each itinerary result includes respective quotes.
The method may be one in which for each respective itinerary result, a respective selectable option is provided on a display screen of the user terminal which is selectable to present the quotes associated with the respective itinerary.
The method may be one in which when the respective selectable option is selected for the respective itinerary, the user terminal requests the quotes for the respective itinerary from the search service server and the search service server returns the quotes for display on the display screen of the user terminal.
The method may be one in which the set of processed distinct responses received are presented in a booking panel.
The method may be one in which the search is resumed if the user indicates a potential willingness to make a booking, such as by clicking on a search result to view more detail.
The method may be one in which the method is used when contacting an external classifier microservice.
The method may be one in which the user terminal is a computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
The method may be one in which the trained machine learning model uses Python or Java.
The method may be one in which the trained machine learning is deployed in all geographic regions in which the search service (e.g. FPS) is present.
The method may be one in which the search service uses a Java client, e.g. one which is already integrated into Itinerary Filtering/Construction.
The method may be one wherein the model travel search parameters derived from the travel search parameters include a query market derived from a destination (e.g. airport).
The method may be one wherein the model travel search parameters derived from the travel search parameters include a booking horizon, which is derived from the departure date.
The method may be one wherein the model travel search parameters derived from the travel search parameters include a trip duration, which is derived from a departure date and a return date.
The method may be one wherein the machine learning model has been trained using a method of any aspect of the second aspect of the invention.
The method may be one wherein the trained machine learning model has been trained using a training database produced by a method of any aspect of the fourth aspect of the invention.
According to a second aspect of the invention, there is provided a computer- implemented method for training a machine learning model, to predict when to stop polling early in a travel search, the method including the steps of:
(i) receiving a plurality of datasets from a database, each dataset comprising a respective model travel search request including respective model travel search parameters, and a respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
(ii) initializing the machine learning model with an initial set of parameters;
(iii) for each dataset, processing the inputs through the machine learning model to generate corresponding predicted outputs;
(iv) calculating a loss function by comparing the corresponding predicted outputs with the corresponding target outputs;
(v) updating the parameters of the machine learning model based on the calculated loss function using an optimization algorithm;
(vi) repeating steps (iii) to (v) until a predefined convergence criterion is met; and
(vii) storing the parameters which correspond to the trained machine learning model which met the predefined convergence criterion in step (vi).
An advantage is that because polling may be stopped early using the trained machine learning model, energy consumption is reduced. An advantage is that because polling may be stopped early using the trained machine learning model, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
The method may be one in which the model travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
The method may be one in which the stored datasets in the database are a sample of all searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed. The method may be one wherein the database is updated after a predefined time interval, e.g. once per week, or once per day, and the machine learning model is correspondingly retrained using the updated database.
The method may be one in which a moving time window is used for the training.
The method may be one in which the number of datasets is at least one thousand, or at least ten thousand, at least one hundred thousand, or at least one million, or at least two million.
The method may be one wherein the trained machine learning model is trained using a training database produced by a method of any of aspect of the fourth aspect of the invention.
According to a third aspect of the invention, there is provided a stored trained machine learning model, produced by a method of any aspect of the second aspect of the invention.
An advantage is that because polling may be stopped early using the trained machine learning model, energy consumption is reduced. An advantage is that because polling may be stopped early using the trained machine learning model, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal. According to a fourth aspect of the invention, there is provided a computer- implemented method of producing and storing a training database for training a machine learning model to predict when to stop polling early in a travel search, the method including the steps of:
(i) receiving a travel search request, the travel search request including travel search parameters;
(ii) polling a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receiving responses from at least one of the plurality of search service servers, and storing the received responses;
(iv) processing the stored received responses to determine a cumulative number of distinct received responses, and storing the determined cumulative number of distinct received responses;
(v) repeating steps (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) storing in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval;
(vii) repeating steps (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) defining a first indicator (e.g. a first Boolean value) which indicates that a search should be continued, and defining a second indicator (e.g. a second Boolean value different to the first Boolean value) different to the first indicator which indicates that a search should be stopped;
(ix) for a respective first dataset, defining a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(x) processing the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including a respective model travel search request including respective model travel search parameters derived from the respective travel search parameters; the respective processed first dataset further including the determined cumulative number of distinct received responses, divided by the respective total number of search results, as a function of successive polling interval; wherein for each stored determined cumulative number of distinct received responses divided by the respective total number of search results, if the value of the cumulative number of distinct received responses divided by the respective total number of search results is less than a threshold, then the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the first indicator, otherwise the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the second indicator;
(xi) repeating steps (ix) to (x) for all the first datasets, to produce respective processed first datasets, and storing the processed first datasets in a database, wherein the stored processed first datasets in the database comprise the training database for training a machine learning model to predict when to stop polling early in a travel search.
The method may be one wherein the predetermined time is in the range of 30s to 120s, or wherein the predetermined time is in the range of 45s to 90s, or wherein the predetermined time is in the range of 55s to 65s, or wherein the predetermined time is 60s.
The method may be one wherein the threshold is a predetermined threshold.
The method may be one wherein the predetermined threshold is in the range 0.40 to 0.95, or wherein the predetermined threshold is in the range 0.60 to 0.90, or wherein the predetermined threshold is in the range 0.70 to 0.85, or wherein the predetermined threshold is 0.8.
The method may be one wherein the threshold is a function of the respective model travel search parameters (e.g. in a simple example it is 0.6 when the target market is UK, or it is 0.8 when the target market is USA).
The method may be one including in step (vi) storing in the first data set the received responses, and in step (x) the respective processed first dataset includes processed received responses derived from the received responses; in which the threshold is a function of the processed received responses. For example the processed received responses may be processed such that the threshold is set so as to include a cheapest result of the processed received responses in order for the second indicator to be stored. For example the processed received responses may be processed to rank the processed received responses, and wherein the threshold is set so as to include a most highly ranked processed received response of the processed received responses in order for the second indicator to be stored.
The method may be one in which the polling intervals are in the range of 0.2s to 50s, or in which the polling intervals are in the range of 0.5s to 45s, or in which the polling intervals are in the range of 1.0s to 30s.
The method may be one in which the stored processed first datasets in the database are a sample of all searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed.
The method may be one wherein the database is updated after a predefined time interval, e.g. once per week, or once per day.
The method may be one in which the number of search requests is at least one thousand, or at least ten thousand, at least one hundred thousand, or at least one million, or at least two million.
The method may be one in which the travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
The method may be one in which the travel search request is a flight search request.
The method may be one in which the flight search request is for a return flight, in which start and destination are specified, outbound and return travel dates are specified, the number of passengers is specified, and cabin class is specified.
The method may be one wherein the respective model travel search parameters derived from the respective travel search parameters include a query market derived from a destination (e.g. airport).
The method may be one wherein the respective model travel search parameters derived from the respective travel search parameters include a booking horizon, which is derived from the departure date.
The method may be one wherein the respective model travel search parameters derived from the respective travel search parameters include a trip duration, which is derived from a departure date and a return date.
The method may be one wherein step (vi) includes deriving and storing an itinerary count from the stored received responses. The method may be one wherein the travel search request is received from a user terminal computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
An advantage is that because polling may be stopped early using the trained machine learning model, energy consumption is reduced. An advantage is that because polling may be stopped early using the trained machine learning model, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
According to a fifth aspect of the invention, there is provided a stored database for training a machine learning model to predict when to stop polling early in a travel search, the database having been produced and stored using a method of any aspect of the fourth aspect of the invention. Advantages includes those of the fourth aspect of the invention.
According to a sixth aspect of the invention, there is provided a computer- implemented method of producing and storing a database for predicting total search results as a function of an input travel search request including input travel search parameters, the method including the steps of:
(i) receiving a travel search request, the travel search request including travel search parameters;
(ii) polling a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receiving responses from at least one of the plurality of search service servers, and storing the received responses;
(iv) processing the stored received responses to determine a cumulative number of distinct received responses, and storing the determined cumulative number of distinct received responses;
(v) repeating steps (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) storing in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval, and a time associated with completing this step;
(vii) repeating steps (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) for a respective first dataset, defining a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(ix) processing the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including the respective travel search request including the respective travel search parameters and the respective total number of search results, and the respective time associated with step (vi);
(x) repeating steps (viii) to (ix) for all the first datasets;
(xi) processing the processed first datasets to produce a database of total number of search results as a function of travel search request including respective travel search parameters, wherein the database entry for a given travel search request including respective travel search parameters is a most recent request for the travel search including respective travel search parameters, wherein the most recent request for the travel search including respective travel search parameters is determined using respective times associated with step (vi).
An advantage is that because polling may be stopped early using the trained machine learning model which uses data from the database as input, energy consumption is reduced. An advantage is that because polling may be stopped early using the trained machine learning model which uses data from the database as input, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
The method may be one in which the travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request. The method may be one wherein the predetermined time is in the range of 30s to 120s, or wherein the predetermined time is in the range of 45s to 90s, or wherein the predetermined time is in the range of 55s to 65s, or wherein the predetermined time is 60s.
The method may be one in which the polling intervals are in the range of 0.2s to 50s, or in which the polling intervals are in the range of 0.5s to 45s, or in which the polling intervals are in the range of 1.0s to 30s.
The method may be one in which the processed first datasets in the database are a sample of all searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed.
The method may be one wherein the database is updated after a predefined time interval, e.g. once per week, or once per day.
The method may be one in which the number of processed first datasets is at least one thousand, or at least ten thousand, at least one hundred thousand, or at least one million, or at least two million.
The method may be one wherein the travel search request is received from a user terminal computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
According to a seventh aspect of the invention, there is provided a stored database for predicting total search results as a function of an input travel search request including input travel search parameters, the database having been produced and stored using a method of any aspect of the sixth aspect of the invention. Advantages include those of the sixth aspect of the invention.
According to an eighth aspect of the invention, there is provided a system including a server and a computer, the computer configured to execute a trained machine learning model to predict when to stop polling early in a travel search, wherein:
(i) the server is configured to receive a travel search request from a user terminal, the travel search request including travel search parameters;
(ii) the server is configured to consult a database, wherein the database returns an expected total number of search results in response to receiving the travel search request including the travel search parameters from the server;
(iii) the server is configured to poll a plurality of search service servers using the travel search request including the travel search parameters;
(iv) the server is configured to receive responses from at least one of the plurality of search service servers, and to store the received responses;
(v) the server is configured to process the stored received responses to determine distinct responses received, and then to determine a cumulative number of distinct responses received;
(vi) the server is configured to convert the travel search request including the travel search parameters into a model travel search request including model travel search parameters, and to input the model travel search request including the model travel search parameters and the determined cumulative number of distinct responses received divided by the expected total number of search results, into the trained machine learning model;
(vii) the computer is configured to execute the trained machine learning model to process the model travel search request including the model travel search parameters, and the determined cumulative number of distinct responses received divided by the expected total number of search results, through the machine learning model, wherein the machine learning model has been previously trained and comprises a set of learned parameters;
(viii) the computer is configured to execute the machine learning model to generate an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by applying the learned parameters to the inputs using the machine learning model;
(ix) proceed to (x) if the predetermined time is reached, or if the output indicator indicates that the polling should be stopped; otherwise repeat (iii) to (viii);
(x) the server is configured to store the distinct responses received;
(xi) the server is configured to process the distinct responses received to provide a set of processed distinct responses received;
(xii) the server is configured to send the set of processed distinct responses received to the user terminal. Advantages include those of the first aspect of the invention.
The system may be configured to perform a method of any aspect of the first aspect of the invention.
According to a ninth aspect of the invention, there is provided a computer configured to train a machine learning model, to predict when to stop polling early in a travel search, wherein the computer is configured to:
(i) receive a plurality of datasets from a database, each dataset comprising a respective model travel search request including respective model travel search parameters, and a respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
(ii) initialize the machine learning model with an initial set of parameters;
(iii) for each dataset, process the inputs through the machine learning model to generate corresponding predicted outputs;
(iv) calculate a loss function by comparing the corresponding predicted outputs with the corresponding target outputs;
(v) update the parameters of the machine learning model based on the calculated loss function using an optimization algorithm;
(vi) repeat (iii) to (v) until a predefined convergence criterion is met; and
(vii) store the parameters which correspond to the trained machine learning model which met the predefined convergence criterion in (vi).
Advantages include those of the second aspect of the invention.
The computer may be configured to perform a method of any aspect of the second aspect of the invention.
According to a tenth aspect of the invention, there is provided a server system configured to produce and to store a training database for training a machine learning model to predict when to stop polling early in a travel search, the server system configured to:
(i) receive a travel search request, the travel search request including travel search parameters;
(ii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receive responses from at least one of the plurality of search service servers, and store the received responses;
(iv) process the stored received responses to determine a cumulative number of distinct received responses, and store the determined cumulative number of distinct received responses;
(v) repeat (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) store in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval;
(vii) repeat (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) define a first indicator (e.g. a first Boolean value) which indicates that a search should be continued, and define a second indicator (e.g. a second Boolean value different to the first Boolean value) different to the first indicator which indicates that a search should be stopped;
(ix) for a respective first dataset, define a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(x) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including a respective model travel search request including respective model travel search parameters derived from the respective travel search parameters; the respective processed first dataset further including the determined cumulative number of distinct received responses, divided by the respective total number of search results, as a function of successive polling interval; wherein for each stored determined cumulative number of distinct received responses divided by the respective total number of search results, if the value of the cumulative number of distinct received responses divided by the respective total number of search results is less than a threshold, then the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the first indicator, otherwise the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the second indicator;
(xi) repeat (ix) to (x) for all the first datasets, to produce respective processed first datasets, and store the processed first datasets in a database, wherein the stored processed first datasets in the database comprise the training database for training a machine learning model to predict when to stop polling early in a travel search.
Advantages include those of the fourth aspect of the invention.
The server system may be configured to perform a method of any aspect of the fourth aspect of the invention.
According to an eleventh aspect of the invention, there is provided a server system configured to produce and to store a database for predicting total search results as a function of an input travel search request including input travel search parameters, the server system configured to:
(i) receive a travel search request, the travel search request including travel search parameters;
(ii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receive responses from at least one of the plurality of search service servers, and store the received responses;
(iv) process the stored received responses to determine a cumulative number of distinct received responses, and store the determined cumulative number of distinct received responses;
(v) repeat (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) store in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval, and a time associated with completing this part;
(vii) repeat (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) for a respective first dataset, define a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(ix) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including the respective travel search request including the respective travel search parameters and the respective total number of search results, and the respective time associated with part (vi);
(x) repeat (viii) to (ix) for all the first datasets;
(xi) process the processed first datasets to produce a database of total number of search results as a function of travel search request including respective travel search parameters, wherein the database entry for a given travel search request including respective travel search parameters is a most recent request for the travel search including respective travel search parameters, wherein the most recent request for the travel search including respective travel search parameters is determined using respective times associated with part (vi).
Advantages include those of the sixth aspect of the invention.
The server system may be configured to perform a method of any aspect of the sixth aspect of the invention.
According to a twelfth aspect of the invention, there is provided a computer program product executable on a server to:
(i) receive a travel search request from a user terminal, the travel search request including travel search parameters;
(ii) consult a database, wherein the database returns an expected total number of search results in response to receiving the travel search request including the travel search parameters from the server; (iii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iv) receive responses from at least one of the plurality of search service servers, and to store the received responses;
(v) process the stored received responses to determine distinct responses received, and then to determine a cumulative number of distinct responses received;
(vi) convert the travel search request including the travel search parameters into a model travel search request including model travel search parameters, and to input the model travel search request including the model travel search parameters and the determined cumulative number of distinct responses received divided by the expected total number of search results, into a trained machine learning model;
(vii) receive from the machine learning model an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by the machine learning model;
(viii) proceed to (ix) if the predetermined time is reached, or if the output indicator indicates that the polling should be stopped; otherwise repeat (iii) to (vii);
(ix) store the distinct responses received;
(x) process the distinct responses received to provide a set of processed distinct responses received;
(xi) send the set of processed distinct responses received to the user terminal.
Advantages include those of the first aspect of the invention.
The computer program product may be executable on the server to perform a method of any aspect of the first aspect of the invention.
According to a thirteenth aspect of the invention, there is provided a computer program product executable on a computer to train a machine learning model, to predict when to stop polling early in a travel search, wherein the computer program product is executable to:
(i) receive a plurality of datasets from a database, each dataset comprising a respective model travel search request including respective model travel search parameters, and a respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
(ii) initialize the machine learning model with an initial set of parameters;
(iii) for each dataset, process the inputs through the machine learning model to generate corresponding predicted outputs;
(iv) calculate a loss function by comparing the corresponding predicted outputs with the corresponding target outputs;
(v) update the parameters of the machine learning model based on the calculated loss function using an optimization algorithm;
(vi) repeat (iii) to (v) until a predefined convergence criterion is met; and
(vii) store the parameters which correspond to the trained machine learning model which met the predefined convergence criterion in (vi).
Advantages include those of the second aspect of the invention.
The computer program product may be executable to perform a method of any aspect of the second aspect of the invention.
According to a fourteenth aspect of the invention, there is provided a computer program product executable on a server system to produce and to store a training database for training a machine learning model to predict when to stop polling early in a travel search, the computer program product executable to:
(i) receive a travel search request, the travel search request including travel search parameters;
(ii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receive responses from at least one of the plurality of search service servers, and store the received responses;
(iv) process the stored received responses to determine a cumulative number of distinct received responses, and store the determined cumulative number of distinct received responses;
(v) repeat (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) store in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval;
(vii) repeat (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) define a first indicator (e.g. a first Boolean value) which indicates that a search should be continued, and define a second indicator (e.g. a second Boolean value different to the first Boolean value) different to the first indicator which indicates that a search should be stopped;
(ix) for a respective first dataset, define a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(x) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including a respective model travel search request including respective model travel search parameters derived from the respective travel search parameters; the respective processed first dataset further including the determined cumulative number of distinct received responses, divided by the respective total number of search results, as a function of successive polling interval; wherein for each stored determined cumulative number of distinct received responses divided by the respective total number of search results, if the value of the cumulative number of distinct received responses divided by the respective total number of search results is less than a threshold, then the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the first indicator, otherwise the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the second indicator;
(xi) repeat (ix) to (x) for all the first datasets, to produce respective processed first datasets, and store the processed first datasets in a database, wherein the stored processed first datasets in the database comprise the training database for training a machine learning model to predict when to stop polling early in a travel search. Advantages include those of the fourth aspect of the invention.
The computer program product may be executable to perform a method of any aspect of the fourth aspect of the invention.
According to a fifteenth aspect of the invention, there is provided a computer program product executable on a server system to produce and to store a database for predicting total search results as a function of an input travel search request including input travel search parameters, the computer program product executable to:
(i) receive a travel search request, the travel search request including travel search parameters;
(ii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receive responses from at least one of the plurality of search service servers, and store the received responses;
(iv) process the stored received responses to determine a cumulative number of distinct received responses, and store the determined cumulative number of distinct received responses;
(v) repeat (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) store in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval, and a time associated with completing this part;
(vii) repeat (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) for a respective first dataset, define a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(ix) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including the respective travel search request including the respective travel search parameters and the respective total number of search results, and the respective time associated with part (vi);
(x) repeat (viii) to (ix) for all the first datasets;
(xi) process the processed first datasets to produce a database of total number of search results as a function of travel search request including respective travel search parameters, wherein the database entry for a given travel search request including respective travel search parameters is a most recent request for the travel search including respective travel search parameters, wherein the most recent request for the travel search including respective travel search parameters is determined using respective times associated with part (vi).
Advantages include those of the sixth aspect of the invention.
The computer program product may be executable to perform a method of any aspect of the sixth aspect of the invention.
According to a sixteenth aspect of the invention, there is provided a computer- implemented method of stopping an internet search early, the internet search being in a class of internet searches, the method using a trained machine learning model, in which the trained machine learning model has been trained on earlier internet searches in the class, and the successive results produced by those searches, to recognize when sufficient search results have been received to make continuing searching not worthwhile, and to output whether or not to continue searching; the method including inputting the internet search into the trained machine learning model, and the search’s present results, and stopping the search in response to the trained machine learning model outputting that the search should be stopped.
An advantage is that because searching is stopped early, energy consumption is reduced. An advantage is that because searching is stopped early, a final set of processed distinct responses is received sooner by the user terminal, compared to always waiting for a predetermined time, before sending the final set of processed distinct responses received to the user terminal.
The method may be one including any aspect of any other aspect of the invention. According to a seventeenth aspect of the invention, there is provided a system configured to perform a method of any aspect of the sixteenth aspect of the invention. Advantages include those of the sixteenth aspect of the invention.
In the above, computer program products may be embodied on a non-transitory storage medium.
Aspects of the invention may be combined.
BRIEF DESCRIPTION OF THE FIGURES
Aspects of the invention will now be described, by way of example(s), with reference to the following Figures, in which:
Figure 1 shows an example within a travel search service, e.g. within a (e.g flight) pricing service (FPS) in which a quote request is stopped early, using a fixed timeout of 15s.
Figure 2 shows an example, within a travel search service, e.g. within a (e.g flight) pricing service (FPS), of a logic within this service for contacting an external classifier microservice, and if that classifier determines that we should stop the session, only then do we stop the session, within a time interval. If the search time is greater than a lower time limit of 12.8s, but less than an upper time limit of 14.5s, then a classifier is used to determine if the search session should be stopped; if the classifier determines that the search should be stopped, then the search is stopped; if the classifier determines that the search should not be stopped, then the search is not stopped.
Figure 3 shows an example of experimentally obtained SHAP values for TTLR, as a function of TTLR. P40 denotes the 40th percentile of the distribution of SHAP values. P80 denotes the 80th percentile of the distribution of SHAP values.
Figure 4 shows an example of first and second itinerary search results displayed at a user terminal.
Figure 5 shows an example of quotes associated with the first itinerary search result of Figure 4 displayed at the user terminal.
Figure 6 shows an example of quotes associated with the second itinerary search result of Figure 4 displayed at the user terminal.
Figure 7 shows an example of a system including a client computer device e.g. including a browser or an application program executing on the device, a search home website, a front end in the Cloud, a search back-end for front-end (BFF) server, and a search service server, which are in connection with the internet.
Figure 8 shows data from Ayala et al, in which for a Galaxy Nexus mobile phone, energy consumption in Joules is reduced as a function of increasing polling interval in milliseconds, for messages of 3000 bytes (This derives from Fig. 3(c) in Ayala et al). DETAILED DESCRIPTION
We describe stopping polling early, e.g. in a travel-based search, e.g. using a classifier approach.
In a travel-based search, in order to reduce the time to last result (TTLR) for travelers, an aim is to predict when the useful results have been returned, and signal that polling can be stopped. The search results obtained up to that point in time may then be loaded, e.g. on a booking panel, which is displayed to a user in a web page, or in an app, on a computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV. By reducing the TTLR, we may reduce the total energy consumed in a search process.
Reducing TTLR is considered to be strongly related to reducing the rate at which users give up on performing a travel search using a travel search website, or in an app, and reducing TTLR is considered to provide a better traveler experience. If we can reduce the time waiting for results that the user will not wait to see, or that the user is unlikely to choose, then we can improve the overall experience for the user. By reducing the TTLR, we may reduce the total energy consumed in a search process.
In a straightforward example, within a (e.g flight) pricing service (FPS) we may stop a quote request early; it is possible to resume the search when the user indicates a potential willingness to make a booking, such as by clicking on a search result to view more detail. In an example, a fixed timeout (e.g. of 15s) is used. If used generally on a travel search website, or if used in an app, or if used in a particular type of search (e.g. for flights, hotels, or rental vehicles), all corresponding search sessions will be curtailed at the fixed timeout (e.g. of 15s) point. An example is shown in Figure 1. In experiments, this has provided positive results, reducing the fraction of cases in which users did not wait for the results to complete, by a statistically significant amount. Unfortunately, we know of searches that do actually require more than a fixed timeout (e.g. of 15s) to return results, and this experiment does not allow for these types of searches, adversely impacting search results coverage and search results price accuracy. In an example, rather than applying a single cut off time to the searching process, we contact a classifier to determine when to stop the searching process. The classifier may be set up to be able to determine if a particular search session is one that should be stopped early, or if the particular search session should be continued through to all results being returned (e.g. in which all partner search services return their results).
In an example, e.g. within a (e.g flight) pricing service (FPS), there is a StopPollingEarlyService, which has the responsibility of identifying if this poll should be stopped early. In an example, within a travel search service, e.g. within a (e.g flight) pricing service (FPS), there is provided a logic within this service to contact an external classifier microservice, and if that classifier determines that we should stop the session, e.g. during a time interval, only then do we stop the session. In an example, if the search time is greater than a lower time limit (e.g. 12.8s), but less than an upper time limit (e.g. 14.5s), then a classifier is used to determine if the search session should be stopped; if the classifier determines that the search should be stopped, then the search is stopped; if the classifier determines that the search should not be stopped, then the search is not stopped. An example is shown in Figure 2.
In an example, if the search time is greater than a lower time limit (e.g. 12.8s), but less than an upper time limit (e.g. 14.5s), then a classifier is used to determine if the search session should be stopped. Why might we use an upper time limit (e.g. 14.5s) that is less than a target cut off time (e.g. 15s)? Our objective is to increase the number of searches that take less than the target cut off time. So if we consider a poll that will be returned after the target cut off time, we've missed our objective target, and in this approach we should continue to allow polling and to complete the search as usual. By only choosing the searches between a lower time limit (e.g. 12.8s) and an upper time limit (e.g. 14.5s) this gives the poll enough time to complete within a search (e.g flight pricing service (FPS)) within the target cut off time (e.g. 15s).
Pre-Computed Results Example Approach
In an example approach to stopping polls early, search results times are obtained in advance, and stored, and a cache is generated which determines when polling should be stopped early, as a function of search parameters. Then when searches are performed, the cache is consulted based on the search parameters of the search, to determine at what time polling should be stopped early, if the search has not yet completed. If the search parameters of the search do not have a good enough match within the cache, then the cache is not used, and polling is not stopped early, or polling is stopped when a target time is reached. In an example, the cache file is fetched periodically and is stored within the memory of the search service, e.g. FPS memory.
Single poll Classifier Example Approach
In an example, the classifier uses heuristics. In an example, the classifier uses a machine learning model. In a machine learning model example, we may use a language more suited to hosting machine learning models e.g. Python, rather than being limited to Java. In a machine learning model example, we use Java. In an example, the classifier is consulted only once per search session. In an example, the machine learning model classifier is consulted only once per search session.
In an example, the classifier is provided as a service, e.g. as a micro-service, e.g. in python. In an example, the service is deployed in all geographic regions that the search service (e.g. FPS) is present; the service receives the current search service (e.g. FPS) state and returns a boolean, e.g. TRUE to stop, FALSE to continue polling.
When the classifier is provided as a service, the service may
• be a reverse proxy infant of individual horizontally scalable instances of each machine learning model, providing some separation.
• use a Java client, e.g. one which is already integrated into Itinerary Filtering/Construction.
• be one in which speed is highly dependent on the speed of the machine learning model. • be one in which 100ms response time is realistic.
Search Service (e.g FPS) interaction with the Relevance Service The Relevance Service may be integrated into itinerary-filtering. The Relevance Service may be integrated into itinerary-construction.
In an example, the model the Relevance Service hosts requires access to the filtered priced itineraries. In an example, the model the Relevance Service hosts requires access to statistics derived from the filtered priced itineraries. In an example, the call to the Relevance Service is made after the call to construction/filtering. In an example, the call to the Relevance Service is made before the state has been written to the State store. In an example, the call to the Relevance Service is a blocking request, so the classifier has tight constraints around latency as it will directly impact the latency of search service (e.g. FPS). In an example, although the call to the Relevance Service may be a single poll, there is no impact on time to first result (TTFR).
Additional Data Required For Classifier
In an example, the classifier may require two new sources of data to be generated. The first is a source of data that can be used for training, generated during the polling process. The second is a source of data emitted a significant time (e.g. 60s) after we stop a poll: this source of data needs to contain the details of what would have been returned had we not stopped early. The two new sources of data can then be used for monitoring our impact on coverage and price accuracy, as well as feedback for the classifier.
In an example, an end result is that when the search service (e.g. FPS) stops a session early we send an event to a worker process, which after a significant time (e.g. 60s) polls the search that was completed early. This gives us two results: we ensure that we're finalising all sessions that we create, which means there will be no impact to the search service (e.g. FPS) cache, and secondly we can determine what would have happened had the search continued a significant time longer (e.g. 60s), which can improve our price accuracy and coverage metrics.
Monitoring the Relevance Service Information display, e.g. dashboards, may be provided showing service metrics, e.g. memory usage, central processing unit (CPU) usage etc. Errors may be monitored from the service itself, but also from the search service’s (e.g. FPS’s) perspective. Circuit breaker status and/or percentage of errors/timeouts may be monitored.
Monitoring the effectiveness of the classifier
In an example, we monitor both the coverage impact and price accuracy of the classifier to ensure that it stays within acceptable thresholds. We make use of the extra event being generated after a significant time (e.g 60s) to compare the full list of itineraries after the significant time to the itineraries that we emitted when we stopped polling early. From the differences between the full list of itineraries after the significant time to the itineraries that we emitted when we stopped polling early, we can determine any itineraries that were completely dropped (a coverage issue) and which itineraries, if any, lost the cheapest pricing option so would have appeared cheaper if polling had not stopped early (a price accuracy issue). These events may be processed at time intervals e.g. in a regularly run data bricks notebook, producing a dashboard we can use to monitor the effectiveness of the Relevance Service algorithm.
Key metrics that may be monitored:
- Number of requests being stopped early
- Number of requests not being stopped early
- Timeouts/Failure rate of classifier service from search service (e.g. FPS) perspective
- Latency of classifier service from search service (e.g. FPS) perspective
Heuristics for deciding if this should be the last poll
In our experiments for travel searches, times for a response to be returned from a poll are as follows. The 50th percentile for the poll response time was 130ms. The 95th percentile for the poll response time was 1.2s. The 99th percentile for the poll response time was 4.6s. In an example, we take the 95th percentile for the poll response time, 1.2s, as a sensible target for cutting off waiting for poll results. Typically the polling interval is 1.0 s.
In an example, we want the final poll to finish before a final time, e.g. 15s. Hence if a poll occurs between (i) the final time minus the 95th percentile for the poll response time minus the polling interval and (ii) the final time minus the 95th percentile for the poll response time, e.g. between (i) 15-1.2-1.0 = 12.8s and (ii) 15-1.2 = 13.8s, then we run the classifier to determine if the poll should be stopped early.
Considerations regarding an Algorithm and Related Data Analysis for Stopping Polls Early
In an example flights search, for a return flight, the start and destination are specified, outbound and return travel dates are specified, the number of passengers is specified, cabin class is specified, etc. The search website polls sources of data for relevant data, and the search website waits for data. For example, the search website can poll for results at regular time intervals, e.g. every second, for example from a system which receives data and constructs itineraries for the user. Then for example the constructed itineraries are sorted and returned to the front end of the website. Or for example the constructed itineraries are returned to the front end of the website where they are sorted and presented. In this pipeline, the algorithm interrupts the process of polling when there is sufficient confidence for this particular search, taking into account the search parameters, that there is insufficient value in continuing polling further. When it is decided to stop polling, in an example a search progress bar is stopped from being presented on a screen to a user who requested the search. In an example, the model has been trained on training data, in which the training data has been selected based on a search time cut off, in which the search time cut off has been determined to be the length of time beyond which users are likely to give up on the search. An example length of time beyond which users are likely to give up on the search is 15s.
In an example, the polling request search time cut off is 14.8s, which leaves 0.2s for final results to be returned, before the search process is terminated at 15.0s. The classifier may be consulted for the previous e.g. one or two pollings (made at 1.0 s intervals), e.g. hence at 12.8s and 13.8s, to determine if polling should be stopped early.
In an example of a flights search, a flights search is requested by a user terminal in communication with a search service server, for flights for one adult in economy class, with the outward flight from London Luton to Tenerife (any airport) on 5 Jan 2025, and the return flight from Tenerife (any airport) to London Luton on 9 Jan 2025, with the flights selected to be restricted to direct flights only. The results are assembled by the search service server and are sent to the user terminal for display on a display screen of the user terminal. In this example, two itineraries are returned. The first itinerary is Easy Jet departing Luton at 0700 and arriving at TFS at 1135 on 5 Jan 2025, and returning Easy Jet from TFS at 1255 to Luton at 1725 on 9 Jan 2025. For the first itinerary, nine quotes are received from partners, hence the result “9 deals” is presented for the first itinerary. The second itinerary is Ryanair departing Luton at 1330 and arriving at TFS at 1750 on 5 Jan 2025, and returning Easy Jet from TFS at 1255 to Luton at 1725 on 9 Jan 2025. For the second itinerary, two quotes are received from partners, hence the result “2 deals” is presented for the second itinerary. An example of search results displayed at the user terminal is shown in Figure 4. Each quote or “deal” is different as it will differ from the other quotes or “deals” e.g. in that the provider is different, or the price is different, or the terms and conditions are different e.g. the baggage allowances are different. In Figure 4 it is stated that 2 of 9 results are shown. The two results shown are the results for direct flights. The other seven results which are not shown are for results involving one stop.
For each respective itinerary, a respective selectable option is provided on the display screen which is selectable to present the quotes, or “deals” associated with the respective itinerary. In an example, when the respective selectable option is selected for the first itinerary, the user terminal requests the quotes for the first itinerary from the search service server and the search service server returns the quotes for display on a display screen of the user terminal. An example is shown in Figure 5, in which the nine quotes or “deals” are listed by provider, namely easyJet, Mytrip, Gotogate, lastminute.com, Booking.com, eDreams, Expedia, Kiwi.com and BudgetAir. In an example, when the respective selectable option is selected for the second itinerary, the user terminal requests the quotes for the second itinerary from the search service server and the search service server returns the quotes for display on a display screen of the user terminal. An example is shown in Figure 6, in which the two quotes or “deals” are listed by provider, namely Ryanair +easyJet, and Kiwi.com.
Loading Classifier Example
In an example, a classifier is consulted to decide when to stop updating displayed search results with new itineraries and options. If we wait for every single source of search results to return results from their Application Programming Interface (API) call, we can potentially wait a long time, e.g. for up to 60 seconds, before completing the displayed search results. Often waiting for this long can result in very little advantage for the traveller so in an example a model, a classifier, is consulted during a specific time window, to decide whether to continue waiting for more results or to stop updating the displayed search results.
In an example, a model has been trained on data from searches, e.g. from small screen devices (e.g. smartphones) based on the state of the search service (e.g. Flights Pricing Service) in a time interval, e.g. 5 - 50 seconds, after the session starts compared to the state of the session, e.g. the search results, when the search completes.
In examples, flight search sessions are stopped early, if the model is called and the model predicts that the session should be stopped early, and if zero, one, or more, or all, of the following set of constraints are met:
• the searching device’s platform is android, or ios, or banana or acorn
• the time of the poll is in between a lower time and an upper time, e.g. between 5 sec and 50 sec
• the trip is one-way or return & the cabin class is economy.
In an example, the model is closed, specific to a single use case.
In an example a travel search request including travel search parameters is converted into a model travel search request including model travel search parameters, for use in querying a trained machine learning model, or for use in generating a training data set (e.g. a database) for training a machine learning model. For example, travel search parameters including a destination airport in Thailand may be converted into model travel search parameters including for example a query market, which is a market code, which here is TH for Thailand, because the destination airport is in Thailand. For example, travel search parameters including a departure date may be converted into model travel search parameters including for example a booking horizon, which is the period of time from the present day to the departure date. For example, travel search parameters including a departure date and a return date may be converted into model travel search parameters including for example a trip duration, which is the period of time from the departure date to the return date.
Example inputs to the trained machine learning model, or for training a machine learning model
An example input variable is a candidate id, which is a candidate id, which is a string, which may or may not be used in the model.
An example input variable is query market, which is a market code eg. TH for Thailand, which is a string, which may be used in the model.
An example input variable is query trip type, which is a search kind eg. ONE WAY, which is a string, which may be used in the model.
An example input variable is query cabin class, which is a cabin class eg ECONOMY, which is a string, which may be used in the model.
An example input variable is query booking horizon, which is a booking horizon e.g. in days eg. 18, which is a number, which may be used in the model. The term "booking horizon" refers to the period of time in advance during which bookings or reservations are being searched for, such as for a flight, hotel room, event, train journey or vehicle rental. This period starts from the current date and extends into the future.
An example input variable is query trip duration, which is a trip duration in days, which is a number, eg. zero for a one-way trip, which may be used in the model.
An example input variable is query adults, which is the number of adults in the search query eg. 1, which is a number, which may be used in the model. An example input variable is stats itineraries itinerary count, which is a count of itineraries in the poll eg. 15, which is a number, which may be used in the model.
An example input variable is stats quote requests current count, which is current quotes eg. 300, which is a number, which may be used in the model.
An example input variable is stats quote requests total count, which is total quotes requested eg. 330, which is a number, which may be used in the model.
Examples of inputs to the model
In an example of input to the model, the input variable stats quote requests current count is the current number of results received from the polling process, during the current search, which is sent to the trained model, for the model to decide if polling should be stopped. So for example, for the current search, if the training data tends to indicate that the number of results that would be received for the present search (or searches similar to it) after 60 s is 300, and at the present time only 100 results have been received (stats quote requests current count =100), then we would tend to expect that the model returns the result to not stop the polling early. In an example, this is irrespective of how long polling has been taking place, provided polling has been taking place for less than a maximum cutoff time (e.g. 60s). And for example, for the current search, if the training data tends to indicate that the number of results that would be received for the present search (or searches similar to it) after 60 s is 300, and at the present time 270 results have been received (stats quote requests current count =270), then we would tend to expect that the model returns the result to stop the polling early. In an example, this is irrespective of how long polling has been taking place, provided polling has been taking place for less than a maximum cutoff time (e.g. 60s).
In an example of input to the model, the input variable stats quote requests total count is a prediction from a database of what the total number of quotes is expected to be, based for example on a search key (e.g. defined as a trip kind (return or one-way trips), route (origin and destination) and inbound and outbound dates). This database can be populated in advance based on stored search results for searches which have been allowed to run for a long time (e.g. 60s) until probably all possible search results have been received. The stored searches in the database which have been allowed to run for a long time (e.g. 60s) may be a sample of all the searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed. For example the number of quotes may be the number of search results, for a given search key. For example the number of quotes may be the number of instances in which providers or partners have provided search results for a given search key. The database may be updated as new stored search results become available, e.g. the database may be updated once per week, or once per day.
In an example of input to the model, the input variable stats quote requests total count is the total number of quotes from providers/partners available for a search key (e.g. defined as a trip kind (return or oneway trips), route (origin and destination) and inbound and outbound dates). In an example, a callable service (e.g. called “WhoToAsk”) maps each search key to the set of partners providing quotes for the search key, e.g. using a database which has been populated in advance based on stored search results for searches which have been allowed to run for a long time (e.g. 60s) until probably all possible search results have been received. The stored searches in the database which have been allowed to run for a long time (e.g. 60s) may be a sample of all the searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed. Typically, the input variable stats quote requests total count doesn’t change during the whole polling process. The database may be updated as new stored search results become available, e.g. the database may be updated once per week, or once per day.
In an example of input/output to/from the model, the input variable stats quote requests current count is the number of quotes from providers/partners returned as the polling process progresses for a given search. For example, let’s assume a polling process lasts for 20 seconds with a 5-second sampling interval: At 5 seconds: Polling returns a Ryanair quote. Then stats quote requests current count will be 1.
At 10 seconds: Polling returns the same Ryanair quote. Then stats quote requests current count will be 1.
At 15 seconds: Polling returns the same Ryanair quote and a Vueling quote. Then stats quote requests current count will be 2.
At 20 seconds: Polling returns a BA quote. Then stats quote requests current count will be 3.
Here, stats quote requests current count counts the number of unique providers/partners quotes returned which is the cumulative count of distinct quotes from partners as the polling process progresses. That’s why on the first two pollings described in the above example stats quote requests current count is 1.
In an example, the model internally builds stats quote requests current count divided by stats quote requests total count to construct a quotient e.g. called frac_partners (which is a fraction of partners). The quotient, e.g. fraction of partners, is a relevant feature (e.g. the most relevant feature) that helps the model decide when to stop a polling process. The model would be expected to stop polling once a large enough fraction of the possible quotes from partners serving that search have been returned.
In an example, the total number of providers/partners quotes available for a search key (e.g. defined as a trip kind (return or one-way trips), route (origin and destination) and inbound and outbound dates) is one hundred. The model may be trained to favour stopping polling when at least 80 percent of the expected providers/partners quotes have been returned as search results, so here when 80 of the 100 providers/partners quotes have been returned as search results it is favoured to stop polling early, and if less than 80 of the 100 providers/partners quotes have been returned as search results then stopping polling early is not favoured.
In an example, on the front end for the particular search key the number of unique providers/partners quotes returned, which is the cumulative count of distinct quotes by partners as the polling process progresses, is counted. A database is queried to provide the total number of quotes by providers/partners available for the particular search key. The front end queries the trained model which includes as input parameters the cumulative count of distinct quotes by partners as the polling process progresses, the total number of quotes by providers/partners available for the particular search key, and the search key, to obtain an output which indicates whether or not the polling should be stopped early. The trained model tends to stop the polling early when the ratio of the cumulative count of distinct partners which have returned search results quotes for the particular search key, to the total number of quotes by providers/partners available for the particular search key, reaches the target ratio, which may be 80%. Even though polling from the front end has been stopped, the search may continue in the back end, e.g. up to a final time (e.g. 60s) to obtain and to store quotes search data which can be used in future training of the model. In an example, the searches which continue in the back end are a sample of all the searches that are performed, e.g. 1% to 20% of all the searches that are performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed, so that not all the searches are continued in the back end.
Example outputs from the model
An example output variable is a candidate id, which is a candidate id, which is a string, which may or may not be used in the model.
An example output variable is prediction, which is a Boolean, in which for example the prediction value can only be 1 (stop session) or 0 (don't stop).
Example Results from the model
In an example, a search is requested for return flights between London, UK and Barcelona, ES on particular dates, and the polling is stopped early after 15s. In an example, a search is requested for return flights between London, UK and Fresno, CA, USA on particular dates, and the polling is stopped early after 35s. Example response codes from the model
Response code 200 has meaning OK.
Response code 404 has meaning Model Not Available.
Response code 412 has meaning Deadline exceeded, which indicates the request to the model has taken too long and a deadline has been exceeded.
In an example, the model is retrained weekly, typically using a new data set compared to the week before. In an example, for the new training data set, a moving time window is used for the training e.g. from one week to the next week. In an example, the model is retrained after a predetermined time interval.
In an example, the number of searches on which the model has been trained is at least one thousand. In an example, the number of searches on which the model has been trained is at least ten thousand. In an example, the number of searches on which the model has been trained is at least one hundred thousand. In an example, the number of searches on which the model has been trained is at least one million. In an example, the number of searches on which the model has been trained is at least two million.
In an example, the number of searches included in the training set on which a model can be trained is at least one thousand. In an example, the number of searches included in the training set on which a model can be trained is at least ten thousand. In an example, the number of searches included in the training set on which a model can be trained is at least one hundred thousand. In an example, the number of searches included in the training set on which a model can be trained is at least one million. In an example, the number of searches included in the training set on which a model can be trained is at least two million.
In an example, the model training procedure follows common approaches, as would be known to those skilled in the art.
The framework can be extended to other applications, such as for car hire, for vehicle hire, for accommodation, for events, for train journeys or for hotels. In website structure, the terms frontend (or sometimes referred to as front end or front-end) and backend (or sometimes referred to as back end or back-end) refer to the separation of concerns between the frontend, e.g. the presentation layer, and the backend, e.g. the data access layer. The terms front end and backend may be used in respect of software. The terms front end and backend may be used in respect of hardware.
Example method of producing and storing a training database for training a machine learning model to predict when to stop polling early in a search (e.g. travel search)
There is provided a computer-implemented method of producing and storing a training database for training a machine learning model to predict when to stop polling early in a search (e.g. travel search), the method including the steps of:
(i) receiving a search (e.g. travel search) request, the search (e.g. travel search) request including search (e.g. travel search) parameters;
(ii) polling a plurality of search service servers using the search (e.g. travel search) request including the search (e.g. travel search) parameters;
(iii) receiving responses from at least one of the plurality of search service servers, and storing the received responses;
(iv) processing the stored received responses to determine a cumulative number of distinct received responses, and storing the determined cumulative number of distinct received responses;
(v) repeating steps (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) storing in a respective first data set the search (e.g. travel search) request including the search (e.g. travel search) parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval;
(vii) repeating steps (i) to (vi) for a plurality of search (e.g. travel search) requests, each respective search (e.g. travel search) request including respective search (e.g. travel search) parameters;
(viii) defining a first indicator (e.g. a first Boolean value) which indicates that a search should be continued, and defining a second indicator (e.g. a second Boolean value different to the first Boolean value) different to the first indicator which indicates that a search should be stopped;
(ix) for a respective first dataset, defining a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(x) processing the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including a respective model search (e.g. travel search) request including respective model search (e.g. travel search) parameters derived from the respective search (e.g. travel search) parameters; the respective processed first dataset further including the determined cumulative number of distinct received responses, divided by the respective total number of search results, as a function of successive polling interval; wherein for each stored determined cumulative number of distinct received responses divided by the respective total number of search results, if the value of the cumulative number of distinct received responses divided by the respective total number of search results is less than a threshold, then the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the first indicator, otherwise the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the second indicator;
(xi) repeating steps (ix) to (x) for all the first datasets, to produce respective processed first datasets, and storing the processed first datasets in a database, wherein the stored processed first datasets in the database comprise the training database for training a machine learning model to predict when to stop polling early in a search (e.g. travel search).
The predetermined time may be in the range of 30s to 120s, or the predetermined time may be in the range of 45s to 90s, or the predetermined time may be in the range of 55s to 65s, or the predetermined time may be 60s.
The threshold may be a predetermined threshold. The predetermined threshold may be in the range 0.40 to 0.95, or the predetermined threshold may be in the range 0.60 to 0.90, or the predetermined threshold may be in the range 0.70 to 0.85, or the predetermined threshold may be 0.8. Example method for training a machine learning model, to predict when to stop polling early in a search (e.g. travel search)
There is provided a computer-implemented method for training a machine learning model, to predict when to stop polling early in a search (e.g. travel search), the method including the steps of:
(i) receiving a plurality of datasets from a database, each dataset comprising a respective model search (e.g. travel search) request including respective model search (e.g. travel search) parameters, and respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
(ii) initializing the machine learning model with an initial set of parameters;
(iii) for each dataset, processing the inputs through the machine learning model to generate corresponding predicted outputs;
(iv) calculating a loss function by comparing the corresponding predicted outputs with the corresponding target outputs;
(v) updating the parameters of the machine learning model based on the calculated loss function using an optimization algorithm;
(vi) repeating steps (iii) to (v) until a predefined convergence criterion is met; and
(vii) storing the parameters which correspond to the trained machine learning model which met the predefined convergence criterion in step (vi).
Example method of predicting when to stop polling early in a search (e.g. travel search)
There is provided a computer implemented method of predicting when to stop polling early in a search (e.g. travel search), the method including the steps of:
(i) a server receiving a search (e.g. travel search) request from a user terminal, the search (e.g. travel search) request including search (e.g. travel search) parameters; (ii) the server consulting a database which returns an expected total number of search results in response to receiving the search (e.g. travel search) request including the search (e.g. travel search) parameters;
(iii) polling a plurality of search service servers using the search (e.g. travel search) request including the search (e.g. travel search) parameters;
(iv) receiving responses from at least one of the plurality of search service servers, and storing the received responses;
(v) processing the stored received responses to determine the distinct responses received, and then determining a cumulative number of distinct received responses;
(vi) converting the search (e.g. travel search) request including the search (e.g. travel search) parameters into a model search (e.g. travel search) request including model search (e.g. travel search) parameters, and inputting the model search (e.g. travel search) request including the model search (e.g. travel search) parameters and the determined cumulative number of distinct received responses divided by the expected total number of search results, into a trained machine learning model;
(vii) processing the model search (e.g. travel search) request including the model search (e.g. travel search) parameters, and the determined cumulative number of distinct received responses divided by the expected total number of search results, through the machine learning model, wherein the machine learning model has been previously trained and comprises a set of learned parameters;
(viii) the machine learning model generating an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by applying the learned parameters to the inputs using the machine learning model;
(ix) proceeding to step (x) if the predetermined time is reached, or if the output indicator indicates that the polling should be stopped; otherwise repeating steps (iii) to (viii);
(x) storing the distinct responses received;
(xi) processing the distinct responses received to provide a set of processed distinct responses received;
(xii) the server sending the set of processed distinct responses received to the user terminal. The predetermined time may be in the range of 30s to 120s, or the predetermined time may be in the range of 45s to 90s, or the predetermined time may be in the range of 55s to 65s, or the predetermined time may be 60s.
Example method of producing and storing a database for predicting total search results as a function of an input search (e.g. travel search) request
There is provided a computer-implemented method of producing and storing a database for predicting total search results as a function of an input search (e.g. travel search) request including input search (e.g. travel search) parameters, the method including the steps of:
(i) receiving a search (e.g. travel search) request, the search (e.g. travel search) request including search (e.g. travel search) parameters;
(ii) polling a plurality of search service servers using the search (e.g. travel search) request including the search (e.g. travel search) parameters;
(iii) receiving responses from at least one of the plurality of search service servers, and storing the received responses;
(iv) processing the stored received responses to determine a cumulative number of distinct received responses, and storing the determined cumulative number of distinct received responses;
(v) repeating steps (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) storing in a respective first data set the search (e.g. travel search) request including the search (e.g. travel search) parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval, and a time associated with completing this step;
(vii) repeating steps (i) to (vi) for a plurality of search (e.g. travel search) requests, each respective search (e.g. travel search) request including respective search (e.g. travel search) parameters;
(viii) for a respective first dataset, defining a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(ix) processing the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including the respective search (e.g. travel search) request including the respective search (e.g. travel search) parameters and the respective total number of search results, and the respective time associated with step (vi);
(x) repeating steps (viii) to (ix) for all the first datasets;
(xi) processing the processed first datasets to produce a database of total number of search results as a function of search (e.g. travel search) request including respective search (e.g. travel search) parameters, wherein the database entry for a given search (e.g. travel search) request including respective search (e.g. travel search) parameters is a most recent request for the search (e.g. travel search) including respective search (e.g. travel search) parameters, wherein the most recent request for the search (e.g. travel search) including respective search (e.g. travel search) parameters is determined using respective times associated with step (vi).
The predetermined time may be in the range of 30s to 120s, or the predetermined time may be in the range of 45s to 90s, or the predetermined time may be in the range of 55s to 65s, or the predetermined time may be 60s.
Speed trade-offs investigation for a travel search service
We have performed investigations to measure trade-off in speed of results, versus usefulness of results to users, for a travel search service.
We used Shapley Additive explanations (SHAP) values to assess latency effects. SHAP values provide a score that measures the impact of a variable/feature on the outcome of a machine learning model. Positive values strengthen the likelihood of the outcome to happen; negative values do the opposite.
The Shapley value provides a principled way to explain the predictions of nonlinear models in the field of machine learning. By interpreting a model trained on a set of features as a value function on a coalition of players, Shapley values provide a natural way to compute which features contribute to a prediction or contribute to the uncertainty of a prediction. SHAP values provide a unified approach to explaining the output of any machine learning model. For the purpose of the analysis we focus on the global interpretability feature for a model whose goal is to predict a redirect in a travel search. The collective SHAP values can show how much each predictor contributes, either positively or negatively, to the target variable. A dependence plot shows the marginal effect one or two features have on the predicted outcome by plotting a feature's SHAP values across its domain. It shows whether the relationship between the target and a feature is linear, monotonic or more complex.
In our analysis of travel search data, the chart of SHAP values for TTLR, as a function of TTLR, shows the TTLR effect on the traveler’s propensity to redirect from a travel search. We found that the distribution of SHAP values follows a negative trend with increasing TTLR, demonstrating an increasing tendency for travelers to redirect away from a travel search with increasing TTLR, although the functonal relationship is largely non-linear with the exception of values located above the 80th percentile of the distribution of TTLR values. In the distribution of TTLR values, the first percentile of the distribution of TTLR values corresponds to the shortest TTLR values, and the 100th percentile of the distribution of TTLR values corresponds to the longest TTLR values.
In the distribution of SHAP values, for TTLR values greater than 14.7 seconds (corresponding to the 80th percentile of the distribution) the SHAP values are mostly negative, with a fairly homogeneous distribution up to 25 seconds.
Observations below the 40th percentile of the distribution of TTLR values show predominantly positive SHAP values, in which the proportion of negative SHAP values gets smaller as TTLR times get shorter. For TTLR times below four seconds, negative SHAP values are practically non-existent. The fraction of positive SHAP values varies as the TTLR increases. Between TTLR values of 5.3 seconds (the 20th percentile of the distribution) and 10.1 seconds (the 40th percentile of the distribution) maximum SHAP values are near to +0.2. Between TTLR values of 1.5 and 5.3 seconds, maximum SHAP values are near to +0.3. For TTLR values lower than 1 second, SHAP values reach all the way up to values of +1 as TTLR tends towards zero seconds.
For TTLR values between 10.1 seconds (the 40th percentile of the distribution) and 14.7 seconds (the 80th percentile of the distribution) SHAP values are slightly negative on average and are distributed reasonably uniformly. For TTLR values between 10.1 seconds (the 40th percentile of the distribution) and 14.7 seconds (the 80th percentile of the distribution), the lower and upper boundaries of the SHAP values are -0.3 and +0.1, respectively.
We conclude that shortening the time-to-last-result has a favourable effect, which is to reduce a user's tendency to leave the search. The most favourable effects are for TTLR times shorter than 10.1 seconds (the 40th percentile of the distribution). As the number of positive SHAP values abruptly overtakes the number of negative SHAP values as TTLR falls below 10.1 seconds, for practical purposes, we can define the 10.1 seconds mark as the best upper limit for time-to-last-result values.
A transition region exists between 10.1 seconds (the 40th percentile of the distribution) and 14.7 seconds (the 80th percentile of the distribution) as SHAP values average near to zero, or -0.1 in more precise terms. This transition region comprises forty per cent of observations, from the 40th percentile of the distribution to the 80th percentile of the distribution. TTLR times in the transition region are expected to provide a beneficial effect if the TTLR times are reduced. Although the TTLR times in the transition region have slightly negative average SHAP values, it is expected that reducing these TTLR values will often change these SHAP values to positive SHAP values. Therefore for TTLR values in the transition region, the goal is to reduce the TTLR values to as close to 10.1 seconds as possible.
Twenty per cent of the TTLR values are greater than 14.7 seconds (the 80th percentile of the distribution) and these TTLR values are associated with larger negative SHAP values. As the distribution of SHAP values for TTLR values greater than 14.7 seconds is particularly uniform, we expect improvements in SHAP values will take place if TTLR values greater than, but close to, 14.7 seconds are reduced. For values at the tail of the distribution, e.g. as the percentile increases from the 90th percentile of the distribution, we don't expect to achieve a noticeable increase in the SHAP values, if the TTLR is reduced by a small amount.
An example of SHAP values for TTLR, as a function of TTLR, is shown in Figure 3.
Example System
In an example system, a client computer device, e.g. including a browser or an application program executing on the device, is in connection with the internet. A search home website (e.g. an HTML server) is in connection with the internet. A front end in the cloud (e.g. a content delivery network (CDN), e.g. a static files server) is in connection with the internet. A back-end for front-end (BFF) application programming interface (API) server is in connection with the internet. A search service server is in connection with the internet. The client computer device (e.g. smartphone), e.g. including a browser or an application program executing on the device, may be in wireless (e.g. cellular or WiFi) connection with the internet. The client computer device (e.g. desktop computer), e.g. including a browser or an application program executing on the device, may be in wired connection with the internet. An example is shown in Figure 7. In an example, the search home website and the BFF server may be in the same server.
OTHER SEARCH SERVICES
Although in the above we have emphasized applications relating to search services for flights, other applications are possible, such as search services for hotel reservations, search services for rental vehicles (e.g. cars, automobiles, vans, trucks, boats, ships, aircraft, motorbikes), search services for home loans, search services for insurance (e.g. car insurance, home insurance, travel insurance, life insurance, pet insurance) search services for money loans, search services for mortgage loans, search services for mobile phone purchase, search services for broadband deals, search services for real estate, search services for rental accommodation, and search services for dating. SERVERS
Where the term “server” is used, this may refer to a single server, to a group of servers, or to servers in the cloud, as would be clear to the person skilled in the art.
Notes re Training
Regarding seeding the machine learning (e.g. neural networks) for training, all the machine learning (e.g. neural network) parameters can be randomized with standard methods (such as Xavier Initialization). Typically, we find that satisfactory results are obtained with sufficiently small learning rates.
Note
It is to be understood that the above-referenced arrangements are only illustrative of the application for the principles of the present invention. Numerous modifications and alternative arrangements can be devised without departing from the spirit and scope of the present invention. While the present invention has been shown in the drawings and fully described above with particularity and detail in connection with what is presently deemed to be the most practical and preferred example(s) of the invention, it will be apparent to those of ordinary skill in the art that numerous modifications can be made without departing from the principles and concepts of the invention as set forth herein.

Claims

1. A computer-implemented method of predicting when to stop polling early in a travel search, the method including the steps of:
(i) a server receiving a travel search request from a user terminal, the travel search request including travel search parameters;
(ii) the server consulting a database, wherein the database returns an expected total number of search results in response to receiving the travel search request including the travel search parameters from the server;
(iii) polling a plurality of search service servers using the travel search request including the travel search parameters;
(iv) receiving responses from at least one of the plurality of search service servers, and storing the received responses;
(v) processing the stored received responses to determine distinct responses received, and then determining a cumulative number of distinct responses received;
(vi) converting the travel search request including the travel search parameters into a model travel search request including model travel search parameters, and inputting the model travel search request including the model travel search parameters and the determined cumulative number of distinct responses received divided by the expected total number of search results, into a trained machine learning model;
(vii) processing the model travel search request including the model travel search parameters, and the determined cumulative number of distinct responses received divided by the expected total number of search results, through the machine learning model, wherein the machine learning model has been previously trained and comprises a set of learned parameters;
(viii) the machine learning model generating an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by applying the learned parameters to the inputs using the machine learning model;
(ix) proceeding to step (x) if the predetermined time is reached, or if the output indicator indicates that the polling should be stopped; otherwise repeating steps (iii) to (viii);
(x) storing the distinct responses received; (xi) processing the distinct responses received to provide a set of processed distinct responses received;
(xii) the server sending the set of processed distinct responses received to the user terminal.
2. The method of Claim 1, in which steps (vi) to (viii) take less than 100ms.
3. The method of Claims 1 or 2, wherein the predetermined time is in the range of 30s to 120s, or wherein the predetermined time is in the range of 45s to 90s, or wherein the predetermined time is in the range of 55s to 65s, or wherein the predetermined time is 60s.
4. The method of any previous Claim, wherein even when polling (e.g. from a front end) has been stopped, the search is continued (e.g. in a back end), up to a final predetermined time (e.g. 60s) to obtain and to store quotes search data which can be used in future training of the model.
5. The method of Claim 4, wherein the searches which continue (e.g. in the back end) are a sample of all the searches that are performed, e.g. in the range of 1% to 20% of all the searches that are performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed, so that not all the searches are continued (e.g. in the back end).
6. The method of any previous Claim, in which the travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
7. The method of Claim 6, in which the travel search request is a flight search request.
8. The method of Claim 7, in which the flight search request is for a return flight, in which start and destination are specified, outbound and return travel dates are specified, the number of passengers is specified, and cabin class is specified.
9. The method of any previous Claim, in which when processing the distinct responses received to provide a set of processed distinct responses received, the distinct responses received are processed into itinerary results, wherein each itinerary result includes respective quotes.
10. The method of Claim 9, in which for each respective itinerary result, a respective selectable option is provided on a display screen of the user terminal which is selectable to present the quotes associated with the respective itinerary.
11. The method of Claim 10, in which when the respective selectable option is selected for the respective itinerary, the user terminal requests the quotes for the respective itinerary from the search service server and the search service server returns the quotes for display on the display screen of the user terminal.
12. The method of Claim 11, in which the set of processed distinct responses received are presented in a booking panel.
13. The method of any previous Claim, in which the search is resumed if the user indicates a potential willingness to make a booking, such as by clicking on a search result to view more detail.
14. The method of any previous Claim, in which the method is used when contacting an external classifier microservice.
15. The method of any previous Claim, in which the user terminal is a computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
16. The method of any previous Claim, in which the trained machine learning model uses Python or Java.
17. The method of any previous Claim, in which the trained machine learning is deployed in all geographic regions in which the search service (e.g. FPS) is present.
18. The method of any previous Claim, in which the search service uses a Java client, e.g. one which is already integrated into Itinerary Filtering/Construction.
19. The method of any previous Claim, wherein the model travel search parameters derived from the travel search parameters include a query market derived from a destination (e.g. airport).
20. The method of any previous Claim, wherein the model travel search parameters derived from the travel search parameters include a booking horizon, which is derived from the departure date.
21. The method of any previous Claim, wherein the model travel search parameters derived from the travel search parameters include a trip duration, which is derived from a departure date and a return date.
22. The method of any previous Claim, wherein the machine learning model has been trained using a method of any of Claims 24 to 30.
23. The method of Claim 22, wherein the trained machine learning model has been trained using a training database produced by a method of any of Claims 32 to 46.
24. A computer-implemented method for training a machine learning model, to predict when to stop polling early in a travel search, the method including the steps of:
(i) receiving a plurality of datasets from a database, each dataset comprising a respective model travel search request including respective model travel search parameters, and a respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
(ii) initializing the machine learning model with an initial set of parameters;
(iii) for each dataset, processing the inputs through the machine learning model to generate corresponding predicted outputs;
(iv) calculating a loss function by comparing the corresponding predicted outputs with the corresponding target outputs;
(v) updating the parameters of the machine learning model based on the calculated loss function using an optimization algorithm;
(vi) repeating steps (iii) to (v) until a predefined convergence criterion is met; and
(vii) storing the parameters which correspond to the trained machine learning model which met the predefined convergence criterion in step (vi).
25. The method of Claim 24, in which the model travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
26. The method of Claims 24 or 25, in which the stored datasets in the database are a sample of all searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed.
27. The method of any of Claims 24 to 26, wherein the database is updated after a predefined time interval, e.g. once per week, or once per day, and the machine learning model is correspondingly retrained using the updated database.
28. The method of any of Claims 24 to 27, in which a moving time window is used for the training.
29. The method of any of Claims 24 to 28, in which the number of datasets is at least one thousand, or at least ten thousand, at least one hundred thousand, or at least one million, or at least two million.
30. The method of any of Claims 24 to 29, wherein the trained machine learning model is trained using a training database produced by a method of any of Claims 32 to 46.
31. A stored trained machine learning model, produced by a method of any of Claims 24 to 30.
32. A computer-implemented method of producing and storing a training database for training a machine learning model to predict when to stop polling early in a travel search, the method including the steps of:
(i) receiving a travel search request, the travel search request including travel search parameters;
(ii) polling a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receiving responses from at least one of the plurality of search service servers, and storing the received responses;
(iv) processing the stored received responses to determine a cumulative number of distinct received responses, and storing the determined cumulative number of distinct received responses;
(v) repeating steps (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) storing in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval;
(vii) repeating steps (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) defining a first indicator (e.g. a first Boolean value) which indicates that a search should be continued, and defining a second indicator (e.g. a second Boolean value different to the first Boolean value) different to the first indicator which indicates that a search should be stopped;
(ix) for a respective first dataset, defining a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(x) processing the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including a respective model travel search request including respective model travel search parameters derived from the respective travel search parameters; the respective processed first dataset further including the determined cumulative number of distinct received responses, divided by the respective total number of search results, as a function of successive polling interval; wherein for each stored determined cumulative number of distinct received responses divided by the respective total number of search results, if the value of the cumulative number of distinct received responses divided by the respective total number of search results is less than a threshold, then the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the first indicator, otherwise the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the second indicator;
(xi) repeating steps (ix) to (x) for all the first datasets, to produce respective processed first datasets, and storing the processed first datasets in a database, wherein the stored processed first datasets in the database comprise the training database for training a machine learning model to predict when to stop polling early in a travel search.
33. The method of Claim 32, wherein the predetermined time is in the range of 30s to 120s, or wherein the predetermined time is in the range of 45s to 90s, or wherein the predetermined time is in the range of 55s to 65s, or wherein the predetermined time is 60s.
34. The method of Claims 32 or 33, wherein the threshold is a predetermined threshold.
35. The method of Claim 34, wherein the predetermined threshold is in the range 0.40 to 0.95, or wherein the predetermined threshold is in the range 0.60 to 0.90, or wherein the predetermined threshold is in the range 0.70 to 0.85, or wherein the predetermined threshold is 0.8.
36. The method of Claims 32 or 33, wherein the threshold is a function of the respective model travel search parameters (e.g. in a simple example it is 0.6 when the target market is UK, or it is 0.8 when the target market is USA).
37. The method of Claims 32, 33 or 36, including in step (vi) storing in the first data set the received responses, and in step (x) the respective processed first dataset includes processed received responses derived from the received responses; in which the threshold is a function of the processed received responses.
38. The method of Claim 37 in which the processed received responses are processed such that the threshold is set so as to include a cheapest result of the processed received responses in order for the second indicator to be stored.
39. The method of Claims 37 or 38 in which the processed received responses are processed to rank the processed received responses, and wherein the threshold is set so as to include a most highly ranked processed received response of the processed received responses in order for the second indicator to be stored.
40. The method of any of Claims 32 to 39, in which the polling intervals are in the range of 0.2s to 50s, or in which the polling intervals are in the range of 0.5s to 45s, or in which the polling intervals are in the range of 1.0s to 30s.
41. The method of any of Claims 32 to 40, in which the stored processed first datasets in the database are a sample of all searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed.
42. The method of any of Claims 32 to 41, wherein the database is updated after a predefined time interval, e.g. once per week, or once per day.
43. The method of any of Claims 32 to 42, in which the number of search requests is at least one thousand, or at least ten thousand, at least one hundred thousand, or at least one million, or at least two million.
44. The method of any of Claims 32 to 43, in which the travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
45. The method of Claim 44, in which the travel search request is a flight search request.
46. The method of Claim 45, in which the flight search request is for a return flight, in which start and destination are specified, outbound and return travel dates are specified, the number of passengers is specified, and cabin class is specified.
47. The method of any of Claims 32 to 46, wherein the respective model travel search parameters derived from the respective travel search parameters include a query market derived from a destination (e.g. airport).
48. The method of any of Claims 32 to 47, wherein the respective model travel search parameters derived from the respective travel search parameters include a booking horizon, which is derived from the departure date.
49. The method of any of Claims 32 to 48, wherein the respective model travel search parameters derived from the respective travel search parameters include a trip duration, which is derived from a departure date and a return date.
50. The method of any of Claims 32 to 49, wherein step (vi) includes deriving and storing an itinerary count from the stored received responses.
51. The method of any of Claims 32 to 50, wherein the travel search request is received from a user terminal computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
52. A stored database for training a machine learning model to predict when to stop polling early in a travel search, the database having been produced and stored using a method of any of Claims 32 to 51.
53. A computer-implemented method of producing and storing a database for predicting total search results as a function of an input travel search request including input travel search parameters, the method including the steps of:
(i) receiving a travel search request, the travel search request including travel search parameters;
(ii) polling a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receiving responses from at least one of the plurality of search service servers, and storing the received responses; (iv) processing the stored received responses to determine a cumulative number of distinct received responses, and storing the determined cumulative number of distinct received responses;
(v) repeating steps (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) storing in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval, and a time associated with completing this step;
(vii) repeating steps (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) for a respective first dataset, defining a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(ix) processing the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including the respective travel search request including the respective travel search parameters and the respective total number of search results, and the respective time associated with step (vi);
(x) repeating steps (viii) to (ix) for all the first datasets;
(xi) processing the processed first datasets to produce a database of total number of search results as a function of travel search request including respective travel search parameters, wherein the database entry for a given travel search request including respective travel search parameters is a most recent request for the travel search including respective travel search parameters, wherein the most recent request for the travel search including respective travel search parameters is determined using respective times associated with step (vi).
54. The method of Claim 53, in which the travel search request is a flight search request, or a hotel search request, or a vehicle search request, or a train journey search request, or an event search request.
55. The method of Claims 53 or 54, wherein the predetermined time is in the range of 30s to 120s, or wherein the predetermined time is in the range of 45s to 90s, or wherein the predetermined time is in the range of 55s to 65s, or wherein the predetermined time is 60s.
56. The method of any of Claims 53 to 55, in which the polling intervals are in the range of 0.2s to 50s, or in which the polling intervals are in the range of 0.5s to 45s, or in which the polling intervals are in the range of 1.0s to 30s.
57. The method of any of Claims 53 to 56, in which the processed first datasets in the database are a sample of all searches performed, e.g. in the range of 1% to 20% of all the searches that have been performed, or in the range of 1% to 50% of all the searches that have been performed, or in the range of 1% to 70% of all the searches that have been performed.
58. The method of any of Claims 53 to 57, wherein the database is updated after a predefined time interval, e.g. once per week, or once per day.
59. The method of any of Claims 53 to 58, in which the number of processed first datasets is at least one thousand, or at least ten thousand, at least one hundred thousand, or at least one million, or at least two million.
60. The method of any of Claims 53 to 59, wherein the travel search request is received from a user terminal computing device e.g. smartphone, tablet computer, desktop computer, laptop computer, or smart TV.
61. A stored database for predicting total search results as a function of an input travel search request including input travel search parameters, the database having been produced and stored using a method of any of Claims 53 to 60.
62. A system including a server and a computer, the computer configured to execute a trained machine learning model to predict when to stop polling early in a travel search, wherein:
(i) the server is configured to receive a travel search request from a user terminal, the travel search request including travel search parameters;
(ii) the server is configured to consult a database, wherein the database returns an expected total number of search results in response to receiving the travel search request including the travel search parameters from the server;
(iii) the server is configured to poll a plurality of search service servers using the travel search request including the travel search parameters;
(iv) the server is configured to receive responses from at least one of the plurality of search service servers, and to store the received responses; (v) the server is configured to process the stored received responses to determine distinct responses received, and then to determine a cumulative number of distinct responses received;
(vi) the server is configured to convert the travel search request including the travel search parameters into a model travel search request including model travel search parameters, and to input the model travel search request including the model travel search parameters and the determined cumulative number of distinct responses received divided by the expected total number of search results, into the trained machine learning model;
(vii) the computer is configured to execute the trained machine learning model to process the model travel search request including the model travel search parameters, and the determined cumulative number of distinct responses received divided by the expected total number of search results, through the machine learning model, wherein the machine learning model has been previously trained and comprises a set of learned parameters;
(viii) the computer is configured to execute the machine learning model to generate an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by applying the learned parameters to the inputs using the machine learning model;
(ix) proceed to (x) if the predetermined time is reached, or if the output indicator indicates that the polling should be stopped; otherwise repeat (iii) to (viii);
(x) the server is configured to store the distinct responses received;
(xi) the server is configured to process the distinct responses received to provide a set of processed distinct responses received;
(xii) the server is configured to send the set of processed distinct responses received to the user terminal.
63. The system of Claim 62, configured to perform a method of any of Claims 1 to 23.
64. A computer configured to train a machine learning model, to predict when to stop polling early in a travel search, wherein the computer is configured to:
(i) receive a plurality of datasets from a database, each dataset comprising a respective model travel search request including respective model travel search parameters, and a respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
(ii) initialize the machine learning model with an initial set of parameters;
(iii) for each dataset, process the inputs through the machine learning model to generate corresponding predicted outputs;
(iv) calculate a loss function by comparing the corresponding predicted outputs with the corresponding target outputs;
(v) update the parameters of the machine learning model based on the calculated loss function using an optimization algorithm;
(vi) repeat (iii) to (v) until a predefined convergence criterion is met; and
(vii) store the parameters which correspond to the trained machine learning model which met the predefined convergence criterion in (vi).
65. The computer of Claim 64, configured to perform a method of any of Claims 24 to 30.
66. A server system configured to produce and to store a training database for training a machine learning model to predict when to stop polling early in a travel search, the server system configured to:
(i) receive a travel search request, the travel search request including travel search parameters;
(ii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receive responses from at least one of the plurality of search service servers, and store the received responses;
(iv) process the stored received responses to determine a cumulative number of distinct received responses, and store the determined cumulative number of distinct received responses; (v) repeat (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) store in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval;
(vii) repeat (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) define a first indicator (e.g. a first Boolean value) which indicates that a search should be continued, and define a second indicator (e.g. a second Boolean value different to the first Boolean value) different to the first indicator which indicates that a search should be stopped;
(ix) for a respective first dataset, define a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(x) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including a respective model travel search request including respective model travel search parameters derived from the respective travel search parameters; the respective processed first dataset further including the determined cumulative number of distinct received responses, divided by the respective total number of search results, as a function of successive polling interval; wherein for each stored determined cumulative number of distinct received responses divided by the respective total number of search results, if the value of the cumulative number of distinct received responses divided by the respective total number of search results is less than a threshold, then the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the first indicator, otherwise the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the second indicator;
(xi) repeat (ix) to (x) for all the first datasets, to produce respective processed first datasets, and store the processed first datasets in a database, wherein the stored processed first datasets in the database comprise the training database for training a machine learning model to predict when to stop polling early in a travel search.
67. The server system of Claim 66, configured to perform a method of any of Claims 32 to 51.
68. A server system configured to produce and to store a database for predicting total search results as a function of an input travel search request including input travel search parameters, the server system configured to:
(i) receive a travel search request, the travel search request including travel search parameters;
(ii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receive responses from at least one of the plurality of search service servers, and store the received responses;
(iv) process the stored received responses to determine a cumulative number of distinct received responses, and store the determined cumulative number of distinct received responses;
(v) repeat (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) store in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval, and a time associated with completing this part;
(vii) repeat (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) for a respective first dataset, define a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(ix) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including the respective travel search request including the respective travel search parameters and the respective total number of search results, and the respective time associated with part (vi);
(x) repeat (viii) to (ix) for all the first datasets;
(xi) process the processed first datasets to produce a database of total number of search results as a function of travel search request including respective travel search parameters, wherein the database entry for a given travel search request including respective travel search parameters is a most recent request for the travel search including respective travel search parameters, wherein the most recent request for the travel search including respective travel search parameters is determined using respective times associated with part (vi).
69. The server system of Claim 68, configured to perform a method of any of Claims 53 to 60.
70. A computer program product executable on a server to:
(i) receive a travel search request from a user terminal, the travel search request including travel search parameters;
(ii) consult a database, wherein the database returns an expected total number of search results in response to receiving the travel search request including the travel search parameters from the server;
(iii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iv) receive responses from at least one of the plurality of search service servers, and to store the received responses;
(v) process the stored received responses to determine distinct responses received, and then to determine a cumulative number of distinct responses received;
(vi) convert the travel search request including the travel search parameters into a model travel search request including model travel search parameters, and to input the model travel search request including the model travel search parameters and the determined cumulative number of distinct responses received divided by the expected total number of search results, into a trained machine learning model;
(vii) receive from the machine learning model an output indicator based on the inputs, wherein the output indicates if the polling should be stopped before a predetermined time, or not, wherein the output is determined by the machine learning model;
(viii) proceed to (ix) if the predetermined time is reached, or if the output indicator indicates that the polling should be stopped; otherwise repeat (iii) to (vii);
(ix) store the distinct responses received;
(x) process the distinct responses received to provide a set of processed distinct responses received;
(xi) send the set of processed distinct responses received to the user terminal.
71. The computer program product of Claim 70 executable on the server to perform a method of any of Claims 1 to 23.
72. A computer program product executable on a computer to train a machine learning model, to predict when to stop polling early in a travel search, wherein the computer program product is executable to:
(i) receive a plurality of datasets from a database, each dataset comprising a respective model travel search request including respective model travel search parameters, and a respective determined cumulative number of distinct received search request responses, divided by a respective total number of search results, as a function of successive polling interval, as inputs, and target outputs comprising a corresponding indicator as a target output for each determined cumulative number of distinct received search request responses divided by the respective total number of search results;
(ii) initialize the machine learning model with an initial set of parameters;
(iii) for each dataset, process the inputs through the machine learning model to generate corresponding predicted outputs;
(iv) calculate a loss function by comparing the corresponding predicted outputs with the corresponding target outputs;
(v) update the parameters of the machine learning model based on the calculated loss function using an optimization algorithm;
(vi) repeat (iii) to (v) until a predefined convergence criterion is met; and
(vii) store the parameters which correspond to the trained machine learning model which met the predefined convergence criterion in (vi).
73. The computer program product of Claim 72, executable to perform a method of any of Claims 24 to 30.
74. A computer program product executable on a server system to produce and to store a training database for training a machine learning model to predict when to stop polling early in a travel search, the computer program product executable to:
(i) receive a travel search request, the travel search request including travel search parameters;
(ii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receive responses from at least one of the plurality of search service servers, and store the received responses;
(iv) process the stored received responses to determine a cumulative number of distinct received responses, and store the determined cumulative number of distinct received responses;
(v) repeat (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) store in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval;
(vii) repeat (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) define a first indicator (e.g. a first Boolean value) which indicates that a search should be continued, and define a second indicator (e.g. a second Boolean value different to the first Boolean value) different to the first indicator which indicates that a search should be stopped;
(ix) for a respective first dataset, define a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(x) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including a respective model travel search request including respective model travel search parameters derived from the respective travel search parameters; the respective processed first dataset further including the determined cumulative number of distinct received responses, divided by the respective total number of search results, as a function of successive polling interval; wherein for each stored determined cumulative number of distinct received responses divided by the respective total number of search results, if the value of the cumulative number of distinct received responses divided by the respective total number of search results is less than a threshold, then the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the first indicator, otherwise the cumulative number of distinct received responses divided by the respective total number of search results is stored together with the second indicator;
(xi) repeat (ix) to (x) for all the first datasets, to produce respective processed first datasets, and store the processed first datasets in a database, wherein the stored processed first datasets in the database comprise the training database for training a machine learning model to predict when to stop polling early in a travel search.
75. The computer program product of Claim 74, executable to perform a method of any of Claims 32 to 51.
76. A computer program product executable on a server system to produce and to store a database for predicting total search results as a function of an input travel search request including input travel search parameters, the computer program product executable to:
(i) receive a travel search request, the travel search request including travel search parameters;
(ii) poll a plurality of search service servers using the travel search request including the travel search parameters;
(iii) receive responses from at least one of the plurality of search service servers, and store the received responses;
(iv) process the stored received responses to determine a cumulative number of distinct received responses, and store the determined cumulative number of distinct received responses;
(v) repeat (ii) to (iv) at successive polling intervals until a predetermined time is reached;
(vi) store in a respective first data set the travel search request including the travel search parameters, and the determined cumulative number of distinct received responses as a function of successive polling interval, and a time associated with completing this part;
(vii) repeat (i) to (vi) for a plurality of travel search requests, each respective travel search request including respective travel search parameters;
(viii) for a respective first dataset, define a respective total number of search results as the determined cumulative number of distinct received responses when the predetermined time was reached;
(ix) process the respective first dataset to produce a respective processed first dataset, the respective processed first dataset including the respective travel search request including the respective travel search parameters and the respective total number of search results, and the respective time associated with part (vi);
(x) repeat (viii) to (ix) for all the first datasets;
(xi) process the processed first datasets to produce a database of total number of search results as a function of travel search request including respective travel search parameters, wherein the database entry for a given travel search request including respective travel search parameters is a most recent request for the travel search including respective travel search parameters, wherein the most recent request for the travel search including respective travel search parameters is determined using respective times associated with part (vi).
77. The computer program product of Claim 76, executable to perform a method of any of Claims 53 to 60.
78. Computer-implemented method of stopping an internet search early, the internet search being in a class of internet searches, the method using a trained machine learning model, in which the trained machine learning model has been trained on earlier internet searches in the class, and the successive results produced by those searches, to recognize when sufficient search results have been received to make continuing searching not worthwhile, and to output whether or not to continue searching; the method including inputting the internet search into the trained machine learning model, and the search’s present results, and stopping the search in response to the trained machine learning model outputting that the search should be stopped.
79. The method of Claim 78, including the method of any previous Claim.
80. System configured to perform the method of Claims 78 or 79.
EP24827102.5A 2024-06-10 2024-12-05 Computer-implemented methods for searching over the internet, databases, methods of producing such databases, and related systems and computer program products Pending EP4689919A1 (en)

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
GBGB2408265.3A GB202408265D0 (en) 2024-06-10 2024-06-10 Method
GBGB2413724.2A GB202413724D0 (en) 2024-09-18 2024-09-18 Method
PCT/GB2024/053037 WO2025257520A1 (en) 2024-06-10 2024-12-05 Computer-implemented methods for searching over the internet, databases, methods of producing such databases, and related systems and computer program products

Publications (1)

Publication Number Publication Date
EP4689919A1 true EP4689919A1 (en) 2026-02-11

Family

ID=93925139

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24827102.5A Pending EP4689919A1 (en) 2024-06-10 2024-12-05 Computer-implemented methods for searching over the internet, databases, methods of producing such databases, and related systems and computer program products

Country Status (2)

Country Link
EP (1) EP4689919A1 (en)
WO (1) WO2025257520A1 (en)

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
AU2014363194C1 (en) * 2013-12-11 2021-01-07 Skyscanner Technology Limited Method and server for providing fare availabilities, such as air fare availabilities
US10726032B2 (en) 2015-12-30 2020-07-28 Palantir Technologies, Inc. Systems and methods for search template generation
CA3249529A1 (en) * 2022-05-20 2023-11-23 Amadeus S.A.S. Computational load mitigation with predictive search request enrichment

Also Published As

Publication number Publication date
WO2025257520A1 (en) 2025-12-18

Similar Documents

Publication Publication Date Title
US12299612B2 (en) Optimizing engagement of transportation providers
US12405961B1 (en) Query completions
CN112819576A (en) Training method and device for charging station recommendation model and recommendation method for charging station
CN110741367A (en) Method and apparatus for real-time interactive recommendation
US11580584B2 (en) Managing item queries
JP2016514325A (en) Managing item queries
US20260023819A1 (en) Task code recommendation model
US20190310888A1 (en) Allocating Resources in Response to Estimated Completion Times for Requests
US20160055505A1 (en) Price elasticity testing
CN116011904A (en) Matching method and device, storage medium, and electronic equipment for transport objects and goods
KR101676219B1 (en) Recommendation engine for interactive search forms
US12574705B2 (en) Transmitting digital transportation requests across modes to limited-eligibility provider devices to improve network coverage and system efficiency
US11206236B2 (en) Systems and methods to prioritize chat rooms using machine learning
EP4689919A1 (en) Computer-implemented methods for searching over the internet, databases, methods of producing such databases, and related systems and computer program products
CN114840559A (en) Travel product query and model training method, device, equipment and storage medium
US20220270126A1 (en) Reinforcement Learning Method For Incentive Policy Based On Historic Data Trajectory Construction
EP4361913A1 (en) Vehicle sharing service optimization
US20210034612A1 (en) Controlling generation of multi-input search results
WO2023241388A1 (en) Model training method and apparatus, energy replenishment intention recognition method and apparatus, device, and medium
CN117278508A (en) Recommended methods, devices and electronic devices for 5G messaging chatbots
US12299609B1 (en) Dynamically transmitting online mode invitations to provider devices in response to detected changes in provider device efficiency
CN109218411A (en) Data processing method and device, computer readable storage medium, electronic equipment
CN114510658B (en) Method, device, server and terminal device for determining recommended boarding point
US20250209561A1 (en) Context-aware matching of provider devices and requester devices based on shared contextual characteristics
CN117227581A (en) Electric automobile charging reminding method and device, electronic equipment and storage medium

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: EXAMINATION IS IN PROGRESS

17P Request for examination filed

Effective date: 20251107

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

17Q First examination report despatched

Effective date: 20260121