WO2023211536A1 - Table visualization based on intelligent conditional formatting - Google Patents

Table visualization based on intelligent conditional formatting Download PDF

Info

Publication number
WO2023211536A1
WO2023211536A1 PCT/US2023/012655 US2023012655W WO2023211536A1 WO 2023211536 A1 WO2023211536 A1 WO 2023211536A1 US 2023012655 W US2023012655 W US 2023012655W WO 2023211536 A1 WO2023211536 A1 WO 2023211536A1
Authority
WO
WIPO (PCT)
Prior art keywords
field
cell
representation
formatting
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2023/012655
Other languages
French (fr)
Inventor
Wei Ji
Mengyu ZHOU
Shi Han
Yining Chen
Daxin Jiang
Dongmei Zhang
Shasha GAO
Fei TENG
Yuelin Zhang
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Microsoft Technology Licensing LLC
Original Assignee
Microsoft Technology Licensing LLC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Microsoft Technology Licensing LLC filed Critical Microsoft Technology Licensing LLC
Publication of WO2023211536A1 publication Critical patent/WO2023211536A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/34Browsing; Visualisation therefor
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/166Editing, e.g. inserting or deleting
    • G06F40/177Editing, e.g. inserting or deleting of tables; using ruled lines
    • G06F40/18Editing, e.g. inserting or deleting of tables; using ruled lines of spreadsheets
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/33Querying
    • G06F16/338Presentation of query results
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/103Formatting, i.e. changing of presentation of documents

Definitions

  • search engines may help users easily acquire information on the network. For example, a user may enter a query into a search box of a search engine. Based on the query, the search engine may retrieve documents relevant to the query from a pre-established index database, and then perform subsequent processing on these documents, e.g., relevance filtering, ranking, etc., and finally present the highest-ranked series of documents to the user through a search result page.
  • Embodiments of the present disclosure propose a methods, apparatus and computer program product for table visualization based on intelligent conditional formatting.
  • a table may be obtained, the table containing a plurality of fields. At least one field representation of at least one field in the plurality of fields may be generated. Conditional formatting corresponding to the field may be automatically recommended based at least on the field representation.
  • the table may be visualized through formatting the field based on the conditional formatting.
  • FIG.l illustrates an exemplary process for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
  • FIG.2 illustrates an exemplary table and an exemplary visualized table corresponding to the table according to an embodiment of the present disclosure.
  • FIG.3 illustrates another exemplary table and an exemplary visualized table corresponding to the table according to an embodiment of the present disclosure.
  • FIG.4 illustrates an exemplary process for automatically recommending conditional formatting according to an embodiment of the present disclosure.
  • FIG.5 is a flowchart of an exemplary method for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
  • FIG.6 illustrates an exemplary apparatus for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
  • FIG.7 illustrates an exemplary apparatus for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
  • a search result page associated with a search engine may typically include relevant information for respective documents.
  • a user may quickly learn about a specific document through viewing information for that document, and decide whether he/she wants to click a corresponding link for deeper knowledge.
  • the information presented on the search result page may include a title, an abstract, a picture, a link, etc. Additionally, in the case where the document includes a table, the table may also be presented on the search result page.
  • tables presented on search result pages are usually extracted directly from documents, which may be in plain text format. For example, a table presented on a search result page may contain only raw text or numbers without any formatting modifications.
  • Embodiments of the present disclosure propose table visualization based on intelligent conditional formatting.
  • Conditional formatting is an analysis tool for tabular data, which may help users select data on a table, and format the selected data. Existing conditional formatting requires a user to provide conditions for selecting data and a format for the selected data.
  • the intelligent conditional formatting according to the embodiments of the present disclosure may automatically recommend at least one conditional formatting corresponding to at least one table field in a table.
  • a column in the vertical direction in a table may be referred to as a table field, or simply a field.
  • Conditional formatting corresponding to a particular field may include, e.g., an operation appropriate for that field and/or a parameter set corresponding to that operation.
  • the recommended conditional formatting may be automatically applied to the corresponding table field for table visualization.
  • the table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure may selectively format the table in an intelligent manner. This process may be performed automatically, and may be widely applied to various types of tables, e.g., local tables, tables retrieved through search engines, etc., without requiring to specially set conditional formatting or determine formatting data for specific tables.
  • an embodiment of the present disclosure proposes to automatically recommend conditional formatting through a machine learning model.
  • a machine learning model used to automatically recommend conditional formatting may be referred to as an Analytical Semantics over Tabular Data Engine.
  • the analytical semantics over tabular data engine may generate at least one field representation of at least one field of the plurality of fields contained in the table, and automatically recommend conditional formatting corresponding to the field based at least on the generated field representation.
  • a field in a table for which conditional formatting is recommended may be referred to as a target field.
  • the analytical semantics over tabular data engine may infer user motivations corresponding to a target field, e.g., user intent, data focus, etc., based on a field representation of the target field.
  • User intent may refer to the intent that motivates the user to create an analysis of the table, e.g. the type of analysis the user might want to perform.
  • the user intent may include, e.g., data comparison, data detection, etc.
  • the data comparison may refer to a user comparing a plurality of sets of cells through making the plurality of sets of cells have different formatting.
  • the data detection may refer to a user selecting a portion of data from a given table field, and highlighting the selected data through a predefined formatting.
  • Data focus may refer to important data features that should be paid attention to in data reference.
  • the data focus may include, e.g., a false value, a null value, a meaningless value, an empirical value, etc.
  • the user intent and the data focus may be collectively referred to as analytical semantics.
  • the analytical semantics over tabular data engine may further recommend conditional formatting corresponding to the target field based on the user intent, the data focus, and the field representation of the target field.
  • the user intent and the data focus may serve as expert knowledge to guide the analytical semantics over tabular data engine to give more accurate recommendation for conditional formatting of the target field, and may make a recommendation result interpretable and easy for humans to understand.
  • an embodiment of the present disclosure proposes to generate a field representation of a target field from a plurality of structural levels and based on different feature types. For example, a cell-level representation and a field-level representation of the target field may be generated, respectively, and a field representation of the target field may be generated based on the generated cell-level representation and field-level representation.
  • the cell-level representation may be generated from a cell-level statistical feature and linguistic feature.
  • the field-level representation may be generated from a field-level statistical feature and linguistic feature.
  • the cell-level statistical feature and the field-level statistical feature may reflect the data distribution characteristics of corresponding levels.
  • the cell-level linguistic feature and the fieldlevel linguistic feature may reflect linguistic or semantic characteristics of corresponding levels. In this way, a field representation containing rich and complete information can be obtained, so that a more accurate recommendation result can be obtained when conditional formatting recommendation is performed subsequently.
  • an embodiment of the present disclosure proposes to automatically recommend a plurality of types of conditional formatting.
  • Operations included in conditional formatting may include e.g., adding an icon, adding a data graph, adjusting a font style, etc.
  • the icon may include, e.g., an image or text composed of graphic and/or text with a sense of design, e.g., a national flag of a country, a logo of a company, a team flag of a team, etc.
  • the data graph may include, e.g., a data bar graph, a data pie graph, etc.
  • Adjusting the font style may include, e.g., changing a font color, bolding the font, italicizing the font, underlining the font, etc.
  • an operation appropriate for the field may be adding a set of icons corresponding to the set of entities; when a set of cells in a target field contains a set of numbers, an operation appropriate for the field may be adding a set of data graphs corresponding to the set of numbers; when some cells in a target field have different characteristics from other cells, an operation appropriate for the field may be adjusting a font style of these cells to highlight these cells; etc.
  • Some operations may have corresponding parameter sets.
  • the recommended conditional formatting may be automatically applied to the corresponding table field for table visualization. Compared to the original table, the visualized table contains richer visual information, thus embodying stronger entity or numerical properties.
  • FIG.l illustrates an exemplary process 100 for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
  • a table 102 may be obtained.
  • the table 102 may be a local table located at a terminal device, e.g., a table authored by a spreadsheet authoring tool at the terminal device. Additionally, the table 102 may be a table retrieved through a search engine. For example, a query may be entered into a search engine, and search results for the query may be obtained from the search engine. The search results may be obtained by the search engine through performing operations such as searching, ranking, etc., for the query. Some search results may contain tables.
  • the table 102 such as a local table, a table retrieved through a search engine, etc., may be a table in plain text format.
  • the table 102 may be a multi-dimension data set containing a plurality of fields.
  • FIG.2 illustrates an exemplary table 200a and an exemplary visualized table 200b corresponding to the table 200a according to an embodiment of the present disclosure.
  • the table 200a may be a table in plain text format.
  • the table 200a may contain 4 fields, i.e., a field 202a, a field 204a, a field 206a, and a field 208a.
  • the field 202a may include a set of rankings.
  • the field 204a may include a set of company names.
  • the field 206a may include a set of sales corresponding to the set of company names in the field 204a.
  • the field 208a may include a set of market shares corresponding to the set of company names in the field 204a.
  • the table 200b may be an exemplary visualized table corresponding to the table 200a.
  • the table 200b may contain 4 fields, i.e., a field 202b, a field 204b, a field 206b, and a field 208b.
  • the field 202b, the field 204b, the field 206b, and the field 208b may correspond to the field 202a, the field 204a, the field 206a, and the field 208a in the table 200a, respectively.
  • the table 102 may be provided to an analytical semantics over tabular data engine 104.
  • the analytical semantics over tabular data engine 104 may be a machine learning model that may automatically recommend at least one conditional formatting 106 corresponding to at least one field in the table 102.
  • the analytical semantics over tabular data engine 104 may generate at least one field representation of at least one of the plurality of fields contained in the table 102, and automatically recommend conditional formatting corresponding to the field based at least on the generated field representation.
  • the analytical semantics over tabular data engine 104 may infer user motivations corresponding to a target field, e.g., user intent, data focus, etc., based on a field representation of the target field.
  • the user intent may include, e.g., data comparison, data detection, etc.
  • the data focus may include, e.g., a false value, a null value, a meaningless value, an empirical value, etc.
  • the user intent and the data focus may be collectively referred to as analytical semantics.
  • the analytical semantics over tabular data engine 104 may further recommend conditional formatting corresponding to the target field based on the user intent, the data focus, and the field representation of the target field.
  • the user intent and the data focus may serve as expert knowledge to guide the analytical semantics over tabular data engine 104 to give more accurate recommendation for conditional formatting, and may make the recommendation result interpretable and easy for humans to understand. An exemplary process for automatically recommending conditional formatting will be described later in conjunction with FIG.4.
  • Conditional formatting for a target field may include an operation appropriate for that field and/or a parameter set.
  • Operations recommended by the analytical semantics over tabular data engine 104 may include, e.g., adding an icon 106-1, adding a data graph 106-2, adjusting a font style 106- 3, etc.
  • the icon may include, e.g., a national flag of a country, a logo of a company, a team flag of a team, etc.
  • the data graph may include, e.g., a data bar graph, a data pie graph, etc.
  • Adjusting the font style may include, e.g., changing a font color, bolding the font, italicizing the font, underlining the font, etc.
  • the table 102 may be visualized through formatting the corresponding fields based on the recommended conditional formatting 106, thereby obtaining a visualized table 112. For example, for a target field, formatted data for the field may be determined through a formatting data enveloper 108 based on the recommended conditional formatting 106. The field may then be rendered by a formatting Tenderer 110 based on the formatting data determined by the formatting data enveloper 108, thereby rendering a field with the corresponding format. For example, when the table 102 is a table retrieved through a search engine, the visualized table 112 may be presented on a search result page by the formatting Tenderer 110 using a programming language such as Hypertext Markup Language (HTML).
  • HTML Hypertext Markup Language
  • the formatting data enveloper 108 may include a plurality of units corresponding to different conditional formatting, e.g., an icon linking unit 108-1, a data graph populating unit 108-2, a font style determining unit 108-3, etc.
  • the icon linking unit 108-1 in the formatting data enveloper 108 may determine an icon corresponding to each cell in the field. For each cell in the field, an entity corresponding to a value in the cell may be identified.
  • a value in a cell may refer to content in the cell, which may include, e.g., number, text, etc.
  • the entity corresponding to the value in the cell may be identified by known entity recognition techniques.
  • an icon corresponding to the identified entity may be extracted from a knowledge graph through known entity linking techniques.
  • the knowledge graph may be, e.g., a knowledge graph corresponding to the category of the identified entity.
  • the extracted icon may be added at the cell.
  • the field 204a may contain a set of company names.
  • the operation included in the conditional formatting corresponding to the field 204a recommended by the analytical semantics over tabular data engine 104 may be adding an icon.
  • the formatting data enveloper 108 may determine a set of logos corresponding to a set of cells in the field 204a. The determined set of logos may be added to the corresponding cells, thereby obtaining the field 204b.
  • the field 204b contains richer visual information, e.g., icons corresponding to the entities in the field 204b, thereby embodying stronger entity properties.
  • the data graph populating unit 108-2 in the formatting data enveloper 108 may, based on a predetermined rule, populate a set of data graphs corresponding to a set of cells in the field.
  • the conditional formatting may also include a parameter set corresponding to the operation, e.g., a global value.
  • a global value may be, e.g., a statistical value of values in a set of cells in a field, e.g., a maximum value, a minimum value, an average value, etc.
  • the global value may be, e.g., a predetermined value related to the values in the set of cells in the field.
  • the global value may be the total global population.
  • a proportion of a prominent portion in a data graph to the data graph may be calculated based on a value in the cell and the global value.
  • the prominent portion in the data graph may be indicated through, e.g., shading, different colors, etc.
  • the length of a shaded bar in the data bar graph may be calculated based on a value in the cell and the global value.
  • the angle of a sector in the data pie graph may be calculated based on a value in the cell and the global value. Subsequently, the data graph may be populated based on the calculated proportion. The populated data graphs may be added to the corresponding cells.
  • the field 206a may contain a set of sales corresponding to a set of countries in the field 204.
  • the operation included in the conditional formatting corresponding to the field 206a recommended by the analytical semantics over tabular data engine 104 may be adding a data bar graph, and the corresponding parameter may be the maximum value in the cells in field 206a, i.e., " 17,098,242" in the cell 210a.
  • a proportion of a prominent portion in a data bar graph to the data bar graph may be calculated based on a value in the cell and the maximum value "17,098,242".
  • the prominent portion in the data bar graph may be indicated through shading.
  • the populated data bar graphs may be added to the corresponding cells in the field 206a, thereby obtaining the field 206b.
  • the field 206b contains richer visual information, e.g., the data bar graphs that may visually present relative magnitude of the values in respective cells, thereby embodying stronger numerical properties.
  • the field 208a may include a set of market shares corresponding to the set of company names in the field 204a.
  • the operation included in the conditional formatting corresponding to the field 208a recommended by the analytical semantics over tabular data engine 104 may be adding a data pie graph, and the corresponding parameter may be the total sales of a similar product in the market. Therefore, for each cell in the field 208a, a proportion of a prominent portion in a data pie graph to the data pie graph may be calculated based on a value in the cell and the total sales of a similar product in the market. The prominent portion in the data pie graph may be indicated through a darker color.
  • the populated data pie graphs may be added to the corresponding cells in the field 208a, thereby obtaining the field 208b.
  • the field 208b contains richer visual information, e.g., the data pie graphs those may visually present relative proportions of the values in respective cells to total sales of a similar product in the market, thereby embodying stronger numerical properties.
  • the font style determining unit 108-3 in the formatting data enveloper 108 may determine a font style of values of cells in the field based on a predetermined rule.
  • the conditional formatting may further include a parameter set corresponding to the operation, e.g., at least one threshold for determining a partition in which each cell in the field is located. For each cell of the field, a partition corresponding to the cell may be determined based on a threshold. Subsequently, a font style of a value in the cell may be adjusted based on the determined partition. In particular, a partition may be highlighted through adjusting its font style to make it significantly different from other partitions.
  • a parameter set corresponding to the operation e.g., at least one threshold for determining a partition in which each cell in the field is located.
  • a partition corresponding to the cell may be determined based on a threshold.
  • a font style of a value in the cell may be adjusted based on the determined partition.
  • a partition may be highlighted through adjusting its font style to make it significantly different from other partitions.
  • FIG.3 illustrates another exemplary table 300a and an exemplary visualized table 300b corresponding to the table 300a according to an embodiment of the present disclosure.
  • the table 300a may contain 4 fields, i.e., a field 302a, a field 304a, a field 306a, and a field 308a.
  • the field 302a may include a set of majors.
  • the field 304a may include a set of enrolments in 2020 corresponding to the set of majors in the field 302a.
  • the field 306a may include a set of enrolments in 2021 corresponding to the set of majors in the field 302a.
  • the field 308a may include a set of enrolment increases in 2021 corresponding to the set of majors in the field 302a.
  • the operation included in the conditional formatting corresponding to the field 308a recommended by the analytical semantics over tabular data engine 104 may be adjusting a font style, and the corresponding parameter may be a threshold "0".
  • the table 300a may be visualized through formatting the field 308a based on the conditional formatting, thereby obtaining a visualized table 300b.
  • the table 300b may contain 4 fields, i.e., a field 302b, a field 304b, a field 306b, and a field 308b.
  • the field 302b, the field 304b, the field 306b, and the field 308b may correspond to the field 302a, the field 304a, the field 306a, and the field 308a in the table 300a, respectively.
  • a partition corresponding to the cell may be determined based on the threshold value "0". For example, when a value in a cell is greater than or equal to the threshold "0", i.e., when the increase is positive or zero, the cell may be determined to be in the first partition; and when the value in the cell is less than the threshold "0", i.e., when the increase is negative, the cell may be determined to be in the second partition.
  • Different font styles may be set for cells in the first partition and cells in the second partition in the field 308a, respectively, thereby obtaining the field 308b.
  • their font styles may remain unchanged, thereby obtaining a cell 310b, a cell 316b, and a cell 318b; and for cells in the second partition, e.g., a cell 312a and a cell 314a, their font styles may be adjusted to bold, thereby obtaining a cell 312b and a cell 314b.
  • a data bar graph is also shown in the field 308b.
  • the length, position, and form of a data bar graph in each cell may correspond to a value in the cell.
  • the field 308b contains richer visual information, e.g., displaying numerical values with a downward trend in bold font, and more intuitively presenting the relative magnitudes and symbols of the values in respective cells with the data bar graphs, thereby embodying stronger numerical properties.
  • process 100 selective formatting may be performed on the table 102 in an intelligent way.
  • the process 100 may be performed automatically.
  • the process 100 may be widely applied to various types of tables, e.g., local tables, tables retrieved through search engines, etc., without requiring to specially set conditional formatting or determine formatting data for specific tables.
  • the process for table visualization based on intelligent conditional formatting described above in conjunction with FIGs.l to 3 is merely exemplary. Depending on actual application requirements, the steps in the process for table visualization based on intelligent conditional formatting may be replaced or modified in any manner, and the process may include more or fewer steps. For example, although only three types of operations including adding an icon, adding a data graph, and adjusting a font style are shown in the conditional formatting 106, and the formatting data enveloper 108 also only includes three types of formatting data determining units corresponding to these three types of operations, but the embodiments of the present disclosure are not limited thereto.
  • the analytical semantics over tabular data engine 104 may also recommend other types of operations, and the formatting data enveloper 108 may accordingly include other types of formatting data determining units.
  • the formatting shown in FIGs.2 and 3 are merely exemplary. Depending on actual application requirements, the table may be formatted in any other form.
  • FIG.4 illustrates an exemplary process 400 for automatically recommending conditional formatting according to an embodiment of the present disclosure.
  • the process 400 may correspond to the operation at the analytical semantics over tabular data engine 104 in FIG. l.
  • At least one conditional formatting corresponding to at least one field in a table 402 may be automatically recommended through the process 400.
  • the process 400 is described below by taking any field in the table 402 as an example.
  • a field for which the process 400 is directed may be referred to as a target field.
  • a field representation of the target field may first be generated. Subsequently, conditional formatting corresponding to the target field may be automatically recommended based at least on the field representation of the target field.
  • the field representation of the target field may be generated from a plurality of structural levels and based on different feature types.
  • a cell-level representation and a field-level representation of the target field may be generated, respectively, and the field representation of the target field may be generated based on the generated cell-level representation and field-level representation.
  • the cell-level representation may be generated from a cell-level statistical feature and linguistic feature.
  • the field-level representation may be generated from a field-level statistical feature and linguistic feature.
  • the cell-level statistical feature and the field-level statistical feature may reflect the data distribution characteristics of corresponding levels.
  • the cell-level linguistic feature and the fieldlevel linguistic feature may reflect linguistic or semantic characteristics of corresponding levels. In this way, a field representation containing rich and complete information can be obtained, so that a more accurate recommendation result can be obtained when conditional formatting recommendation is performed subsequently.
  • the table 402 may be a multi-dimension data set containing a plurality of fields.
  • the table 402 may be denoted as fl) , where n > 2 is used to denote the number of fields contained in the table 402.
  • the target field in the table 402 may be denoted as /j T (l ⁇ i ⁇ ri).
  • a table/field-level input 406 Tab t may be extracted from the table 402, where Cellsi E Tab t .
  • the table/field-level input 406 may contain information of all fields in the table 402, and may contain two-dimensional structure information of the table 402.
  • a cell-level representation of the target field fl may be generated based on a statistical feature and a linguistic feature corresponding to the cell set 404.
  • the statistical feature corresponding to the cell set 404 may reflect data statistical information corresponding to each cell in the cell set 404.
  • Data statistical information corresponding to a specific cell may include, e.g., the ranking of a value in the cell in the field in which it is located, whether the cell contains a null value, etc.
  • the linguistic feature corresponding to the cell set 404 may reflect linguistic or semantic characteristics corresponding to each cell in the cell set 404.
  • a cell subset 418 Cellsi' of the target field fl may be first sampled from the cell set 404 of the target field fl. For example, when the number of cells included in the cell set 404 exceeds a predetermined threshold, the cell subset 418 Cells ⁇ may be sampled from the cell set 404 first. The cell-level representation of the target field fl may then be generated based on a statistical feature and a linguistic feature corresponding to the cell subset 418 Cells-.
  • the cell subset 418 Cellsi' may be sampled from the cell set 404 based on a statistical feature corresponding to the cell set 404 and/or a statistical feature corresponding to the target field fl .
  • the statistical feature corresponding to the cell set 404 may be characterized as a cell signature 410 of the target field fl .
  • the cell signature 410 of the target field fl may be calculated based on the cell set 404 through a cell signature calculating unit 408.
  • the statistical feature corresponding to the target field ft may reflect data distribution characteristics of the target field ff .
  • the statistical feature corresponding to the target field ff may be characterized as a field signature 414 of the target field ff .
  • the field signature 414 of the target field ff may be calculated based on the table/field-level input 406 through a field signature calculating unit 412.
  • the field signature 414 may be denoted as F .
  • a cell sampling unit 416 may sample the cell subset 418 Cells' from the cell set 404 based on the cell signature 410 and/or the field signature 414. Through sampling the cell set 404, the number of cells provided to a subsequent model can be reduced, and the size of a candidate pool used to generate parameters can be reduced, which help to improve the efficiency and performance of the model.
  • the cell signature 410 may be updated to obtain an updated cell signature 420.
  • the updated cell signature 420 may correspond to the cell subset 418 Cells ⁇ .
  • the updated cell signature 420 may be denoted as . It should be appreciated that when the number of cells included in the cell set 404 does not exceed a predetermined threshold, the sampling operation may not be performed. In this case, the cell subset 418 may be consistent with the cell set 404, and the updated cell signature 420 may be consistent with the cell signature 410.
  • the cell subset 418 Cells may be provided to a Pre-trained Language Model (PLM) 422.
  • the pre-trained language model 422 may generate a linguistic feature corresponding to the cell subset 418 Cells .
  • the pre-trained language model 422 may be, e.g., a Bidirectional Encoder Representations from Transformers (BERT) model, a Generative Pre-trained Transformer (GPT) model, etc.
  • the pre-trained language model 422 may generate a cell-level linguistic representation 424 CLR of the target field ft based on the cell subset 418 Cells'.
  • the pre-trained language model 422 may regard content of each cell in the cell subset 418 Cellst' as a sentence sample. For each cell, an embedding for the starting token of the cell output by the pre-trained language model 422 may be regarded as a linguistic representation of the cell.
  • the process described above may be as shown by the following equation:
  • CLR t PLM(Cellsf) (1) where CLR t E JR fcxe , b is the number of cells in the cell subset 418 Cells ⁇ ', and e is the embedding size of the pre-trained language model 422.
  • a cell-level representation 428 MCR t of the target field ff may be generated through a combining unit 426.
  • the combining unit 426 may first concatenate the updated cell signature 420 fP and the cell-level linguistic representation 424 CLR t .
  • the concatenated representation may then be transformed with a linear layer and an activation function.
  • the process described above may be as shown by the following equation: where MCR t G IR z ’ xD , and D is the embedding size of a field encoder 450 used subsequently.
  • a field-level representation of the target field ff may be generated based on a statistical feature and a linguistic feature corresponding to the target field f[.
  • the table/field level input 406 may be provided to a Pre-trained Tabular Model (PTM) 430.
  • the pre-trained tabular model 430 may generate a linguistic feature corresponding to the target field f[ .
  • the pre-trained tabular model 430 may be, e.g., a Tabular Information Embedding (TABBIE) model, a Table Parser (TAPAS) model, etc.
  • the pre-trained tabular model 430 may generate a field-level linguistic representation 432 FLRt of the target field ff based on the table/field-level input 406 containing two-dimensional structural information of the table 102.
  • a predetermined token e.g. [CLS]
  • CLS a predetermined token
  • the process through which the pre-trained tabular model 430 generates the field-level linguistic representation 432 FLRt may be represented by the following equation:
  • FLRt PTM(Tabt) (3) where FLRt E IR lx£ , and E is the embedding size of the pre-trained tabular model 430.
  • the fieldlevel linguistic representation 432 FLRt may contain overall tabular context information of the table 102 and information of the target field ff itself.
  • a field-level representation 436 MFRt of the target field ff may be generated through a combining unit 434.
  • the combining unit 434 may first concatenate the field signature 414 F and the field-level representation 432 FLRt.
  • the concatenated representation may then be transformed with a linear layer and an activation function.
  • the process described above may be as shown by the following equation: Through the processing of the combining unit 426 and the combining unit 434, the cell-level representation 428 MCR t and the field-level representation 436 MFR t may be located in a same feature space.
  • a merged representation 442 MSf of the target field ff may be generated based on the cell-level representation 428 MCR t and the field-level representation 436 MFR t of the target field ff through a combining unit 440.
  • the combining unit 440 may append the cell-level representation 428 MCR t to the field-level representation 436 MFR t , to obtain the merged representation 442 MSf.
  • the process described above may be as shown by the following equation:
  • MSf MFRi MCRi (5) where MSf E IR z ’ xD .
  • a token type representation 444 TE t corresponding to the merged representation 442 MSf may be obtained.
  • the token type representation 444 TE t may indicate the type to which each embedding in the merged representation 442 MSf corresponds, including, e.g., a field-level representation, a cell-level representation, etc.
  • a field encoder 450 may generate a field representation 452 of the target field ff based on the merged representation 442 MSf and the token type representation 444 TEi.
  • the field representation 452 may be considered as a final representation of the target field fl.
  • the field encoder 450 may be a machine learning model based on a transformer structure.
  • the process through which the field encoder 450 generates the field representation 452 Q L may be represented by the following equation:
  • conditional formatting corresponding to the target field fl may be recommended through performing a plurality of tasks.
  • the plurality of tasks may include analytical semantics tasks, e.g., a user intent classification task and a data focus classification task. Through the user intent classification task, user intent corresponding to the target field fl may be inferred. Through the data focus classification task, data focus corresponding to the target field fl may be inferred.
  • the plurality of tasks may also include an operation classification task and/or a parameter generation task. Through the operation classification task, an operation corresponding to the target field fl may be recommended. Through the parameter generation task, a parameter set corresponding to the target field fl may be recommended.
  • the operation and/or the parameter set may be combined into conditional formatting corresponding to the target field fl . Since the process 400 aims to recommend conditional formatting corresponding to the target field fl , the operation classification task and/or the parameter generation task may be considered as target tasks, while the user intent classification task and the data focus classification task may be considered as additional tasks for assisting the target tasks.
  • the user intent corresponding to the target field fl may be inferred based on the field representation 452 Q L .
  • the first element Q ⁇ extracted from the field representation 452 Qi may be provided to a user intent classification layer 460.
  • the user intent classification layer 460 may infer a user intent 470 corresponding to the target field f based on Q .
  • the user intent 470 may include, e.g., data comparison, data detection, etc.
  • the user intent classification layer 460 may predict probabilities y i it: for different types of user intents based on
  • a prediction loss corresponding to the user intent classification layer 460 may be calculated through a binary cross entropy loss function, as shown in the following equation:
  • T it -Wi,i t [(Pi t yi,it ⁇ log (o-(yijt)) + (i - yi,u) ⁇ log (1 - ff(y £ ,i t )))] CO where w i it is a trainable model weight.
  • a data focus classification layer 462 may infer a data focus 472 c cn corresponding to the target field based on QI .
  • the data focus may include, e.g., a false value, a null value, a meaningless value, an empirical value, etc.
  • the data focus classification layer 462 may predict probabilities y i d for different types of data focuses based on Q®.
  • a prediction loss corresponding to the data focus classification layer 462 may be calculated through a binary cross entropy loss function, as shown in the following equation:
  • T d -w i ,d[(Pdy i ,d ⁇ log O(y £ , d )) + (1 - yt'd) ⁇ log (1 - ff(y £ ,d)))] ( 8 ) where w i d is a trainable model weight.
  • conditional formatting corresponding to the target field f? may be automatically recommended based on the field representation 452 and at least one of the inferred user intent 470 and the data focus 472.
  • the conditional formatting may be automatically recommended through an operation classification layer 464 and a parameter generation layer 466.
  • the user intent 470 and the data focus 472 may be provided, as expert knowledge 474 that provides guidance information, to the operation classification layer 464.
  • the operation classification layer 464 may automatically recommend an operation 476 corresponding to the target field fT based on the expert knowledge 474 and the field representation 452 Q t .
  • the expert knowledge 474 and the first element Q® extracted from the field representation 452 may be provided to the operation classification layer 464.
  • the operation classification layer 464 may be a task layer for performing multi-label classification task.
  • the operation classification layer 464 may predict probabilities y i op for different types of operations based on the expert knowledge 474 and Q® .
  • a prediction loss corresponding to the operation classification layer 464 may be calculated through a binary cross entropy loss function, as shown in the following equation: where w i op is a trainable model weight.
  • the operation 476 may include e.g., adding an icon, adding a data graph, adjusting a font style, etc.
  • parameter sets corresponding to the operations may be further recommended.
  • a parameter set 478 corresponding to the target field ft may be recommended based on the field representation 452 Qi and the operation 476 through the parameter generation layer 466.
  • the parameter generation layer 466 may be a task layer for performing a multi-label classification task.
  • the last hidden layer embedding in the field representation 452 Q t may be considered as a cell representation embedding.
  • These embeddings and the operation 476 may be provided to the parameter generation layer 466.
  • the parameter generation layer 466 may accordingly predict probabilities yi :Param that each cell may be recommended for parameters of the corresponding operation.
  • a prediction loss corresponding to the parameter generation layer 466 may be calculated through a binary cross entropy loss function, as shown in the following equation: log (1 - where w i param is a trainable model weight.
  • the operation 476 and/or the parameter set 478 may be combined into conditional formatting 480 corresponding to the target field ff .
  • the user intent classification layer 460, the data focus classification layer 462, the operation classification layer 464, and the parameter generation layer 466 may be trained simultaneously.
  • a final prediction loss may be calculated through, e.g., the following equation: final - aT it + f d + yT op + 8T param (11) where a, /?, yand 6 are scaling factors for respective task.
  • the process for automatically recommending conditional formatting described above in conjunction with FIG.4 is merely exemplary.
  • the steps in the process for automatically recommending conditional formatting may be replaced or modified in any manner, and the process may include more or fewer steps.
  • both the user intent 470 and the data focus 472 may be provided, as expert knowledge, to the operation classification layer 464, but in some embodiments, only one of the user intent 470 and the data focus 472 may be provided to the operation classification layer 464.
  • the specific order or hierarchy of the steps in the process 400 is merely exemplary, and the process for automatically recommending conditional formatting may be performed in an order different from the described one.
  • FIG.5 is a flowchart of an exemplary method 500 for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
  • a table may be obtained, the table containing a plurality of fields.
  • at least one field representation of at least one field in the plurality of fields may be generated.
  • conditional formatting corresponding to the field may be automatically recommended based at least on the field representation.
  • the table may be visualized through formatting the field based on the conditional formatting.
  • the table may include a local table or a table retrieved through a search engine.
  • the generating at least one field representation may comprise: generating a cell-level representation of the field; generating a field-level representation of the field; and generating the field representation based on the cell-level representation and the field-level representation.
  • the field may include a cell set.
  • the generating a cell-level representation may comprise: generating the cell-level representation based on a statistical feature and a linguistic feature corresponding to the cell set.
  • the field may include a cell set.
  • the generating a cell-level representation may comprise: sampling a cell subset from the cell set based on a statistical feature corresponding to the cell set and/or a statistical feature corresponding to the field; and generating the cell-level representation based on a statistical feature and a linguistic feature corresponding to the cell subset.
  • the generating a field-level representation may comprise: generating the field-level representation based on a statistical feature and a linguistic feature corresponding to the field.
  • the automatically recommending conditional formatting may comprise: inferring at least one of user intent and data focus corresponding to the field based on the field representation; and automatically recommending the conditional formatting based on the field representation and at least one of the user intent and the data focus.
  • the automatically recommending the conditional formatting may comprise: automatically recommending an operation corresponding to the field based on the field representation and at least one of the user intent and the data focus.
  • the operation may comprise at least one of adding an icon, adding a data graph, and adjusting a font style.
  • the data graph may include a data bar graph and/or a data pie graph.
  • the method 500 may further comprise: automatically recommending a parameter set corresponding to the field based on the field representation and the operation.
  • conditional formatting may include an operation.
  • the operation may include adding an icon.
  • the formatting the field may comprise, for each cell in the field: identifying an entity corresponding to a value in the cell; extracting an icon corresponding to the entity from a knowledge graph; and adding the icon at the cell.
  • the conditional formatting may include an operation and a parameter set.
  • the operation may include adding a data graph.
  • the parameter set may include a global value.
  • the formatting the field may comprise, for each cell in the field: calculating a proportion of a prominent portion in a data graph to the data graph based on a value in the cell and the global value; populating the data graph based on the calculated proportion; and adding the data graph at the cell.
  • conditional formatting may include an operation and a parameter set.
  • the operation may include adjusting a font style.
  • the parameter set may include at least one threshold.
  • the formatting the field may comprise, for each cell in the field: determining a partition corresponding to the cell based on the at least one threshold; and adjusting a font style of a value in the cell based on the determined partition.
  • the method 500 may further comprise any step/process for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above.
  • FIG.6 illustrates an exemplary apparatus 600 for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
  • the apparatus 600 may comprise: a table obtaining module 610, for obtaining a table, the table containing a plurality of fields; a field representation generating module 620, for generating at least one field representation of at least one field in the plurality of fields; a conditional formatting recommending module 630, for automatically recommending conditional formatting corresponding to the field based at least on the field representation; and a table visualizing module 640, for visualizing the table through formatting the field based on the conditional formatting.
  • the apparatus 600 may further comprise any other modules configured for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above.
  • FIG.7 illustrates an exemplary apparatus 700 for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
  • the apparatus 700 may comprise at least one processor 710 and a memory 720 storing computerexecutable instructions.
  • the computer-executable instructions when executed, may cause the at least one processor to: obtain a table, the table containing a plurality of fields; generate at least one field representation of at least one field in the plurality of fields; automatically recommend conditional formatting corresponding to the field based at least on the field representation; and visualize the table through formatting the field based on the conditional formatting.
  • the generating at least one field representation may comprise: generating a cell-level representation of the field; generating a field-level representation of the field; and generating the field representation based on the cell-level representation and the field-level representation.
  • the automatically recommending conditional formatting may comprise: inferring at least one of user intent and data focus corresponding to the field based on the field representation; and automatically recommending the conditional formatting based on the field representation and at least one of the user intent and the data focus.
  • the automatically recommending the conditional formatting may comprise: automatically recommending an operation corresponding to the field based on the field representation and at least one of the user intent and the data focus.
  • the computer-executable instructions when executed, may further cause the at least one processor 710 to: automatically recommending a parameter set corresponding to the field based on the field representation and the operation.
  • processor 710 may further perform any other step/process of the method for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above.
  • the embodiments of the present disclosure propose a computer program product for table visualization based on intelligent conditional formatting, comprising a computer program that is executed by at least one processor for: obtaining a table, the table containing a plurality of fields; generating at least one field representation of at least one field in the plurality of fields; automatically recommending conditional formatting corresponding to the field based at least on the field representation; and visualizing the table through formatting the field based on the conditional formatting.
  • the computer program may be further executed for implementing any other steps/processes of the method for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above.
  • the embodiments of the present disclosure may be embodied in a non-transitory computer- readable medium.
  • the non-transitory computer readable medium may comprise instructions that, when executed, cause one or more processors to perform any operation of the method for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above.
  • modules in the apparatuses described above may be implemented in various approaches. These modules may be implemented as hardware, software, or a combination thereof. Moreover, any of these modules may be further functionally divided into sub-modules or combined together.
  • processors have been described in connection with various apparatuses and methods. These processors may be implemented using electronic hardware, computer software, or any combination thereof. Whether such processors are implemented as hardware or software will depend upon the particular application and overall design constraints imposed on the system.
  • a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with a microprocessor, microcontroller, digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gated logic, discrete hardware circuits, and other suitable processing components configured for performing the various functions described throughout the present disclosure.
  • DSP digital signal processor
  • FPGA field-programmable gate array
  • PLD programmable logic device
  • the functionality of a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with software being executed by a microprocessor, microcontroller, DSP, or other suitable platform.
  • a computer-readable medium may include, by way of example, memory such as a magnetic storage device (e.g., hard disk, floppy disk, magnetic strip), an optical disk, a smart card, a flash memory device, random access memory (RAM), read only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), a register, or a removable disk.
  • memory is shown separate from the processors in the various aspects presented throughout the present disclosure, the memory may be internal to the processors, e.g., cache or register.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

The present disclosure proposes a method, apparatus and computer program products for table visualization based on intelligent conditional formatting. A table may be obtained, the table containing a plurality of fields. At least one field representation of at least one field in the plurality of fields may be generated. Conditional formatting corresponding to the field may be automatically recommended based at least on the field representation. The table may be visualized through formatting the field based on the conditional formatting.

Description

TABLE VISUALIZATION BASED ON INTELLIGENT CONDITIONAL FORMATTING
BACKGROUND
With the development of computer technology and network technology, people are increasingly acquiring the information they need through the network. Applications such as search engines may help users easily acquire information on the network. For example, a user may enter a query into a search box of a search engine. Based on the query, the search engine may retrieve documents relevant to the query from a pre-established index database, and then perform subsequent processing on these documents, e.g., relevance filtering, ranking, etc., and finally present the highest-ranked series of documents to the user through a search result page.
SUMMARY
This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. It is not intended to identify key features or essential features of the claimed subj ect matter, nor is it intended to be used to limit the scope of the claimed subj ect matter. Embodiments of the present disclosure propose a methods, apparatus and computer program product for table visualization based on intelligent conditional formatting. A table may be obtained, the table containing a plurality of fields. At least one field representation of at least one field in the plurality of fields may be generated. Conditional formatting corresponding to the field may be automatically recommended based at least on the field representation. The table may be visualized through formatting the field based on the conditional formatting.
It should be noted that the above one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the drawings set forth in detail certain illustrative features of the one or more aspects. These features are only indicative of the various ways in which the principles of various aspects may be employed, and this disclosure is intended to include all such aspects and their equivalents.
BRIEF DESCRIPTION OF THE DRAWINGS
The disclosed aspects will hereinafter be described in connection with the appended drawings that are provided to illustrate and not to limit the disclosed aspects.
FIG.l illustrates an exemplary process for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
FIG.2 illustrates an exemplary table and an exemplary visualized table corresponding to the table according to an embodiment of the present disclosure.
FIG.3 illustrates another exemplary table and an exemplary visualized table corresponding to the table according to an embodiment of the present disclosure. FIG.4 illustrates an exemplary process for automatically recommending conditional formatting according to an embodiment of the present disclosure.
FIG.5 is a flowchart of an exemplary method for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
FIG.6 illustrates an exemplary apparatus for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
FIG.7 illustrates an exemplary apparatus for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
DETAILED DESCRIPTION
The present disclosure will now be discussed with reference to several example implementations. It is to be understood that these implementations are discussed only for enabling those skilled in the art to better understand and thus implement the embodiments of the present disclosure, rather than suggesting any limitations on the scope of the present disclosure.
A search result page associated with a search engine may typically include relevant information for respective documents. A user may quickly learn about a specific document through viewing information for that document, and decide whether he/she wants to click a corresponding link for deeper knowledge. The information presented on the search result page may include a title, an abstract, a picture, a link, etc. Additionally, in the case where the document includes a table, the table may also be presented on the search result page. Currently, tables presented on search result pages are usually extracted directly from documents, which may be in plain text format. For example, a table presented on a search result page may contain only raw text or numbers without any formatting modifications.
Embodiments of the present disclosure propose table visualization based on intelligent conditional formatting. Conditional formatting is an analysis tool for tabular data, which may help users select data on a table, and format the selected data. Existing conditional formatting requires a user to provide conditions for selecting data and a format for the selected data. The intelligent conditional formatting according to the embodiments of the present disclosure may automatically recommend at least one conditional formatting corresponding to at least one table field in a table. Herein, a column in the vertical direction in a table may be referred to as a table field, or simply a field. Conditional formatting corresponding to a particular field may include, e.g., an operation appropriate for that field and/or a parameter set corresponding to that operation. The recommended conditional formatting may be automatically applied to the corresponding table field for table visualization. The table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure may selectively format the table in an intelligent manner. This process may be performed automatically, and may be widely applied to various types of tables, e.g., local tables, tables retrieved through search engines, etc., without requiring to specially set conditional formatting or determine formatting data for specific tables.
In an aspect, an embodiment of the present disclosure proposes to automatically recommend conditional formatting through a machine learning model. Herein, a machine learning model used to automatically recommend conditional formatting may be referred to as an Analytical Semantics over Tabular Data Engine. The analytical semantics over tabular data engine may generate at least one field representation of at least one field of the plurality of fields contained in the table, and automatically recommend conditional formatting corresponding to the field based at least on the generated field representation. Herein, a field in a table for which conditional formatting is recommended may be referred to as a target field. The analytical semantics over tabular data engine may infer user motivations corresponding to a target field, e.g., user intent, data focus, etc., based on a field representation of the target field. User intent may refer to the intent that motivates the user to create an analysis of the table, e.g. the type of analysis the user might want to perform. The user intent may include, e.g., data comparison, data detection, etc. The data comparison may refer to a user comparing a plurality of sets of cells through making the plurality of sets of cells have different formatting. The data detection may refer to a user selecting a portion of data from a given table field, and highlighting the selected data through a predefined formatting. Data focus may refer to important data features that should be paid attention to in data reference. The data focus may include, e.g., a false value, a null value, a meaningless value, an empirical value, etc. The user intent and the data focus may be collectively referred to as analytical semantics. After inferring the user intent and the data focus corresponding to the target field, the analytical semantics over tabular data engine may further recommend conditional formatting corresponding to the target field based on the user intent, the data focus, and the field representation of the target field. The user intent and the data focus may serve as expert knowledge to guide the analytical semantics over tabular data engine to give more accurate recommendation for conditional formatting of the target field, and may make a recommendation result interpretable and easy for humans to understand.
In another aspect, an embodiment of the present disclosure proposes to generate a field representation of a target field from a plurality of structural levels and based on different feature types. For example, a cell-level representation and a field-level representation of the target field may be generated, respectively, and a field representation of the target field may be generated based on the generated cell-level representation and field-level representation. The cell-level representation may be generated from a cell-level statistical feature and linguistic feature. The field-level representation may be generated from a field-level statistical feature and linguistic feature. The cell-level statistical feature and the field-level statistical feature may reflect the data distribution characteristics of corresponding levels. The cell-level linguistic feature and the fieldlevel linguistic feature may reflect linguistic or semantic characteristics of corresponding levels. In this way, a field representation containing rich and complete information can be obtained, so that a more accurate recommendation result can be obtained when conditional formatting recommendation is performed subsequently.
In yet another aspect, an embodiment of the present disclosure proposes to automatically recommend a plurality of types of conditional formatting. Operations included in conditional formatting may include e.g., adding an icon, adding a data graph, adjusting a font style, etc. The icon may include, e.g., an image or text composed of graphic and/or text with a sense of design, e.g., a national flag of a country, a logo of a company, a team flag of a team, etc. The data graph may include, e.g., a data bar graph, a data pie graph, etc. Adjusting the font style may include, e.g., changing a font color, bolding the font, italicizing the font, underlining the font, etc. For example, when a set of cells in a target field contains a set of entities, an operation appropriate for the field may be adding a set of icons corresponding to the set of entities; when a set of cells in a target field contains a set of numbers, an operation appropriate for the field may be adding a set of data graphs corresponding to the set of numbers; when some cells in a target field have different characteristics from other cells, an operation appropriate for the field may be adjusting a font style of these cells to highlight these cells; etc. Some operations may have corresponding parameter sets. The recommended conditional formatting may be automatically applied to the corresponding table field for table visualization. Compared to the original table, the visualized table contains richer visual information, thus embodying stronger entity or numerical properties.
FIG.l illustrates an exemplary process 100 for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
At first, a table 102 may be obtained. The table 102 may be a local table located at a terminal device, e.g., a table authored by a spreadsheet authoring tool at the terminal device. Additionally, the table 102 may be a table retrieved through a search engine. For example, a query may be entered into a search engine, and search results for the query may be obtained from the search engine. The search results may be obtained by the search engine through performing operations such as searching, ranking, etc., for the query. Some search results may contain tables. The table 102, such as a local table, a table retrieved through a search engine, etc., may be a table in plain text format.
The table 102 may be a multi-dimension data set containing a plurality of fields. FIG.2 illustrates an exemplary table 200a and an exemplary visualized table 200b corresponding to the table 200a according to an embodiment of the present disclosure. The table 200a may be a table in plain text format. The table 200a may contain 4 fields, i.e., a field 202a, a field 204a, a field 206a, and a field 208a. The field 202a may include a set of rankings. The field 204a may include a set of company names. The field 206a may include a set of sales corresponding to the set of company names in the field 204a. The field 208a may include a set of market shares corresponding to the set of company names in the field 204a. The table 200b may be an exemplary visualized table corresponding to the table 200a. The table 200b may contain 4 fields, i.e., a field 202b, a field 204b, a field 206b, and a field 208b. The field 202b, the field 204b, the field 206b, and the field 208b may correspond to the field 202a, the field 204a, the field 206a, and the field 208a in the table 200a, respectively.
The table 102 may be provided to an analytical semantics over tabular data engine 104. The analytical semantics over tabular data engine 104 may be a machine learning model that may automatically recommend at least one conditional formatting 106 corresponding to at least one field in the table 102. The analytical semantics over tabular data engine 104 may generate at least one field representation of at least one of the plurality of fields contained in the table 102, and automatically recommend conditional formatting corresponding to the field based at least on the generated field representation. For example, the analytical semantics over tabular data engine 104 may infer user motivations corresponding to a target field, e.g., user intent, data focus, etc., based on a field representation of the target field. The user intent may include, e.g., data comparison, data detection, etc. The data focus may include, e.g., a false value, a null value, a meaningless value, an empirical value, etc. The user intent and the data focus may be collectively referred to as analytical semantics. After inferring the user intent and the data focus corresponding to the target field, the analytical semantics over tabular data engine 104 may further recommend conditional formatting corresponding to the target field based on the user intent, the data focus, and the field representation of the target field. The user intent and the data focus may serve as expert knowledge to guide the analytical semantics over tabular data engine 104 to give more accurate recommendation for conditional formatting, and may make the recommendation result interpretable and easy for humans to understand. An exemplary process for automatically recommending conditional formatting will be described later in conjunction with FIG.4.
Conditional formatting for a target field may include an operation appropriate for that field and/or a parameter set. Operations recommended by the analytical semantics over tabular data engine 104 may include, e.g., adding an icon 106-1, adding a data graph 106-2, adjusting a font style 106- 3, etc. The icon may include, e.g., a national flag of a country, a logo of a company, a team flag of a team, etc. The data graph may include, e.g., a data bar graph, a data pie graph, etc. Adjusting the font style may include, e.g., changing a font color, bolding the font, italicizing the font, underlining the font, etc. Some operations may also have corresponding parameter sets (not shown in the figure). After obtaining the conditional formatting 106 automatically recommended by the analytical semantics over tabular data engine 104, the table 102 may be visualized through formatting the corresponding fields based on the recommended conditional formatting 106, thereby obtaining a visualized table 112. For example, for a target field, formatted data for the field may be determined through a formatting data enveloper 108 based on the recommended conditional formatting 106. The field may then be rendered by a formatting Tenderer 110 based on the formatting data determined by the formatting data enveloper 108, thereby rendering a field with the corresponding format. For example, when the table 102 is a table retrieved through a search engine, the visualized table 112 may be presented on a search result page by the formatting Tenderer 110 using a programming language such as Hypertext Markup Language (HTML).
The formatting data enveloper 108 may include a plurality of units corresponding to different conditional formatting, e.g., an icon linking unit 108-1, a data graph populating unit 108-2, a font style determining unit 108-3, etc.
When an operation included in the conditional formatting corresponding to the field in the table 102 is adding an icon, the icon linking unit 108-1 in the formatting data enveloper 108 may determine an icon corresponding to each cell in the field. For each cell in the field, an entity corresponding to a value in the cell may be identified. Herein, a value in a cell may refer to content in the cell, which may include, e.g., number, text, etc. The entity corresponding to the value in the cell may be identified by known entity recognition techniques. Subsequently, an icon corresponding to the identified entity may be extracted from a knowledge graph through known entity linking techniques. The knowledge graph may be, e.g., a knowledge graph corresponding to the category of the identified entity. The extracted icon may be added at the cell.
Referring to FIG.2, the field 204a may contain a set of company names. The operation included in the conditional formatting corresponding to the field 204a recommended by the analytical semantics over tabular data engine 104 may be adding an icon. Accordingly, the formatting data enveloper 108 may determine a set of logos corresponding to a set of cells in the field 204a. The determined set of logos may be added to the corresponding cells, thereby obtaining the field 204b. Compared to the field 204a, the field 204b contains richer visual information, e.g., icons corresponding to the entities in the field 204b, thereby embodying stronger entity properties.
When the operation included in the conditional formatting corresponding to the field in the table 102 is adding a data graph, the data graph populating unit 108-2 in the formatting data enveloper 108 may, based on a predetermined rule, populate a set of data graphs corresponding to a set of cells in the field. When the operation is adding the data graph, the conditional formatting may also include a parameter set corresponding to the operation, e.g., a global value. A global value may be, e.g., a statistical value of values in a set of cells in a field, e.g., a maximum value, a minimum value, an average value, etc. Additionally, the global value may be, e.g., a predetermined value related to the values in the set of cells in the field. As an example, when values in a set of cells in a field are the population numbers of some countries, the global value may be the total global population. For each cell in the field, a proportion of a prominent portion in a data graph to the data graph may be calculated based on a value in the cell and the global value. The prominent portion in the data graph may be indicated through, e.g., shading, different colors, etc. For example, for a data bar graph, the length of a shaded bar in the data bar graph may be calculated based on a value in the cell and the global value. For a data pie graph, the angle of a sector in the data pie graph may be calculated based on a value in the cell and the global value. Subsequently, the data graph may be populated based on the calculated proportion. The populated data graphs may be added to the corresponding cells.
Referring to the table 200a in FIG.2, the field 206a may contain a set of sales corresponding to a set of countries in the field 204. The operation included in the conditional formatting corresponding to the field 206a recommended by the analytical semantics over tabular data engine 104 may be adding a data bar graph, and the corresponding parameter may be the maximum value in the cells in field 206a, i.e., " 17,098,242" in the cell 210a. Thus, for each cell in the field 206a, a proportion of a prominent portion in a data bar graph to the data bar graph may be calculated based on a value in the cell and the maximum value "17,098,242". The prominent portion in the data bar graph may be indicated through shading. The populated data bar graphs may be added to the corresponding cells in the field 206a, thereby obtaining the field 206b. Compared with the field 206a, the field 206b contains richer visual information, e.g., the data bar graphs that may visually present relative magnitude of the values in respective cells, thereby embodying stronger numerical properties.
Additionally, referring to the field 208a in the table 200a, the field 208a may include a set of market shares corresponding to the set of company names in the field 204a. The operation included in the conditional formatting corresponding to the field 208a recommended by the analytical semantics over tabular data engine 104 may be adding a data pie graph, and the corresponding parameter may be the total sales of a similar product in the market. Therefore, for each cell in the field 208a, a proportion of a prominent portion in a data pie graph to the data pie graph may be calculated based on a value in the cell and the total sales of a similar product in the market. The prominent portion in the data pie graph may be indicated through a darker color. The populated data pie graphs may be added to the corresponding cells in the field 208a, thereby obtaining the field 208b. Compared with the field 208a, the field 208b contains richer visual information, e.g., the data pie graphs those may visually present relative proportions of the values in respective cells to total sales of a similar product in the market, thereby embodying stronger numerical properties. When the operation included in the conditional formatting corresponding to the field in the table 102 is adjusting a font style, the font style determining unit 108-3 in the formatting data enveloper 108 may determine a font style of values of cells in the field based on a predetermined rule. When the operation is adjusting a font style, the conditional formatting may further include a parameter set corresponding to the operation, e.g., at least one threshold for determining a partition in which each cell in the field is located. For each cell of the field, a partition corresponding to the cell may be determined based on a threshold. Subsequently, a font style of a value in the cell may be adjusted based on the determined partition. In particular, a partition may be highlighted through adjusting its font style to make it significantly different from other partitions.
FIG.3 illustrates another exemplary table 300a and an exemplary visualized table 300b corresponding to the table 300a according to an embodiment of the present disclosure. The table 300a may contain 4 fields, i.e., a field 302a, a field 304a, a field 306a, and a field 308a. The field 302a may include a set of majors. The field 304a may include a set of enrolments in 2020 corresponding to the set of majors in the field 302a. The field 306a may include a set of enrolments in 2021 corresponding to the set of majors in the field 302a. The field 308a may include a set of enrolment increases in 2021 corresponding to the set of majors in the field 302a. The operation included in the conditional formatting corresponding to the field 308a recommended by the analytical semantics over tabular data engine 104 may be adjusting a font style, and the corresponding parameter may be a threshold "0". The table 300a may be visualized through formatting the field 308a based on the conditional formatting, thereby obtaining a visualized table 300b. The table 300b may contain 4 fields, i.e., a field 302b, a field 304b, a field 306b, and a field 308b. The field 302b, the field 304b, the field 306b, and the field 308b may correspond to the field 302a, the field 304a, the field 306a, and the field 308a in the table 300a, respectively. For each cell in the field 308a, a partition corresponding to the cell may be determined based on the threshold value "0". For example, when a value in a cell is greater than or equal to the threshold "0", i.e., when the increase is positive or zero, the cell may be determined to be in the first partition; and when the value in the cell is less than the threshold "0", i.e., when the increase is negative, the cell may be determined to be in the second partition. Different font styles may be set for cells in the first partition and cells in the second partition in the field 308a, respectively, thereby obtaining the field 308b. For example, for cells in the first partition, e.g., a cell 310a, a cell 316a, and a cell 318a, their font styles may remain unchanged, thereby obtaining a cell 310b, a cell 316b, and a cell 318b; and for cells in the second partition, e.g., a cell 312a and a cell 314a, their font styles may be adjusted to bold, thereby obtaining a cell 312b and a cell 314b. Furthermore, in order to further enrich visual information of the field 308b, a data bar graph is also shown in the field 308b. The length, position, and form of a data bar graph in each cell may correspond to a value in the cell. Compared with the field 308a, the field 308b contains richer visual information, e.g., displaying numerical values with a downward trend in bold font, and more intuitively presenting the relative magnitudes and symbols of the values in respective cells with the data bar graphs, thereby embodying stronger numerical properties.
Through the process 100, selective formatting may be performed on the table 102 in an intelligent way. The process 100 may be performed automatically. Furthermore, the process 100 may be widely applied to various types of tables, e.g., local tables, tables retrieved through search engines, etc., without requiring to specially set conditional formatting or determine formatting data for specific tables.
It should be appreciated that the process for table visualization based on intelligent conditional formatting described above in conjunction with FIGs.l to 3 is merely exemplary. Depending on actual application requirements, the steps in the process for table visualization based on intelligent conditional formatting may be replaced or modified in any manner, and the process may include more or fewer steps. For example, although only three types of operations including adding an icon, adding a data graph, and adjusting a font style are shown in the conditional formatting 106, and the formatting data enveloper 108 also only includes three types of formatting data determining units corresponding to these three types of operations, but the embodiments of the present disclosure are not limited thereto. The analytical semantics over tabular data engine 104 may also recommend other types of operations, and the formatting data enveloper 108 may accordingly include other types of formatting data determining units. Furthermore, it should be appreciated that the formatting shown in FIGs.2 and 3 are merely exemplary. Depending on actual application requirements, the table may be formatted in any other form.
FIG.4 illustrates an exemplary process 400 for automatically recommending conditional formatting according to an embodiment of the present disclosure. The process 400 may correspond to the operation at the analytical semantics over tabular data engine 104 in FIG. l. At least one conditional formatting corresponding to at least one field in a table 402 may be automatically recommended through the process 400. The process 400 is described below by taking any field in the table 402 as an example. A field for which the process 400 is directed may be referred to as a target field.
In the process 400, a field representation of the target field may first be generated. Subsequently, conditional formatting corresponding to the target field may be automatically recommended based at least on the field representation of the target field. The field representation of the target field may be generated from a plurality of structural levels and based on different feature types. In an implementation, a cell-level representation and a field-level representation of the target field may be generated, respectively, and the field representation of the target field may be generated based on the generated cell-level representation and field-level representation. The cell-level representation may be generated from a cell-level statistical feature and linguistic feature. The field-level representation may be generated from a field-level statistical feature and linguistic feature. The cell-level statistical feature and the field-level statistical feature may reflect the data distribution characteristics of corresponding levels. The cell-level linguistic feature and the fieldlevel linguistic feature may reflect linguistic or semantic characteristics of corresponding levels. In this way, a field representation containing rich and complete information can be obtained, so that a more accurate recommendation result can be obtained when conditional formatting recommendation is performed subsequently.
The table 402 may be a multi-dimension data set containing a plurality of fields. The table 402 may be denoted as
Figure imgf000012_0001
fl) , where n > 2 is used to denote the number of fields contained in the table 402. The target field in the table 402 may be denoted as /jT(l < i < ri). A cell set 404 Cellsi = (ceZZ1( ceZZk) of the target field fl may be extracted from the table 402, where k is used to denote the number of cells in the cell set 404 of the target field fl . Additionally, a table/field-level input 406 Tabt may be extracted from the table 402, where Cellsi E Tabt. The table/field-level input 406 may contain information of all fields in the table 402, and may contain two-dimensional structure information of the table 402.
A cell-level representation of the target field fl may be generated based on a statistical feature and a linguistic feature corresponding to the cell set 404. The statistical feature corresponding to the cell set 404 may reflect data statistical information corresponding to each cell in the cell set 404. Data statistical information corresponding to a specific cell may include, e.g., the ranking of a value in the cell in the field in which it is located, whether the cell contains a null value, etc. The linguistic feature corresponding to the cell set 404 may reflect linguistic or semantic characteristics corresponding to each cell in the cell set 404. Preferably, when generating the cell-level representation of the target field fl, a cell subset 418 Cellsi' of the target field fl may be first sampled from the cell set 404 of the target field fl. For example, when the number of cells included in the cell set 404 exceeds a predetermined threshold, the cell subset 418 Cells^ may be sampled from the cell set 404 first. The cell-level representation of the target field fl may then be generated based on a statistical feature and a linguistic feature corresponding to the cell subset 418 Cells-.
In an implementation, the cell subset 418 Cellsi' may be sampled from the cell set 404 based on a statistical feature corresponding to the cell set 404 and/or a statistical feature corresponding to the target field fl . The statistical feature corresponding to the cell set 404 may be characterized as a cell signature 410 of the target field fl . The cell signature 410 of the target field fl may be calculated based on the cell set 404 through a cell signature calculating unit 408. The statistical feature corresponding to the target field ft may reflect data distribution characteristics of the target field ff . The statistical feature corresponding to the target field ff may be characterized as a field signature 414 of the target field ff . The field signature 414 of the target field ff may be calculated based on the table/field-level input 406 through a field signature calculating unit 412. The field signature 414 may be denoted as F .
A cell sampling unit 416 may sample the cell subset 418 Cells' from the cell set 404 based on the cell signature 410 and/or the field signature 414. Through sampling the cell set 404, the number of cells provided to a subsequent model can be reduced, and the size of a candidate pool used to generate parameters can be reduced, which help to improve the efficiency and performance of the model. After the cell subset 418 Cells' is sampled, the cell signature 410 may be updated to obtain an updated cell signature 420. The updated cell signature 420 may correspond to the cell subset 418 Cells^. The updated cell signature 420 may be denoted as
Figure imgf000013_0001
. It should be appreciated that when the number of cells included in the cell set 404 does not exceed a predetermined threshold, the sampling operation may not be performed. In this case, the cell subset 418 may be consistent with the cell set 404, and the updated cell signature 420 may be consistent with the cell signature 410.
The cell subset 418 Cells may be provided to a Pre-trained Language Model (PLM) 422. The pre-trained language model 422 may generate a linguistic feature corresponding to the cell subset 418 Cells . The pre-trained language model 422 may be, e.g., a Bidirectional Encoder Representations from Transformers (BERT) model, a Generative Pre-trained Transformer (GPT) model, etc. For example, the pre-trained language model 422 may generate a cell-level linguistic representation 424 CLR of the target field ft based on the cell subset 418 Cells'. The pre-trained language model 422 may regard content of each cell in the cell subset 418 Cellst' as a sentence sample. For each cell, an embedding for the starting token of the cell output by the pre-trained language model 422 may be regarded as a linguistic representation of the cell. The process described above may be as shown by the following equation:
CLRt = PLM(Cellsf) (1) where CLRt E JRfcxe, b is the number of cells in the cell subset 418 Cells^', and e is the embedding size of the pre-trained language model 422.
After the updated cell signature 420 fP and the cell-level linguistic representation 424 CLRt of the target field ff are obtained, a cell-level representation 428 MCRt of the target field ff may be generated through a combining unit 426. The combining unit 426 may first concatenate the updated cell signature 420 fP and the cell-level linguistic representation 424 CLRt . The concatenated representation may then be transformed with a linear layer and an activation function. The process described above may be as shown by the following equation:
Figure imgf000014_0001
where MCRt G IRzxD, and D is the embedding size of a field encoder 450 used subsequently. Moreover, a field-level representation of the target field ff may be generated based on a statistical feature and a linguistic feature corresponding to the target field f[. The table/field level input 406 may be provided to a Pre-trained Tabular Model (PTM) 430. The pre-trained tabular model 430 may generate a linguistic feature corresponding to the target field f[ . The pre-trained tabular model 430 may be, e.g., a Tabular Information Embedding (TABBIE) model, a Table Parser (TAPAS) model, etc. For example, the pre-trained tabular model 430 may generate a field-level linguistic representation 432 FLRt of the target field ff based on the table/field-level input 406 containing two-dimensional structural information of the table 102. As an example, for a TABBIE model, an embedding of a predetermined token, e.g. [CLS], at the start of the representation for the target field ft output by the model may be regarded as the field-level linguistic representation 432 FLRt of the target field ff . The process through which the pre-trained tabular model 430 generates the field-level linguistic representation 432 FLRt may be represented by the following equation:
FLRt = PTM(Tabt) (3) where FLRt E IRlx£, and E is the embedding size of the pre-trained tabular model 430. The fieldlevel linguistic representation 432 FLRt may contain overall tabular context information of the table 102 and information of the target field ff itself.
After the field signature 414 Ff and the field-level linguistic representation 432 FLRt °f the target field f[ are obtained, a field-level representation 436 MFRt of the target field ff may be generated through a combining unit 434. The combining unit 434 may first concatenate the field signature 414 F and the field-level representation 432 FLRt. The concatenated representation may then be transformed with a linear layer and an activation function. The process described above may be as shown by the following equation:
Figure imgf000014_0002
Through the processing of the combining unit 426 and the combining unit 434, the cell-level representation 428 MCRt and the field-level representation 436 MFRt may be located in a same feature space. A merged representation 442 MSf of the target field ff may be generated based on the cell-level representation 428 MCRt and the field-level representation 436 MFRt of the target field ff through a combining unit 440. The combining unit 440 may append the cell-level representation 428 MCRt to the field-level representation 436 MFRt , to obtain the merged representation 442 MSf. The process described above may be as shown by the following equation:
MSf = MFRi MCRi (5) where MSf E IRzxD.
A token type representation 444 TEt corresponding to the merged representation 442 MSf may be obtained. The token type representation 444 TEt may indicate the type to which each embedding in the merged representation 442 MSf corresponds, including, e.g., a field-level representation, a cell-level representation, etc. A field encoder 450 may generate a field representation 452
Figure imgf000015_0001
of the target field ff based on the merged representation 442 MSf and the token type representation 444 TEi. The field representation 452
Figure imgf000015_0002
may be considered as a final representation of the target field fl. The field encoder 450 may be a machine learning model based on a transformer structure. The process through which the field encoder 450 generates the field representation 452 QL may be represented by the following equation:
Qi = Trans MSI , TE ) (6)
After the field representation 452
Figure imgf000015_0003
of the target field fl is obtained, conditional formatting corresponding to the target field fl may be recommended through performing a plurality of tasks. The plurality of tasks may include analytical semantics tasks, e.g., a user intent classification task and a data focus classification task. Through the user intent classification task, user intent corresponding to the target field fl may be inferred. Through the data focus classification task, data focus corresponding to the target field fl may be inferred. The plurality of tasks may also include an operation classification task and/or a parameter generation task. Through the operation classification task, an operation corresponding to the target field fl may be recommended. Through the parameter generation task, a parameter set corresponding to the target field fl may be recommended. The operation and/or the parameter set may be combined into conditional formatting corresponding to the target field fl . Since the process 400 aims to recommend conditional formatting corresponding to the target field fl , the operation classification task and/or the parameter generation task may be considered as target tasks, while the user intent classification task and the data focus classification task may be considered as additional tasks for assisting the target tasks.
The user intent corresponding to the target field fl may be inferred based on the field representation 452 QL. For example, the first element Q^ extracted from the field representation 452 Qi may be provided to a user intent classification layer 460. The user intent classification layer 460 may infer a user intent 470 corresponding to the target field f based on Q . The user intent 470 may include, e.g., data comparison, data detection, etc. For example, the user intent classification layer 460 may predict probabilities yi it: for different types of user intents based on
In an implementation, a prediction loss corresponding to the user intent classification layer 460 may be calculated through a binary cross entropy loss function, as shown in the following equation:
Tit = -Wi,it[(Pityi,it ■ log (o-(yijt)) + (i - yi,u) ■ log (1 - ff(y£,it)))] CO where wi it is a trainable model weight.
Alternatively or additionally, a data focus classification layer 462 may infer a data focus 472 c cn corresponding to the target field
Figure imgf000016_0001
based on QI . The data focus may include, e.g., a false value, a null value, a meaningless value, an empirical value, etc. For example, the data focus classification layer 462 may predict probabilities yi d for different types of data focuses based on Q®. In an implementation, a prediction loss corresponding to the data focus classification layer 462 may be calculated through a binary cross entropy loss function, as shown in the following equation:
Td = -wi,d[(Pdyi,d ■ log O(y£,d)) + (1 - yt'd) ■ log (1 - ff(y£,d)))] (8) where wi d is a trainable model weight.
After the user intent 470 and the data focus 472 are inferred, conditional formatting corresponding to the target field f? may be automatically recommended based on the field representation 452 and at least one of the inferred user intent 470 and the data focus 472. The conditional formatting may be automatically recommended through an operation classification layer 464 and a parameter generation layer 466.
The user intent 470 and the data focus 472 may be provided, as expert knowledge 474 that provides guidance information, to the operation classification layer 464. The operation classification layer 464 may automatically recommend an operation 476 corresponding to the target field fT based on the expert knowledge 474 and the field representation 452 Qt. The expert knowledge 474 and the first element Q® extracted from the field representation 452
Figure imgf000016_0002
may be provided to the operation classification layer 464. The operation classification layer 464 may be a task layer for performing multi-label classification task. The operation classification layer 464 may predict probabilities yi op for different types of operations based on the expert knowledge 474 and Q® . In an implementation, a prediction loss corresponding to the operation classification layer 464 may be calculated through a binary cross entropy loss function, as shown in the following equation:
Figure imgf000016_0003
where wi op is a trainable model weight. The operation 476 may include e.g., adding an icon, adding a data graph, adjusting a font style, etc. For some operations, parameter sets corresponding to the operations may be further recommended. For example, when the operation 476 is adding a data graph, adjusting a font style, etc., a parameter set corresponding to the operation may be further recommended. A parameter set 478 corresponding to the target field ft may be recommended based on the field representation 452 Qi and the operation 476 through the parameter generation layer 466. The parameter generation layer 466 may be a task layer for performing a multi-label classification task. The last hidden layer embedding in the field representation 452 Qt may be considered as a cell representation embedding. These embeddings and the operation 476 may be provided to the parameter generation layer 466. The parameter generation layer 466 may accordingly predict probabilities yi:Param that each cell may be recommended for parameters of the corresponding operation. In an implementation, a prediction loss corresponding to the parameter generation layer 466 may be calculated through a binary cross entropy loss function, as shown in the following equation: log (1 -
Figure imgf000017_0001
where wi param is a trainable model weight. The operation 476 and/or the parameter set 478 may be combined into conditional formatting 480 corresponding to the target field ff .
The user intent classification layer 460, the data focus classification layer 462, the operation classification layer 464, and the parameter generation layer 466 may be trained simultaneously. A final prediction loss may be calculated through, e.g., the following equation: final - aTit + f d + yTop + 8Tparam (11) where a, /?, yand 6 are scaling factors for respective task.
It should be appreciated that the process for automatically recommending conditional formatting described above in conjunction with FIG.4 is merely exemplary. Depending on actual application requirements, the steps in the process for automatically recommending conditional formatting may be replaced or modified in any manner, and the process may include more or fewer steps. For example, in the process 400, both the user intent 470 and the data focus 472 may be provided, as expert knowledge, to the operation classification layer 464, but in some embodiments, only one of the user intent 470 and the data focus 472 may be provided to the operation classification layer 464. In addition, the specific order or hierarchy of the steps in the process 400 is merely exemplary, and the process for automatically recommending conditional formatting may be performed in an order different from the described one.
FIG.5 is a flowchart of an exemplary method 500 for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
At 510, a table may be obtained, the table containing a plurality of fields. At 520, at least one field representation of at least one field in the plurality of fields may be generated.
At 530, conditional formatting corresponding to the field may be automatically recommended based at least on the field representation.
At 540, the table may be visualized through formatting the field based on the conditional formatting.
In an implementation, the table may include a local table or a table retrieved through a search engine.
In an implementation, the generating at least one field representation may comprise: generating a cell-level representation of the field; generating a field-level representation of the field; and generating the field representation based on the cell-level representation and the field-level representation.
The field may include a cell set. The generating a cell-level representation may comprise: generating the cell-level representation based on a statistical feature and a linguistic feature corresponding to the cell set.
The field may include a cell set. The generating a cell-level representation may comprise: sampling a cell subset from the cell set based on a statistical feature corresponding to the cell set and/or a statistical feature corresponding to the field; and generating the cell-level representation based on a statistical feature and a linguistic feature corresponding to the cell subset.
The generating a field-level representation may comprise: generating the field-level representation based on a statistical feature and a linguistic feature corresponding to the field.
In an implementation, the automatically recommending conditional formatting may comprise: inferring at least one of user intent and data focus corresponding to the field based on the field representation; and automatically recommending the conditional formatting based on the field representation and at least one of the user intent and the data focus.
The automatically recommending the conditional formatting may comprise: automatically recommending an operation corresponding to the field based on the field representation and at least one of the user intent and the data focus.
The operation may comprise at least one of adding an icon, adding a data graph, and adjusting a font style.
The data graph may include a data bar graph and/or a data pie graph.
The method 500 may further comprise: automatically recommending a parameter set corresponding to the field based on the field representation and the operation.
In an implementation, the conditional formatting may include an operation. The operation may include adding an icon. The formatting the field may comprise, for each cell in the field: identifying an entity corresponding to a value in the cell; extracting an icon corresponding to the entity from a knowledge graph; and adding the icon at the cell.
In an implementation, the conditional formatting may include an operation and a parameter set. The operation may include adding a data graph. The parameter set may include a global value. The formatting the field may comprise, for each cell in the field: calculating a proportion of a prominent portion in a data graph to the data graph based on a value in the cell and the global value; populating the data graph based on the calculated proportion; and adding the data graph at the cell.
In an implementation, the conditional formatting may include an operation and a parameter set. The operation may include adjusting a font style. The parameter set may include at least one threshold. The formatting the field may comprise, for each cell in the field: determining a partition corresponding to the cell based on the at least one threshold; and adjusting a font style of a value in the cell based on the determined partition.
It should be appreciated that the method 500 may further comprise any step/process for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above.
FIG.6 illustrates an exemplary apparatus 600 for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
The apparatus 600 may comprise: a table obtaining module 610, for obtaining a table, the table containing a plurality of fields; a field representation generating module 620, for generating at least one field representation of at least one field in the plurality of fields; a conditional formatting recommending module 630, for automatically recommending conditional formatting corresponding to the field based at least on the field representation; and a table visualizing module 640, for visualizing the table through formatting the field based on the conditional formatting. Moreover, the apparatus 600 may further comprise any other modules configured for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above.
FIG.7 illustrates an exemplary apparatus 700 for table visualization based on intelligent conditional formatting according to an embodiment of the present disclosure.
The apparatus 700 may comprise at least one processor 710 and a memory 720 storing computerexecutable instructions. The computer-executable instructions, when executed, may cause the at least one processor to: obtain a table, the table containing a plurality of fields; generate at least one field representation of at least one field in the plurality of fields; automatically recommend conditional formatting corresponding to the field based at least on the field representation; and visualize the table through formatting the field based on the conditional formatting. In an implementation, the generating at least one field representation may comprise: generating a cell-level representation of the field; generating a field-level representation of the field; and generating the field representation based on the cell-level representation and the field-level representation.
In an implementation, the automatically recommending conditional formatting may comprise: inferring at least one of user intent and data focus corresponding to the field based on the field representation; and automatically recommending the conditional formatting based on the field representation and at least one of the user intent and the data focus.
The automatically recommending the conditional formatting may comprise: automatically recommending an operation corresponding to the field based on the field representation and at least one of the user intent and the data focus.
The computer-executable instructions, when executed, may further cause the at least one processor 710 to: automatically recommending a parameter set corresponding to the field based on the field representation and the operation.
It should be appreciated that the processor 710 may further perform any other step/process of the method for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above.
The embodiments of the present disclosure propose a computer program product for table visualization based on intelligent conditional formatting, comprising a computer program that is executed by at least one processor for: obtaining a table, the table containing a plurality of fields; generating at least one field representation of at least one field in the plurality of fields; automatically recommending conditional formatting corresponding to the field based at least on the field representation; and visualizing the table through formatting the field based on the conditional formatting. Furthermore, the computer program may be further executed for implementing any other steps/processes of the method for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above. The embodiments of the present disclosure may be embodied in a non-transitory computer- readable medium. The non-transitory computer readable medium may comprise instructions that, when executed, cause one or more processors to perform any operation of the method for table visualization based on intelligent conditional formatting according to the embodiments of the present disclosure as mentioned above.
It should be appreciated that all the operations in the methods described above are merely exemplary, and the present disclosure is not limited to any operations in the methods or sequence orders of these operations, and should cover all other equivalents under the same or similar concepts. In addition, the articles “a” and “an” as used in this specification and the appended claims should generally be construed to mean “one” or “one or more” unless specified otherwise or clear from the context to be directed to a singular form.
It should also be appreciated that all the modules in the apparatuses described above may be implemented in various approaches. These modules may be implemented as hardware, software, or a combination thereof. Moreover, any of these modules may be further functionally divided into sub-modules or combined together.
Processors have been described in connection with various apparatuses and methods. These processors may be implemented using electronic hardware, computer software, or any combination thereof. Whether such processors are implemented as hardware or software will depend upon the particular application and overall design constraints imposed on the system. By way of example, a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with a microprocessor, microcontroller, digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gated logic, discrete hardware circuits, and other suitable processing components configured for performing the various functions described throughout the present disclosure. The functionality of a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with software being executed by a microprocessor, microcontroller, DSP, or other suitable platform.
Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, threads of execution, procedures, functions, etc. The software may reside on a computer-readable medium. A computer-readable medium may include, by way of example, memory such as a magnetic storage device (e.g., hard disk, floppy disk, magnetic strip), an optical disk, a smart card, a flash memory device, random access memory (RAM), read only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), a register, or a removable disk. Although memory is shown separate from the processors in the various aspects presented throughout the present disclosure, the memory may be internal to the processors, e.g., cache or register.
The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described throughout the present disclosure that are known or later come to be known to those of ordinary skilled in the art are expressly incorporated herein and intended to be encompassed by the claims.

Claims

1. A method for table visualization based on intelligent conditional formatting, comprising: obtaining a table, the table containing a plurality of fields; generating at least one field representation of at least one field in the plurality of fields; automatically recommending conditional formatting corresponding to the field based at least on the field representation; and visualizing the table through formatting the field based on the conditional formatting.
2. The method of claim 1, wherein the generating at least one field representation comprises: generating a cell-level representation of the field; generating a field-level representation of the field; and generating the field representation based on the cell-level representation and the field-level representation.
3. The method of claim 2, wherein the field includes a cell set, and the generating a celllevel representation comprises: generating the cell-level representation based on a statistical feature and a linguistic feature corresponding to the cell set.
4. The method of claim 2, wherein the field includes a cell set, and the generating a celllevel representation comprises: sampling a cell subset from the cell set based on a statistical feature corresponding to the cell set and/or a statistical feature corresponding to the field; and generating the cell-level representation based on a statistical feature and a linguistic feature corresponding to the cell subset.
5. The method of claim 2, wherein the generating a field-level representation comprises: generating the field-level representation based on a statistical feature and a linguistic feature corresponding to the field.
6. The method of claim 1, wherein the automatically recommending conditional formatting comprises: inferring at least one of user intent and data focus corresponding to the field based on the field representation; and automatically recommending the conditional formatting based on the field representation and at least one of the user intent and the data focus.
7. The method of claim 6, wherein the automatically recommending the conditional formatting comprises: automatically recommending an operation corresponding to the field based on the field representation and at least one of the user intent and the data focus.
8. The method of claim 7, wherein the operation includes at least one of adding an icon, adding a data graph, and adjusting a font style.
9. The method of claim 8, wherein the data graph includes a data bar graph and/or a data pie graph.
10. The method of claim 7, further comprising: automatically recommending a parameter set corresponding to the field based on the field representation and the operation.
11. The method of claim 1, wherein the conditional formatting includes an operation, the operation includes adding an icon, and the formatting the field comprises, for each cell in the field: identifying an entity corresponding to a value in the cell; extracting an icon corresponding to the entity from a knowledge graph; and adding the icon at the cell.
12. The method of claim 1, wherein the conditional formatting includes an operation and a parameter set, the operation includes adding a data graph, the parameter set including a global value, and the formatting the field comprises, for each cell in the field: calculating a proportion of a prominent portion in a data graph to the data graph based on a value in the cell and the global value; populating the data graph based on the calculated proportion; and adding the data graph at the cell.
13. The method of claim 1, wherein the conditional formatting includes an operation and a parameter set, the operation includes adjusting a font style, the parameter set includes at least one threshold, and the formatting the field comprises, for each cell in the field: determining a partition corresponding to the cell based on the at least one threshold; and adjusting a font style of a value in the cell based on the determined partition.
14. An apparatus for table visualization based on intelligent conditional formatting, comprising: at least one processor; and a memory storing computer-executable instructions that, when executed, cause the at least one processor to: obtain a table, the table containing a plurality of fields; generate at least one field representation of at least one field in the plurality of fields; automatically recommend conditional formatting corresponding to the field based at least on the field representation; and visualize the table through formatting the field based on the conditional formatting.
15. A computer program product for table visualization based on intelligent conditional formatting, comprising a computer program that is executed by at least one processor for: obtaining a table, the table containing a plurality of fields; generating at least one field representation of at least one field in the plurality of fields; automatically recommending conditional formatting corresponding to the field based at least on the field representation; and visualizing the table through formatting the field based on the conditional formatting.
PCT/US2023/012655 2022-04-28 2023-02-09 Table visualization based on intelligent conditional formatting Ceased WO2023211536A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202210469254.2A CN117009498A (en) 2022-04-28 2022-04-28 Table visualization based on smart conditional formatting
CN202210469254.2 2022-04-28

Publications (1)

Publication Number Publication Date
WO2023211536A1 true WO2023211536A1 (en) 2023-11-02

Family

ID=85510836

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2023/012655 Ceased WO2023211536A1 (en) 2022-04-28 2023-02-09 Table visualization based on intelligent conditional formatting

Country Status (2)

Country Link
CN (1) CN117009498A (en)
WO (1) WO2023211536A1 (en)

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8549392B2 (en) * 2005-08-30 2013-10-01 Microsoft Corporation Customizable spreadsheet table styles
CN111428457B (en) * 2018-12-21 2024-03-22 微软技术许可有限责任公司 Automatic formatting of data tables
CN113850249A (en) * 2021-12-01 2021-12-28 深圳市迪博企业风险管理技术有限公司 Method for formatting and extracting chart information

Non-Patent Citations (6)

* Cited by examiner, † Cited by third party
Title
DONG HAOYU DONG HADONG@MICROSOFT COM ET AL: "Neural Formatting for Spreadsheet Tables", PROCEEDINGS OF THE 7TH ACM CONFERENCE ON INFORMATION-CENTRIC NETWORKING, ACMPUB27, NEW YORK, NY, USA, 19 October 2020 (2020-10-19), pages 305 - 314, XP058626844, ISBN: 978-1-4503-8312-7, DOI: 10.1145/3340531.3411943 *
DONG HAOYU HADONG@MICROSOFT COM ET AL: "Learning Formatting Style Transfer and Structure Extraction for Spreadsheet Tables with a Hybrid Neural Network Architecture", PROCEEDINGS OF THE 7TH ACM CONFERENCE ON INFORMATION-CENTRIC NETWORKING, ACMPUB27, NEW YORK, NY, USA, 19 October 2020 (2020-10-19), pages 2389 - 2396, XP058626356, ISBN: 978-1-4503-8312-7, DOI: 10.1145/3340531.3412718 *
DU LUN LUN DU@MICROSOFT COM ET AL: "TabularNet A Neural Network Architecture for Understanding Semantic Structures of Tabular Data", ACM SYMPOSIUM ON APPLIED PERCEPTION 2020, ACMPUB27, NEW YORK, NY, USA, 14 August 2021 (2021-08-14), pages 322 - 331, XP058749820, ISBN: 978-1-4503-8332-5, DOI: 10.1145/3447548.3467228 *
LINGBO LI ET AL: "ASTA: Learning Analytical Semantics over Tables for Intelligent Data Analysis and Visualization", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 1 August 2022 (2022-08-01), XP091285683 *
REMA ANANTHANARAYANAN ET AL: "DataVizard: Recommending Visual Presentations for Structured Data", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 14 November 2017 (2017-11-14), XP081288158 *
ZHOU MENGYU MEZHO@MICROSOFT COM ET AL: "Table2Charts Recommending Charts by Learning Shared Table Representations", ACM SYMPOSIUM ON APPLIED PERCEPTION 2020, ACMPUB27, NEW YORK, NY, USA, 14 August 2021 (2021-08-14), pages 2389 - 2399, XP058613838, ISBN: 978-1-4503-8332-5, DOI: 10.1145/3447548.3467279 *

Also Published As

Publication number Publication date
CN117009498A (en) 2023-11-07

Similar Documents

Publication Publication Date Title
US12282504B1 (en) Systems and methods for graph-based dynamic information retrieval and synthesis
KR101448325B1 (en) A system that facilitates delivering improved query results, a computer implemented method and a computer implemented system that facilitate providing contextual query results
JP6309644B2 (en) Method, system, and storage medium for realizing smart question answer
Fried et al. Maps of computer science
Huistra et al. Phrasing history: Selecting sources in digital repositories
KR102128659B1 (en) System and Method for Extracting Keyword and Generating Abstract
US20120102390A1 (en) Method and apparatus for generating widget
CN107704621A (en) A kind of internet public feelings map visualization methods of exhibiting
CN120541310B (en) Multi-source information retrieval and fusion methods, apparatus, devices and readable storage media
CN111694930A (en) Dynamic knowledge hotspot evolution and trend analysis method
CN120764676A (en) An intelligent question-answering system based on the Silk Road knowledge base and a large language model
JP6529698B2 (en) Data analyzer and data analysis method
CN119088935A (en) Intelligent question answering method, system, computer device and storage medium
WO2023211536A1 (en) Table visualization based on intelligent conditional formatting
Liang et al. Detecting novel business blogs
CN121256012A (en) Document processing method and device, electronic equipment and storage medium
CN119271802A (en) Inspection report generation method, equipment, medium and product for cloud network networking scenarios
JP2007279978A (en) Document search apparatus and document search method
KR20020061443A (en) Method and system for data gathering, processing and presentation using computer network
CN113176878B (en) Automatic query method, device and equipment
US20240152531A1 (en) Query-based table visualization
KR101279753B1 (en) Search service providing apparatus and method for reconstructing search result based on user&#39;s response for search result
Bartík Text-based web page classification with use of visual information
Wu et al. Empirical Analysis on User Profile in Personalized LLMs
Xiao et al. Policy narratives’ salience: A comparative analysis of artificial intelligence policy responsiveness to public attention in China and the United States

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23709848

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23709848

Country of ref document: EP

Kind code of ref document: A1