WO2018176758A1 - 用于生成文章的方法和装置 - Google Patents
用于生成文章的方法和装置 Download PDFInfo
- Publication number
- WO2018176758A1 WO2018176758A1 PCT/CN2017/102620 CN2017102620W WO2018176758A1 WO 2018176758 A1 WO2018176758 A1 WO 2018176758A1 CN 2017102620 W CN2017102620 W CN 2017102620W WO 2018176758 A1 WO2018176758 A1 WO 2018176758A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- article
- outline
- sub
- established
- rich media
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/903—Querying
- G06F16/9035—Filtering based on additional data, e.g. user or group profiles
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/31—Indexing; Data structures therefor; Storage structures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/901—Indexing; Data structures therefor; Storage structures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/103—Formatting, i.e. changing of presentation of documents
- G06F40/106—Display of layout of documents; Previewing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/12—Use of codes for handling textual entities
- G06F40/131—Fragmentation of text files, e.g. creating reusable text-blocks; Linking to fragments, e.g. using XInclude; Namespaces
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/12—Use of codes for handling textual entities
- G06F40/137—Hierarchical processing, e.g. outlines
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/166—Editing, e.g. inserting or deleting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/166—Editing, e.g. inserting or deleting
- G06F40/186—Templates
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/189—Automatic justification
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/253—Grammatical analysis; Style critique
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
Definitions
- the present application relates to the field of computer technologies, and in particular, to the field of computer network technologies, and in particular, to a method and apparatus for generating an article.
- the way to automate writing through machines to generate articles is basically to stay in special topics in special fields. Most of them use techniques that fill in rules or templates to generate articles. For example, you can filter the original article and directly reference it; or, simply transform the original article directly; or, combine the original article in a certain order and extract the abstract; or, organize and display the data through the template.
- the purpose of the present application is to propose an improved method and apparatus for generating an article to solve the technical problems mentioned in the background section above.
- an embodiment of the present application provides a method for generating an article, the method comprising: generating an article outline based on an input article theme and any one of the following: an outline model, and an outline established according to user behavior data corresponding to the topic of the article The database, and the outline of the manual setting; extracting the material associated with the feature of the article outline from the pre-established material library; inserting the extracted material into the outline of the article to obtain the generated article.
- the outline database established according to the user behavior data corresponding to the topic of the article includes: retrieving the sub-topics surrounding the topic of the article on the whole network, establishing a sub-topic database; and clicking the user according to the sub-topics in the sub-topic database and/or Or the semantic progression order of the sub-topics in the sub-topic database, sorting the sub-topics in the sub-topic database; excluding the sub-topics in the sub-topic database that do not meet the predetermined logic rules, and obtaining the sub-topics that meet the predetermined logic rules;
- the sub-themes of the logic rules are used as outlines to get the outline database.
- the pre-established material library is established by acquiring the feature of the material, and the material is obtained by filtering the content of the existing article according to the screening rule and/or transforming the content of the existing article;
- the feature builds an index structure to get the material library.
- the method further includes: optimizing the generated article to obtain an optimized generated article, and the optimization process includes one or more of the following: a touch-up process, an insert rich media data process, and a typesetting optimization process.
- the touch-up process includes one or more of the following: a grammatical style of the uniformly generated article; a statement that is inconsistent with the preceding and following statements; and a statement that is inconsistent with the preceding and following statements.
- inserting rich media data processing includes extracting rich media data associated with features of the generated article from a pre-established resource repository; and inserting the extracted rich media data into the generated article.
- extracting rich media data associated with features of the generated article from a pre-established resource library includes: extracting rich media data from the pre-established resource library to generate candidate rich media according to one or more of the following: List: article topic, article outline, summary of each paragraph of the generated article, and keywords of each paragraph of the generated article; using quality filtering to extract rich media data associated with the features of the generated article from the list of candidate rich media.
- the pre-established resource library is established by acquiring features of the rich media data, and establishing an index structure according to the characteristics of the rich media data to obtain a resource library.
- the quality screening is performed according to one or more of the following: graphic relevance, image resolution, image aspect ratio, image source authority, advertisement filtering strategy, anti Cheating filtering strategy, anti-yellow filtering strategy and watermark filtering strategy.
- the method further includes: inputting the article theme and the article outline into the title model to obtain a title of the generated article.
- the method further includes: performing attribute expansion on the core words in the title; and replacing and rewriting the core words in the expanded title to obtain the updated title.
- an embodiment of the present application provides an apparatus for generating an article, where the apparatus includes: an outline generating unit, configured to generate an article outline based on the input article theme and any one of the following: an outline model, according to a corresponding article theme An outline database for establishing user behavior data, and a manually set outline; a material extraction unit for extracting a material associated with a feature of the article outline from a pre-established material library; a material insertion unit for arranging the article , insert the extracted material and get the generated article.
- the outline database established in the outline generation unit according to the user behavior data of the corresponding article theme includes: retrieving the sub-topics surrounding the article topic of the entire network, establishing a sub-topic database; and sub-topics in the sub-topic database according to the user Sub-topics in the sub-topic database in the order of clicks and/or sub-topics in the sub-topic database; sub-topics in the sub-topic database that do not conform to the predetermined logical rules are eliminated, and sub-topics that meet predetermined logical rules are obtained The sub-themes that meet the predetermined logic rules are used as an outline to obtain an outline database.
- the pre-established material library in the material extraction unit is established by: acquiring features of the material, the material is filtering the content of the existing article according to the screening rules and/or transforming the content of the existing article. Get; build an index structure based on the characteristics of the material to get the material library.
- the apparatus further includes: an article optimization unit, configured to optimize the generated article to obtain the optimized generated article, and the optimization process includes one or more of the following: retouching processing, inserting rich media data processing And typesetting optimization processing.
- the retouching process in the article optimization unit includes one or more of the following: a grammatical style of the uniformly generated article; a statement that is inconsistent with the preceding and following statements; and a statement that replaces the inconsistent statement with the preceding and following statements.
- inserting rich media data processing in the article optimization unit includes extracting rich media data associated with features of the generated article from a pre-established resource library; The extracted rich media data is inserted into the generated article.
- extracting, from the pre-established resource library, the rich media data associated with the features of the generated article in the article optimization unit comprises: extracting rich media from the pre-established resource library according to one or more of the following: Data generation candidate rich media list: article topic, article outline, summary of each paragraph of the generated article, and keywords of each paragraph of the generated article; using quality filtering to extract from the candidate rich media list associated with the features of the generated article Rich media data.
- the pre-established resource library in the article optimization unit is established by acquiring features of the rich media data, and establishing an index structure according to the characteristics of the rich media data to obtain a resource library.
- the quality screening in the article optimization unit is performed according to one or more of the following: graphic relevance, image resolution, image aspect ratio, image source authority, advertisement filtering strategy, anti-cheat filtering strategy, Anti-yellow filtering strategy and watermark filtering strategy.
- the apparatus further includes: a title generating unit configured to input the article theme and the article outline into the title model to obtain a title of the generated article.
- the apparatus further includes: an attribute extension unit configured to perform attribute expansion on the core word in the title; and a title update unit configured to replace and rewrite the core word in the attribute extended title to obtain an update title.
- an embodiment of the present application provides an apparatus, including: one or more processors; a storage device, configured to store one or more programs; and when one or more programs are executed by one or more processors, Having one or more processors implement the method for generating an article as described above.
- the embodiment of the present application provides a computer readable storage medium, where a computer program is stored thereon, and when the program is executed by the processor, the method for generating an article as described above is implemented.
- the method and apparatus for generating an article provided by the embodiment of the present application first generates an article outline based on the input article theme and any one of the following: an outline model; an outline database established according to user behavior data corresponding to the topic of the article; and manual setting After that, the material associated with the feature of the article outline is extracted from the pre-built material library; after that, the extracted material is inserted into the article outline to obtain the generated article.
- This embodiment is implemented According to the input article theme, the outline is generated, which improves the quality of the article outline, ensures that the generated article has reasonable logic and rich form, and enriches the content of the article according to the article outline inserting the material associated with the feature of the article outline. Make the generated article logical and informative.
- FIG. 1 is a schematic flow diagram of one embodiment of a method for generating an article in accordance with the present application
- FIG. 2 is a schematic flow chart of still another embodiment of a method for generating an article according to the present application
- FIG. 3 is an exemplary application scenario of one embodiment of a method for generating an article to which the present application is applied;
- FIG. 4 is an exemplary structural diagram of one embodiment of an apparatus for generating an article in accordance with the present application.
- FIG. 5 is a schematic structural diagram of a computer system suitable for implementing a terminal device or a server of an embodiment of the present application.
- FIG. 1 illustrates a flow 100 of one embodiment of a method for generating an article in accordance with the present application.
- the methods used to generate articles include:
- an article is generated based on the input article topic and any of the following Outline: Outline model; an outline database based on user behavior data corresponding to the topic of the article; and an artificially defined outline.
- the input article theme may be a machine topic or a manually input article theme.
- An outline model usually refers to a function with an article subject as an independent variable.
- you can set the article model f (topic, outline, material), that is, the article model is obtained from the independent variables (themes, outlines, and materials) in the function f, and by using the article model, one can be used for The method of generating the article, that is, selecting the theme, mining the outline and sorting through the outline model, and mounting the material through the material library; finally, the article is obtained through the drawing, typesetting, and retouching.
- the outline database established according to the user behavior data corresponding to the topic of the article refers to determining the article directory from the perspective of the article theme, and sorting and screening the article directory according to the user behavior data to obtain an outline database. It should be understood that the outline generated by the outline generation strategy herein has a certain logical order to ensure the rationality of the text.
- the foregoing outline database established according to the user behavior data corresponding to the topic of the article includes: retrieving a sub-topic of the entire network around the topic of the article, establishing a sub-topic database; and sub-topic according to the user Sub-topics in the sub-topics in the database and/or sub-topics in the sub-topic database, sub-topics in the sub-topic database; sub-topics in the sub-topic database that do not meet the predetermined logic rules are obtained Sub-themes of logical rules; sub-topics that meet predetermined logical rules are used as outlines to obtain an outline database.
- the user's behavior data is fully considered to establish an outline, and the aim of the established outline can be improved, thereby enhancing the generated article and the user. Interaction ability.
- step 120 the material associated with the feature of the article outline is extracted from the pre-built material library.
- the pre-established material library refers to a material library obtained by establishing an index structure according to the characteristics of the material.
- the material can be extracted for later use.
- a predetermined number of materials whose features are most relevant to the features of the article outline may be extracted from the plurality of materials for later use.
- the pre-established material library is established by acquiring the feature of the material, and the material is filtering the content of the existing article according to the screening rule and/or transforming the existing article.
- the content is obtained; the index structure is built according to the characteristics of the material, and the material library is obtained.
- the generation of the material library includes material with a clear theme and material without a clear theme, and the latter needs to extract the theme using the article abstraction technique.
- Obtaining the characteristics of the material can be understood as extracting features from the text material. These features can describe the theme, keywords, core semantics and other information of the text material, and are used for correlation calculation and sorting with the article outline and the article theme.
- the filtering according to the screening rule may include filtering according to one or more of the following content: the content length of the article, the content quality score of the article, the content satisfaction rating of the article, the page view of the article, and the timeliness of the article. and many more.
- the above-mentioned transformation of the existing article content is mainly to control the granularity of the material, and can be completed by using predetermined rules. For example, a paragraph with a number of words greater than a predetermined value is disassembled and segmented.
- a material is a raw corpus. After filtering, it can be sorted according to the outline. If a material is a paragraph, you need to consider the topic relevance of the paragraph, sorting between paragraphs, etc. Similarly, you can assume that the material is a sentence. A word, the smaller the particle size of the material, the more difficult it is to disassemble and/or transform.
- step 130 the extracted material is inserted into the article outline to obtain the generated article.
- the material extracted in step 120 can be inserted into the article outline obtained in step 110, thereby obtaining the generated article.
- the method for generating an article provided by the above embodiment of the present application generates an article outline, extracts a material associated with the feature of the article outline, inserts the extracted material, obtains the generated article, and generates an article outline according to the input article theme.
- the material of the inserted article outline is extremely rich, so the generated article is logically ordered, rich in form and content, close to the articles written by professionals, thus abandoning the limitations of current machine writing.
- FIG. 2 shows a schematic flow diagram of yet another embodiment of a method for generating an article in accordance with the present application.
- the method 200 for generating an article includes:
- an article outline is generated based on the input article theme and any of the following: an outline model; an outline database established based on user behavior data corresponding to the article theme; and an artificially set outline.
- the input article theme may be a machine topic or a manually input article theme.
- An outline model usually refers to a function with an article subject as an independent variable.
- you can set the article model f (topic, outline, material), that is, the article model is obtained from the independent variables (themes, outlines, and materials) in the function f, and by using the article model, one can be used for The method of generating the article, that is, selecting the theme, mining the outline and sorting through the outline model, and mounting the material through the material library; finally, the article is obtained through the drawing, typesetting, and retouching.
- the outline database established according to the user behavior data corresponding to the topic of the article refers to determining the article directory from the perspective of the article theme, and sorting and screening the article directory according to the user behavior data to obtain an outline database. It should be understood that the outline generated by the outline generation strategy herein has a certain logical order to ensure the rationality of the text.
- step 220 the material associated with the feature of the article outline is extracted from the pre-built material library.
- the pre-established material library refers to a material library obtained by establishing an index structure according to the characteristics of the material.
- the material can be extracted for later use.
- a predetermined number of materials whose features are most relevant to the features of the article outline may be extracted from the plurality of materials for later use.
- step 230 the extracted material is inserted into the article outline to obtain the generated article.
- the material extracted in step 220 may be inserted into the outline of the article obtained in step 210, thereby obtaining a generated article that is initially prototyped.
- step 240 the generated article is optimized to obtain an optimized generated article.
- the optimization process includes one or more of the following: a touch-up process, an insert rich media data process, and a typesetting optimization process.
- the generated article because there are different grammatical styles in the material library, and And the front and back connections may not be coherent, so the generated articles can be polished, that is, the grammar styles and sentences of the article are processed.
- the grammar that is, the writing regulations of the article, is generally used to refer to the complete sentence composed of words, words, short sentences, and sentences, and the rational organization of the articles.
- the style here refers to the performance that is unique to other articles, with a comprehensive overall characteristics.
- the retouching process includes one or more of the following: a grammatical style of the uniformly generated article; a statement that is inconsistent with the preceding and following statements; and a statement that replaces the inconsistent statement with the preceding and following statements.
- the grammatical style of the uniformly generated article can be realized by replacing and transforming a specific vocabulary and a specific sentence pattern, thereby making the grammatic style of the article consistent. Deleting a statement that is inconsistent with the preceding and following statements, or replacing a statement that is inconsistent with the preceding and following statements, can improve the inconsistency of the statement.
- inserting rich media data processing includes: extracting rich media data associated with features of the generated article from a pre-established resource library, and inserting the extracted rich into the generated article Media data.
- inserting the extracted rich media data into the generated article includes: first searching for rich media data according to one or more of the theme, the outline, the paragraph summary, and the keyword, and then selecting the quality through the quality screening.
- Rich media database and according to the number of words or paragraphs between pictures, to ensure that the inserted rich media data is relatively uniform. For example, if there are 1000 words between two images in the article and 10 words between the other two images, the inserted rich media data is not uniform and does not meet the reading habits of the user group.
- Rich media data is one or a combination of several types of programming languages that can include streaming media, sound, Flash, and Java, Javascript, dynamic HTML, and the like.
- Rich media data can be applied to a variety of web services, such as website design, email, banner for website pages, buttons, pop-up ads, interstitial ads, and more. It should be understood that rich media data can enhance information, and a more accurate orientation of information will have better interaction.
- extracting the rich media data associated with the features of the retouched article from the pre-established resource library includes: extracting from the pre-established resource library according to one or more of the following Rich media data generation candidate rich media list: article topic, article outline, summary of each paragraph of the polished article, and polished article Keywords for each paragraph; quality media is used to extract rich media data associated with the features of the retouched article from the list of candidate rich media.
- a rich media list is generated by extracting rich media data according to one or more of the article topic, the outline of the article, the abstract of each paragraph of the polished article, and the keywords of each paragraph of the polished article.
- the use of quality filtering to extract rich media data associated with the features of the retouched article from the rich media list can improve the quality of the rich media data in the repository.
- the pre-established resource library may be established by acquiring the feature of the rich media data, and establishing an index structure according to the feature of the rich media data to obtain the resource library.
- the foregoing quality screening may be performed according to one or more of the following: graphic relevance, image resolution, image aspect ratio, image source authority, advertisement filtering strategy, and anti- Cheating filtering strategy, anti-yellow filtering strategy and watermark filtering strategy.
- the advertisement filtering policy may include an advertisement filtering rule and an advertisement filtering model
- the anti-cheat filtering policy may include an anti-cheat filtering rule and an anti-cheat filtering model
- the anti-yellow filtering policy may include an anti-yellow filtering rule and an anti-yellow filtering model.
- the watermark filtering strategy may include a watermark filtering strategy and a watermark filtering model.
- the typesetting optimization process may be implemented by using a typographic optimization method in the prior art or a technology developed in the future, which is not limited in this application.
- the typographic optimization process can select an item that needs to be highlighted after determining various articles to be presented, and finally match the appropriate color layout to obtain an optimized article.
- the typesetting optimization process can also determine the typesetting corresponding to the generated article based on the analysis result of the article sample data and the behavior data of the user for the article sample data, thereby obtaining the optimized article.
- step 250 the article topic and article outline are entered into the title model to generate the title of the article.
- the article theme and the article outline can be input into the title model to generate the topic of the article.
- the heading model here is a function of the arguments for the topic of the article and the outline of the article.
- the topic of the article can be output.
- it may be a title model that the machine learns based on the article theme, the article outline, and the title of the article included in the existing article sample, or may be an artificially set title model.
- the method further includes: performing attribute expansion on the core word in the title; and replacing and rewriting the core word in the expanded title to obtain the updated title.
- the core word in the title may be first mined, then the core word is expanded, and the core word in the expanded title is replaced and rewritten to obtain the updated title.
- the core word in the title is XXX.
- the attribute of XXX can be obtained from the emperor who was born in Niu Wa. Therefore, the introduction of Emperor XXX can be replaced and rewritten as: Emperor born from Niu Wa. who is it?
- FIG. 2 is only an exemplary description of the method for generating an article in the embodiments of the present application, and does not represent a limitation on the present application.
- the method for generating an article in the embodiment of the present application may not include the above step 240, or may not include the above step 250, thereby obtaining a new method for generating an article.
- Step 210, step 220, and step 230 in FIG. 2 respectively correspond to step 110, step 120, and step 130 in FIG. 1, and therefore, the operations and features described in FIG. 1 for step 110, step 120, and step 130 are equally applicable.
- step 210, step 220 and step 230 details are not described herein again.
- the method for generating an article adds step 240 and step 250 by comparing with the method for generating an article described in FIG. 1, and according to step 240 and step 250, optimization can be obtained.
- the generated article and the title of the generated article so that the content of the generated article is more comprehensive, contains more information, the title of the article is more attractive, and the content and title of the article are more suitable for the user group. reading habit.
- the material 330 associated with the features of the article outlines 321 to 323 is extracted, including the following materials: material 331 "regime problem”, material 332 "who wants to squat", material 333 “wise decision”, material 334 "literatures can't make a difference”, material 335 "resistance outside the group”, material 336 “resistance within the group”, material 337 “external resistance”, material 338 "military war” and material 339 "most critical A little bit.” Then, into the article outline, the extracted material 330 (including the materials 331-339) is inserted to obtain the generated article.
- the generated article is polished 340, specifically including in step 341, unifying the style of the article, and in step 342, the consecutive sentence to obtain the polished article.
- the rich media 350 associated with the features of the retouched article is extracted, including the picture numbered 351, the picture 2 labeled 352, and the picture 3 labeled 353.
- the extracted rich media 350 is inserted into the retouched article to obtain an article after inserting the rich media; then, in the generating step of the title 360, the article theme and the article outline are input into the title model.
- the layout-optimized processing is performed on the article inserted into the rich media, for example, the specific operation 371 is performed, the key points are highlighted, and the color layout is adjusted, thereby obtaining the layout-optimized article.
- the operation 381 can be specifically performed to output the layout-optimized article.
- the method for generating an article provided in the above application scenario of the present application improves the efficiency of generating articles and enriches the content of the article, so that the generated article is consistent with the prior art in logic and grammar style. Form, content is richer and more reasonable.
- an embodiment of the present application provides an embodiment of an apparatus for generating an article, and an embodiment of the method for generating an article is shown in FIG. 1 to FIG.
- the embodiments of the method for generating an article correspond, whereby the operations and features described above with respect to the method for generating articles in FIGS. 1 through 3 are equally applicable to the device 400 for generating articles and the units contained therein , will not repeat them here.
- the apparatus 400 for generating an article includes: the apparatus includes: an outline generating unit 410, configured to generate an article outline based on the input article theme and any one of the following: an outline model; and a user according to the corresponding article theme An outline database of the behavior data creation; and an outline of the manual setting; the material extraction unit 420 is configured to extract the material associated with the feature of the article outline from the pre-established material library; the material insertion unit 430 is used to outline the article Insert the extracted material to get the generated article.
- an outline generating unit 410 configured to generate an article outline based on the input article theme and any one of the following: an outline model; and a user according to the corresponding article theme
- An outline database of the behavior data creation and an outline of the manual setting
- the material extraction unit 420 is configured to extract the material associated with the feature of the article outline from the pre-established material library
- the material insertion unit 430 is used to outline the article Insert the extracted material to get the generated article.
- the outline database established in the outline generation unit according to the user behavior data of the corresponding article theme includes: retrieving the sub-topics surrounding the article topic of the entire network, establishing a sub-topic database; and sub-topics in the sub-topic database according to the user Sub-topics in the sub-topic database in the order of clicks and/or sub-topics in the sub-topic database; sub-topics in the sub-topic database that do not conform to the predetermined logical rules are eliminated, and sub-topics that meet predetermined logical rules are obtained The sub-themes that meet the predetermined logic rules are used as an outline to obtain an outline database.
- the pre-established material library in the material extraction unit is established by: acquiring features of the material, the material is filtering the content of the existing article according to the screening rules and/or transforming the content of the existing article. Get; build an index structure based on the characteristics of the material to get the material library.
- the apparatus further includes: an article optimization unit 440, configured to perform optimization processing on the generated article to obtain an optimized generated article, and the optimization process includes one or more of the following: retouching processing, inserting rich media data Processing and layout optimization processing.
- the retouching process in the article optimization unit includes one or more of the following: a grammatical style of the uniformly generated article; a statement that is inconsistent with the preceding and following statements; and a statement that replaces the inconsistent statement with the preceding and following statements.
- the inserting rich media data processing in the article optimization unit includes: extracting rich media data associated with features of the generated article from a pre-established resource library; and inserting the extracted rich media into the generated article data.
- extracting, from the pre-established resource library, the rich media data associated with the features of the generated article in the article optimization unit comprises: extracting rich media from the pre-established resource library according to one or more of the following: Data generation candidate rich media list: article topic, article outline, summary of each paragraph of the generated article, and paragraphs of the generated article Key words; quality filtering is used to extract rich media data associated with the features of the generated article from the list of candidate rich media.
- the pre-established resource library in the article optimization unit is established by acquiring features of the rich media data, and establishing an index structure according to the characteristics of the rich media data to obtain a resource library.
- the quality screening in the article optimization unit is performed according to one or more of the following: graphic relevance, image resolution, image aspect ratio, image source authority, advertisement filtering strategy, anti-cheat filtering strategy, Anti-yellow filtering strategy and watermark filtering strategy.
- the apparatus further includes: a title generating unit 450, configured to input the article theme and the article outline into the title model to obtain a title of the generated article.
- a title generating unit 450 configured to input the article theme and the article outline into the title model to obtain a title of the generated article.
- the apparatus further includes: an attribute extension unit (not shown) for performing attribute expansion on the core words in the title; and a title update unit (not shown) for expanding the attributes The core words in the title are replaced and rewritten to get the updated title.
- the application also provides an embodiment of a device comprising: one or more processors; a storage device for storing one or more programs; and one or more programs being executed by one or more processors such that one Or a plurality of processors implement the method for generating an article as described above.
- the present application also provides an embodiment of a computer readable storage medium having stored thereon a computer program that, when executed by a processor, implements the method for generating an article as described above.
- FIG. 5 there is shown a block diagram of a computer system 500 suitable for use in implementing a terminal device or server of an embodiment of the present application.
- the terminal device shown in FIG. 5 is merely an example, and should not impose any limitation on the function and scope of use of the embodiments of the present application.
- computer system 500 includes a central processing unit (CPU) 501 that can be loaded into a program in random access memory (RAM) 503 according to a program stored in read only memory (ROM) 502 or from storage portion 508. And perform various appropriate actions and processes.
- RAM random access memory
- ROM read only memory
- RAM 503 various programs and data required for the operation of the system 500 are also stored.
- the CPU 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504.
- An input/output (I/O) interface 505 is also coupled to bus 504.
- the following components are connected to the I/O interface 505: an input portion 506 including a keyboard, a mouse, etc.; an output portion 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a storage portion 508 including a hard disk or the like. And a communication portion 509 including a network interface card such as a LAN card, a modem, or the like. The communication section 509 performs communication processing via a network such as the Internet.
- Driver 510 is also coupled to I/O interface 505 as needed.
- a removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory or the like is mounted on the drive 510 as needed so that a computer program read therefrom is installed into the storage portion 508 as needed.
- an embodiment of the present disclosure includes a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for executing the method illustrated by the flowchart.
- the computer program can be downloaded and installed from the network via the communication portion 509, and/or installed from the removable medium 511.
- CPU central processing unit
- the computer readable medium described herein may be a computer readable signal medium or a computer readable storage medium or any combination of the two.
- the computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of computer readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read only memory (ROM), erasable Programmable read only memory (EPROM or flash memory), optical fiber, portable compact disk read only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
- a computer readable storage medium may be any tangible medium that can contain or store a program, which can be used by or in connection with an instruction execution system, apparatus or device.
- a computer readable signal medium may include a data signal that is propagated in the baseband or as part of a carrier, carrying computer readable program code. Such propagated data signals can take a variety of forms including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the foregoing.
- the computer readable signal medium can also be any computing other than a computer readable storage medium
- Program code embodied on a computer readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, fiber optic cable, RF, etc., or any suitable combination of the foregoing.
- each block of the flowchart or block diagrams can represent a unit, a program segment, or a portion of code that includes one or more logic for implementing the specified.
- Functional executable instructions can also occur in a different order than that illustrated in the drawings. For example, two successively represented blocks may in fact be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending upon the functionality involved.
- each block of the block diagrams and/or flowcharts, and combinations of blocks in the block diagrams and/or flowcharts can be implemented in a dedicated hardware-based system that performs the specified function or operation. Or it can be implemented by a combination of dedicated hardware and computer instructions.
- the units involved in the embodiments of the present application may be implemented by software or by hardware.
- the described unit may also be disposed in the processor, for example, as a processor including an outline generation unit, a material extraction unit, and a material insertion unit.
- the names of these units do not constitute a limitation on the unit itself under certain circumstances.
- the outline generation unit may also be described as “a unit that generates an article outline based on an input article theme and an outline generation strategy”.
- the present application further provides a non-volatile computer storage medium, which may be a non-volatile computer storage medium included in the apparatus described in the foregoing embodiments; It may be a non-volatile computer storage medium that exists alone and is not assembled into the terminal.
- the non-volatile computer storage medium stores one or more programs, when the one or more programs are executed by a device, causing the device to: generate an article outline based on the input article theme and any of the following: outline a model; an outline database established based on user behavior data corresponding to the topic of the article; and an artificially set outline; extracting material associated with the feature of the article outline from the pre-established material library; and inserting the extracted material into the article outline , get the generated article.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Databases & Information Systems (AREA)
- Data Mining & Analysis (AREA)
- Software Systems (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Machine Translation (AREA)
Abstract
本申请公开了用于生成文章的方法和装置。方法的一具体实施方式包括:基于输入的文章主题和以下任意一项生成文章提纲:提纲模型;根据对应文章主题的用户行为数据建立的提纲数据库;以及人工设定的提纲;从预先建立的素材库中,提取与文章提纲的特征相关联的素材;向文章提纲中,插入提取的素材,得到生成的文章。该实施方式实现了根据输入的文章主题生成提纲,提高了文章提纲的质量,保证了生成的文章的行文逻辑合理、形式丰富,并根据文章提纲插入与文章提纲的特征相关联的素材,丰富了文章的内容,从而使得生成的文章逻辑合理并且形式、内容丰富。
Description
相关申请的交叉引用
本申请要求于2017年3月31日提交的中国专利申请号为“201710206961.1”的优先权,其全部内容作为整体并入本申请中。
本申请涉及计算机技术领域,具体涉及计算机网络技术领域,尤其涉及用于生成文章的方法和装置。
目前,通过机器实现自动化写作,从而生成文章的方式,基本停留在特殊领域的特殊题材,多是采用将素材填入规则或模板的技术来生成文章。例如,可以筛选原始文章后直接引用;或者,对原始文章进行简单的变换直接发布;或者,将原始文章以一定的顺序进行组合和摘要提取;又或者,通过模板对数据进行组织并展现。
然而,目前的生成文章的方式,由于题材和方法的限制,产出的文章形式和内容比较单调,并且行文可能出现前后逻辑不合理、文法风格不一致等情况,机器写作的痕迹较重。
发明内容
本申请的目的在于提出一种改进的用于生成文章的方法和装置,来解决以上背景技术部分提到的技术问题。
第一方面,本申请实施例提供了一种用于生成文章的方法,方法包括:基于输入的文章主题和以下任意一项生成文章提纲:提纲模型,根据对应文章主题的用户行为数据建立的提纲数据库,以及人工设定的提纲;从预先建立的素材库中,提取与文章提纲的特征相关联的素材;向文章提纲中,插入提取的素材,得到生成的文章。
在一些实施例中,根据对应文章主题的用户行为数据建立的提纲数据库包括:检索全网围绕文章主题的子主题,建立子主题数据库;根据用户对子主题数据库中的子主题的点击顺序和/或子主题数据库中的子主题的语义递进顺序,排序子主题数据库中的子主题;剔除子主题数据库中不符合预定逻辑规则的子主题,得到符合预定逻辑规则的子主题;将各符合预定逻辑规则的子主题作为提纲,得到提纲数据库。
在一些实施例中,预先建立的素材库通过以下步骤建立:获取素材的特征,素材为将现有的文章的内容根据筛选规则筛选得到和/或变换现有的文章的内容得到;根据素材的特征建立索引结构,得到素材库。
在一些实施例中,方法还包括:对生成的文章进行优化处理,得到优化后的生成的文章,优化处理包括以下一项或多项:润色处理、插入富媒体数据处理以及排版优化处理。
在一些实施例中,润色处理包括以下一项或多项:统一生成的文章的文法风格;删除与前后语句不连贯的语句;以及替换与前后语句不连贯的语句。
在一些实施例中,插入富媒体数据处理包括:从预先建立的资源库,提取与生成的文章的特征相关联的富媒体数据;向生成的文章中,插入提取的富媒体数据。
在一些实施例中,从预先建立的资源库,提取与生成的文章的特征相关联的富媒体数据包括:根据以下一项或多项从预先建立的资源库中提取富媒体数据生成候选富媒体列表:文章主题、文章提纲、生成的文章的各段落的摘要以及生成的文章的各段落的关键词;采用质量筛选从候选富媒体列表中提取与生成的文章的特征相关联的富媒体数据。
在一些实施例中,预先建立的资源库通过以下步骤建立:获取富媒体数据的特征;根据富媒体数据的特征建立索引结构,得到资源库。
在一些实施例中,质量筛选根据以下一项或多项进行:图文相关性、图片分辨率、图片长宽比、图片来源权威度、广告过滤策略、反
作弊过滤策略、反黄过滤策略和水印过滤策略。
在一些实施例中,方法还包括:将文章主题和文章提纲输入标题模型,得到生成的文章的标题。
在一些实施例中,方法还包括:对标题中的核心词进行属性扩展;对属性扩展后的标题中的核心词进行替换和改写,得到更新后的标题。
第二方面,本申请实施例提供了一种用于生成文章的装置,装置包括:提纲生成单元,用于基于输入的文章主题和以下任意一项生成文章提纲:提纲模型,根据对应文章主题的用户行为数据建立的提纲数据库,以及人工设定的提纲;素材提取单元,用于从预先建立的素材库中,提取与文章提纲的特征相关联的素材;素材插入单元,用于向文章提纲中,插入提取的素材,得到生成的文章。
在一些实施例中,提纲生成单元中的根据对应文章主题的用户行为数据建立的提纲数据库包括:检索全网围绕文章主题的子主题,建立子主题数据库;根据用户对子主题数据库中的子主题的点击顺序和/或子主题数据库中的子主题的语义递进顺序,排序子主题数据库中的子主题;剔除子主题数据库中不符合预定逻辑规则的子主题,得到符合预定逻辑规则的子主题;将各符合预定逻辑规则的子主题作为提纲,得到提纲数据库。
在一些实施例中,素材提取单元中的预先建立的素材库通过以下步骤建立:获取素材的特征,素材为将现有的文章的内容根据筛选规则筛选得到和/或变换现有的文章的内容得到;根据素材的特征建立索引结构,得到素材库。
在一些实施例中,装置还包括:文章优化单元,用于对生成的文章进行优化处理,得到优化后的生成的文章,优化处理包括以下一项或多项:润色处理、插入富媒体数据处理以及排版优化处理。
在一些实施例中,文章优化单元中的润色处理包括以下一项或多项:统一生成的文章的文法风格;删除与前后语句不连贯的语句;以及替换与前后语句不连贯的语句。
在一些实施例中,文章优化单元中的插入富媒体数据处理包括:从预先建立的资源库,提取与生成的文章的特征相关联的富媒体数据;
向生成的文章中,插入提取的富媒体数据。
在一些实施例中,文章优化单元中的从预先建立的资源库,提取与生成的文章的特征相关联的富媒体数据包括:根据以下一项或多项从预先建立的资源库中提取富媒体数据生成候选富媒体列表:文章主题、文章提纲、生成的文章的各段落的摘要以及生成的文章的各段落的关键词;采用质量筛选从候选富媒体列表中提取与生成的文章的特征相关联的富媒体数据。
在一些实施例中,文章优化单元中的预先建立的资源库通过以下步骤建立:获取富媒体数据的特征;根据富媒体数据的特征建立索引结构,得到资源库。
在一些实施例中,文章优化单元中的质量筛选根据以下一项或多项进行:图文相关性、图片分辨率、图片长宽比、图片来源权威度、广告过滤策略、反作弊过滤策略、反黄过滤策略和水印过滤策略。
在一些实施例中,装置还包括:标题生成单元,用于将文章主题和文章提纲输入标题模型,得到生成的文章的标题。
在一些实施例中,装置还包括:属性扩展单元,用于对标题中的核心词进行属性扩展;标题更新单元,用于对属性扩展后的标题中的核心词进行替换和改写,得到更新后的标题。
第三方面,本申请实施例提供了一种设备,包括:一个或多个处理器;存储装置,用于存储一个或多个程序;当一个或多个程序被一个或多个处理器执行,使得一个或多个处理器实现如上任一所述的用于生成文章的方法。
第四方面,本申请实施例提供了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现如上任一所述的用于生成文章的方法。
本申请实施例提供的用于生成文章的方法和装置,首先基于输入的文章主题和以下任意一项生成文章提纲:提纲模型;根据对应文章主题的用户行为数据建立的提纲数据库;以及人工设定的提纲;之后,从预先建立的素材库中,提取与文章提纲的特征相关联的素材;之后,向文章提纲中,插入提取的素材,得到生成的文章。本实施例实现了
根据输入的文章主题生成提纲,提高了文章提纲的质量,保证了生成的文章的行文逻辑合理、形式丰富,并根据文章提纲插入与文章提纲的特征相关联的素材,丰富了文章的内容,从而使得生成的文章逻辑合理并且内容丰富。
通过阅读参照以下附图所作的对非限制性实施例所作的详细描述,本申请的其它特征、目的和优点将会变得更明显:
图1是根据本申请的用于生成文章的方法的一个实施例的示意性流程图;
图2是根据本申请的用于生成文章的方法的又一个实施例的示意性流程图;
图3是应用本申请的用于生成文章的方法的一个实施例的示例性应用场景;
图4是根据本申请的用于生成文章的装置的一个实施例的示例性结构图;
图5是适于用来实现本申请实施例的终端设备或服务器的计算机系统的结构示意图。
下面结合附图和实施例对本申请作进一步的详细说明。可以理解的是,此处所描述的具体实施例仅仅用于解释相关发明,而非对该发明的限定。另外还需要说明的是,为了便于描述,附图中仅示出了与有关发明相关的部分。
需要说明的是,在不冲突的情况下,本申请中的实施例及实施例中的特征可以相互组合。下面将参考附图并结合实施例来详细说明本申请。
图1示出了根据本申请的用于生成文章的方法的一个实施例的流程100。该用于生成文章的方法包括:
在步骤110中,基于输入的文章主题和以下任意一项生成文章提
纲:提纲模型;根据对应文章主题的用户行为数据建立的提纲数据库;以及人工设定的提纲。
在本实施例中,输入的文章主题可以为机器挖掘或人工输入的文章主题。
提纲模型通常是指以文章主题为自变量的函数。首先,可以设定文章模型=f(主题,提纲,素材),也即文章模型由函数f中的自变量(主题、提纲和素材)得到,并借由该文章模型,可以得到一种用于生成文章的方法,即选定主题,通过提纲模型挖掘提纲并排序,通过素材库来挂载素材;最后通过配图、排版和润色得到文章。
根据对应文章主题的用户行为数据建立的提纲数据库是指从文章主题角度确定文章目录,并根据用户行为数据对文章目录进行合理排序和筛选,得到提纲数据库。应当理解,这里的提纲生成策略生成的提纲具有一定的逻辑顺序,以保障行文的合理性。
在本实施例的一些可选实现方式中,上述的根据对应所述文章主题的用户行为数据建立的提纲数据库包括:检索全网围绕文章主题的子主题,建立子主题数据库;根据用户对子主题数据库中的子主题的点击顺序和/或子主题数据库中的子主题的语义递进顺序,排序子主题数据库中的子主题;剔除子主题数据库中不符合预定逻辑规则的子主题,得到符合预定逻辑规则的子主题;将各符合预定逻辑规则的子主题作为提纲,得到提纲数据库。
在本实现方式中,根据对应所述文章主题的用户行为数据建立的提纲数据库,充分考虑了用户的行为数据来建立提纲,可以提高建立的提纲的针对性,进而增强了生成的文章与用户的交互能力。
在步骤120中,从预先建立的素材库中,提取与文章提纲的特征相关联的素材。
在本实施例中,预先建立的素材库,是指根据素材的特征建立索引结构得到的素材库。当素材的特征与文章提纲的特征相关联时,可以提取该素材以备后续使用。当多个素材的特征均与文章提纲的特征相关联时,可以从多个素材中,提取特征与文章提纲的特征最为相关的预定数量个素材,以备后续使用。
在本实施例的一些可选实现方式中,预先建立的素材库通过以下步骤建立:获取素材的特征,素材为将现有的文章的内容根据筛选规则筛选得到和/或变换现有的文章的内容得到;根据素材的特征建立索引结构,得到素材库。
在本实现方式中,素材库的生成包括有明确主题的素材和无明确主题的素材,后者需要使用文章摘要技术提取主题。获取素材的特征,可以理解为从文本素材中提取特征,这些特征可以说明文本素材的主题、关键词、核心语义等信息,用于和文章提纲、文章主题进行相关性计算和排序。
具体地,上述的根据筛选规则筛选得到可以包括根据以下一项或多项内容进行筛选:文章的内容长度、文章的内容质量评分、文章的内容满意度评分、文章的浏览量、文章的时效性等等。而上述的变换现有的文章内容主要是为了控制素材的粒度,可以采用预定规则来完成变换。例如,将字数大于预定值的段落进行拆解分段。假设一个素材是一篇原始语料,筛选后根据提纲排序组合就可以了;假设一个素材是一段,就需要考虑段落的主题相关性、段落间排序等;同理,还可以假设素材是一句话、一个词,当素材的粒度越小时,拆解和/或变换的难度越大。
在步骤130中,向文章提纲中,插入提取的素材,得到生成的文章。
在本实施例中,可以向步骤110中得到的文章提纲中,插入步骤120中提取的素材,从而得到生成的文章。
本申请的上述实施例提供的用于生成文章的方法,通过生成文章提纲,提取与文章提纲的特征相关联的素材,插入提取的素材,得到生成的文章,可以根据输入的文章主题生成文章提纲,并且插入文章提纲的素材极为丰富,因此生成的文章的逻辑顺序合理、形式和内容更为丰富,接近于专业人士写的文章,从而摒弃了目前机器写作的局限性。
进一步参考图2,图2示出了根据本申请的用于生成文章的方法的又一个实施例的示意性流程图。该用于生成文章的方法200包括:
在步骤210中,基于输入的文章主题和以下任意一项生成文章提纲:提纲模型;根据对应文章主题的用户行为数据建立的提纲数据库;以及人工设定的提纲。
在本实施例中,在本实施例中,输入的文章主题可以为机器挖掘或人工输入的文章主题。
提纲模型通常是指以文章主题为自变量的函数。首先,可以设定文章模型=f(主题,提纲,素材),也即文章模型由函数f中的自变量(主题、提纲和素材)得到,并借由该文章模型,可以得到一种用于生成文章的方法,即选定主题,通过提纲模型挖掘提纲并排序,通过素材库来挂载素材;最后通过配图、排版和润色得到文章。
根据对应文章主题的用户行为数据建立的提纲数据库是指从文章主题角度确定文章目录,并根据用户行为数据对文章目录进行合理排序和筛选,得到提纲数据库。应当理解,这里的提纲生成策略生成的提纲具有一定的逻辑顺序,以保障行文的合理性。
在步骤220中,从预先建立的素材库中,提取与文章提纲的特征相关联的素材。
在本实施例中,预先建立的素材库,是指根据素材的特征建立索引结构得到的素材库。当素材的特征与文章提纲的特征相关联时,可以提取该素材以备后续使用。当多个素材的特征均与文章提纲的特征相关联时,可以从多个素材中,提取特征与文章提纲的特征最为相关的预定数量个素材,以备后续使用。
在步骤230中,向文章提纲中,插入提取的素材,得到生成的文章。
在本实施例中,可以向步骤210中得到的文章提纲中,插入步骤220中提取的素材,从而得到初具雏形的生成的文章。
在步骤240中,对生成的文章进行优化处理,得到优化后的生成的文章。
在本实施例中,优化处理包括以下一项或多项:润色处理、插入富媒体数据处理以及排版优化处理。
对于生成的文章,由于素材库中存在不同的文法风格的素材,并
且前后连接处可能并不连贯,因此可以对生成的文章进行润色处理,也即对文章的文法风格和语句等进行处理。这里的文法,即文章的书写法规,一般用来指以文字、词语、短句、句子的编排而组成的完整语句和文章的合理性组织。这里的风格,是指具有独特于其他文章的表现,带有综合性的总体特点。
在本实施例的一些可选实现方式中,进行润色处理包括以下一项或多项:统一生成的文章的文法风格;删除与前后语句不连贯的语句;以及替换与前后语句不连贯的语句。
在本实现方式中,统一生成的文章的文法风格,可以通过对于特定词汇、特定句式的替换和变换实现,从而使得文章的文法风格一致。而删除与前后语句不连贯的语句,或者替换与前后语句不连贯的语句,均可改善语句的不连贯现象。
在本实施例的一些可选实现方式中,插入富媒体数据处理包括:从预先建立的资源库,提取与生成的文章的特征相关联的富媒体数据,向生成的文章中,插入提取的富媒体数据。
在本实施例中,向生成的文章中,插入提取的富媒体数据包括:首先根据主题、提纲、段落摘要和关键词中的一项或多项查找富媒体数据,之后通过质量筛选挑选出优质富媒体数据库,并根据图片间字数或段落数,保证插入的富媒体数据相对均匀。例如,若文章中有两张图之间1000字,而另外两个图间10个字,那么插入的富媒体数据不均匀,并不符合用户群体的阅读习惯。富媒体数据为可以包含流媒体、声音、Flash、以及Java、Javascript、动态的HTML等程序设计语言的形式之一或者几种的组合。富媒体数据可以应用于各种网络服务中,如网站设计、电子邮件、网站页面的横幅、按钮、弹出式广告、插播式广告等。应当理解,富媒体数据可以加强信息,而信息更准确的定向会具有更好的交互效果。
在本实施例的一些可选实现方式中,从预先建立的资源库,提取与润色后的文章的特征相关联的富媒体数据包括:根据以下一项或多项从预先建立的资源库中提取富媒体数据生成候选富媒体列表:文章主题、文章提纲、润色后的文章的各段落的摘要以及润色后的文章的
各段落的关键词;采用质量筛选从候选富媒体列表中提取与润色后的文章的特征相关联的富媒体数据。
在本实现方式中,通过根据文章主题、文章提纲、润色后的文章的各段落的摘要以及润色后的文章的各段落的关键词中的一项或多项提取富媒体数据,生成富媒体列表;之后采用质量筛选从富媒体列表中提取与润色后的文章的特征相关联的富媒体数据,可以提高资源库中的富媒体数据的质量。
在本实施例的一些可选实现方式中,预先建立的资源库可以通过以下步骤建立:获取富媒体数据的特征;根据富媒体数据的特征建立索引结构,得到资源库。
在本实施例的一些可选实现方式中,上述的质量筛选可以根据以下一项或多项进行:图文相关性、图片分辨率、图片长宽比、图片来源权威度、广告过滤策略、反作弊过滤策略、反黄过滤策略和水印过滤策略。
在本实现方式中,广告过滤策略可以包括广告过滤规则和广告过滤模型;反作弊过滤策略可以包括反作弊过滤规则和反作弊过滤模型;反黄过滤策略可以包括反黄过滤规则和反黄过滤模型;水印过滤策略则可以包括水印过滤策略和水印过滤模型。
在本实施例中,排版优化处理可以采用现有技术或未来发展的技术中的排版优化方法来完成,本申请对此不做限定。例如,排版优化处理可以为在确定各种需要呈现的文章内容之后,选择需要重点突出的内容,最后搭配恰当的颜色版式,从而得到优化后的文章。这里的排版优化处理,也可以根据对文章样本数据和用户针对文章样本数据的行为数据的分析结果来确定与生成的文章相适应的排版,从而得到优化后的文章。
在步骤250中,将文章主题和文章提纲输入标题模型,生成文章的标题。
在本实施例中,在得到生成的文章之后,可以将文章主题和文章提纲输入标题模型,以便生成文章的主题。这里的标题模型,是自变量为文章主题和文章提纲的函数,当接收到文章主题和文章提纲时,
根据该函数即可输出文章的主题。例如,可以为机器根据现有的文章样本中包括的文章主题、文章提纲和文章的标题学习得到的标题模型,也可以为人为设定的标题模型。
在本实施例的一些可选实现方式中,方法还包括:对标题中的核心词进行属性扩展;对属性扩展后的标题中的核心词进行替换和改写,得到更新后的标题。
在本实现方式中,可以首先挖掘标题中的核心词,之后对核心词进行属性扩展,再对属性扩展后的标题中的核心词进行替换和改写,得到更新后的标题。例如,对于皇帝XXX的介绍,挖掘出标题中的核心词为XXX,之后可以得到XXX的属性是放牛娃出身的皇帝,因此可以将皇帝XXX的介绍替换和改写为:放牛娃出身的皇帝是谁?
应当理解,上述图2中的描述仅为本申请实施例的用于生成文章的方法的一个示例性描述,并不代表对本申请的限定。例如,本申请实施例中的用于生成文章的方法,也可以不包括上述步骤240,或者不包括上述步骤250,从而得到新的用于生成文章的方法。图2中的步骤210、步骤220和步骤230分别与图1中的步骤110、步骤120和步骤130相对应,因此,图1中针对步骤110、步骤120和步骤130描述的操作和特征同样适用于步骤210、步骤220和步骤230,在此不再赘述。
本申请的上述实施例提供的用于生成文章的方法,通过与图1中描述的用于生成文章的方法相比,增加了步骤240和步骤250,根据步骤240和步骤250,可以得到优化后的生成的文章以及得到生成的文章的标题,从而使得生成的文章的内容更为全面,包含的信息更为丰富,文章的标题更具有吸引力,并且文章的内容和标题更为适应用户群体的阅读习惯。
以下结合图3,描述本申请实施例的用于生成文章的方法的一个示例性应用场景。
如图3所示,根据本申请实施例的用于生成文章的方法,首先,根据输入的文章主题310的具体实施例311“诸葛亮称帝”,可以生成文章提纲320的具体实施例,也即包括提纲321:刘备托孤时为什
么让诸葛亮称帝;提纲322:诸葛亮为什么不称帝;以及提纲323:诸葛亮如果称帝会怎么样。之后,从预先建立的素材库中,提取与文章提纲321至323的特征相关联的素材330,包括以下素材:素材331“政权问题”、素材332“欲擒故纵”、素材333“明智决定”、素材334“文人是造不了反的”、素材335“集团外部的阻力”、素材336“集团内部的阻力”、素材337“外部方面的阻力”、素材338“兵民厌战”以及素材339“最关键的一点”。之后,向文章提纲中,插入提取的素材330(包括素材331-339),得到生成的文章。之后,对生成的文章进行润色340,具体包括在步骤341中,统一文章的文风,以及在步骤342中,连贯语句,得到润色后的文章。然后,从预先建立的资源库,提取与润色后的文章的特征相关联的富媒体350,包括标号为351的图片1、标号为352的图片2以及标号为353的图片3。之后,向润色后的文章中,插入提取的富媒体350(包括富媒体351-353),得到插入富媒体后的文章;之后,在标题360的生成步骤,将文章主题和文章提纲输入标题模型,得到初始标题,并对初始标题中的核心词进行属性扩展,对属性扩展后的初始标题中的核心词进行替换和改写,得到更新后的标题361“有颜有实力,集尽万千追捧的男神为何终未加冕?”。之后,在排版370的处理步骤中,对插入富媒体后的文章进行排版优化处理,例如进行具体操作371,突出重点,并进行颜色版式调整,从而得到排版优化后的文章。最后,在输出380的处理步骤中,可以具体进行操作381,输出排版优化后的文章。
本申请的上述应用场景中提供的用于生成文章的方法,提高了文章的生成效率,并丰富了文章的内容,使得生成的文章的行文与现有技术相比,前后逻辑、文法风格一致,形式、内容更为丰富且更为合理。
进一步参考图4,作为对上述方法的实现,本申请实施例提供了一种用于生成文章的装置的一个实施例,该用于生成文章的方法的实施例与图1至图3所示的用于生成文章的方法的实施例相对应,由此,上文针对图1至图3中用于生成文章的方法描述的操作和特征同样适用于用于生成文章的装置400及其中包含的单元,在此不再赘述。
如图4所示,该配置用于生成文章的装置400包括:装置包括:提纲生成单元410,用于基于输入的文章主题和以下任意一项生成文章提纲:提纲模型;根据对应文章主题的用户行为数据建立的提纲数据库;以及人工设定的提纲;素材提取单元420,用于从预先建立的素材库中,提取与文章提纲的特征相关联的素材;素材插入单元430,用于向文章提纲中,插入提取的素材,得到生成的文章。
在一些实施例中,提纲生成单元中的根据对应文章主题的用户行为数据建立的提纲数据库包括:检索全网围绕文章主题的子主题,建立子主题数据库;根据用户对子主题数据库中的子主题的点击顺序和/或子主题数据库中的子主题的语义递进顺序,排序子主题数据库中的子主题;剔除子主题数据库中不符合预定逻辑规则的子主题,得到符合预定逻辑规则的子主题;将各符合预定逻辑规则的子主题作为提纲,得到提纲数据库。
在一些实施例中,素材提取单元中的预先建立的素材库通过以下步骤建立:获取素材的特征,素材为将现有的文章的内容根据筛选规则筛选得到和/或变换现有的文章的内容得到;根据素材的特征建立索引结构,得到素材库。
在一些实施例中,装置还包括:文章优化单元440,用于对生成的文章进行优化处理,得到优化后的生成的文章,优化处理包括以下一项或多项:润色处理、插入富媒体数据处理以及排版优化处理。
在一些实施例中,文章优化单元中的润色处理包括以下一项或多项:统一生成的文章的文法风格;删除与前后语句不连贯的语句;以及替换与前后语句不连贯的语句。
在一些实施例中,文章优化单元中的插入富媒体数据处理包括:从预先建立的资源库,提取与生成的文章的特征相关联的富媒体数据;向生成的文章中,插入提取的富媒体数据。
在一些实施例中,文章优化单元中的从预先建立的资源库,提取与生成的文章的特征相关联的富媒体数据包括:根据以下一项或多项从预先建立的资源库中提取富媒体数据生成候选富媒体列表:文章主题、文章提纲、生成的文章的各段落的摘要以及生成的文章的各段落
的关键词;采用质量筛选从候选富媒体列表中提取与生成的文章的特征相关联的富媒体数据。
在一些实施例中,文章优化单元中的预先建立的资源库通过以下步骤建立:获取富媒体数据的特征;根据富媒体数据的特征建立索引结构,得到资源库。
在一些实施例中,文章优化单元中的质量筛选根据以下一项或多项进行:图文相关性、图片分辨率、图片长宽比、图片来源权威度、广告过滤策略、反作弊过滤策略、反黄过滤策略和水印过滤策略。
在一些实施例中,装置还包括:标题生成单元450,用于将文章主题和文章提纲输入标题模型,得到生成的文章的标题。
在一些实施例中,装置还包括:属性扩展单元(图中未示出),用于对标题中的核心词进行属性扩展;标题更新单元(图中未示出),用于对属性扩展后的标题中的核心词进行替换和改写,得到更新后的标题。
本申请还提供了一种设备的实施例,包括:一个或多个处理器;存储装置,用于存储一个或多个程序;当一个或多个程序被一个或多个处理器执行,使得一个或多个处理器实现如上任一所述的用于生成文章的方法。
本申请还提供了一种计算机可读存储介质的实施例,其上存储有计算机程序,该程序被处理器执行时实现如上任一所述的用于生成文章的方法。
下面参考图5,其示出了适于用来实现本申请实施例的终端设备或服务器的计算机系统500的结构示意图。图5示出的终端设备仅仅是一个示例,不应对本申请实施例的功能和使用范围带来任何限制。
如图5所示,计算机系统500包括中央处理单元(CPU)501,其可以根据存储在只读存储器(ROM)502中的程序或者从存储部分508加载到随机访问存储器(RAM)503中的程序而执行各种适当的动作和处理。在RAM 503中,还存储有系统500操作所需的各种程序和数据。CPU 501、ROM 502以及RAM 503通过总线504彼此相连。输入/输出(I/O)接口505也连接至总线504。
以下部件连接至I/O接口505:包括键盘、鼠标等的输入部分506;包括诸如阴极射线管(CRT)、液晶显示器(LCD)等以及扬声器等的输出部分507;包括硬盘等的存储部分508;以及包括诸如LAN卡、调制解调器等的网络接口卡的通信部分509。通信部分509经由诸如因特网的网络执行通信处理。驱动器510也根据需要连接至I/O接口505。可拆卸介质511,诸如磁盘、光盘、磁光盘、半导体存储器等等,根据需要安装在驱动器510上,以便于从其上读出的计算机程序根据需要被安装入存储部分508。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在计算机可读介质上的计算机程序,所述计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信部分509从网络上被下载和安装,和/或从可拆卸介质511被安装。在该计算机程序被中央处理单元(CPU)501执行时,执行本申请的方法中限定的上述功能。
需要说明的是,本申请所述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本申请中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本申请中,计算机可读的信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读的信号介质还可以是计算机可读存储介质以外的任何计算
机可读介质,该计算机可读介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:无线、电线、光缆、RF等等,或者上述的任意合适的组合。
附图中的流程图和框图,图示了按照本申请各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个单元、程序段、或代码的一部分,所述单元、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本申请实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。所描述的单元也可以设置在处理器中,例如,可以描述为:一种处理器包括提纲生成单元、素材提取单元、素材插入单元。其中,这些单元的名称在某种情况下并不构成对该单元本身的限定,例如,提纲生成单元还可以被描述为“基于输入的文章主题和提纲生成策略,生成文章提纲的单元”。
作为另一方面,本申请还提供了一种非易失性计算机存储介质,该非易失性计算机存储介质可以是上述实施例中所述装置中所包含的非易失性计算机存储介质;也可以是单独存在,未装配入终端中的非易失性计算机存储介质。上述非易失性计算机存储介质存储有一个或者多个程序,当所述一个或者多个程序被一个设备执行时,使得所述设备:基于输入的文章主题和以下任意一项生成文章提纲:提纲模型;根据对应文章主题的用户行为数据建立的提纲数据库;以及人工设定的提纲;从预先建立的素材库中,提取与文章提纲的特征相关联的素材;向文章提纲中,插入提取的素材,得到生成的文章。
以上描述仅为本申请的较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本申请中所涉及的发明范围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述发明构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上述特征与本申请中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。
Claims (24)
- 一种用于生成文章的方法,其特征在于,所述方法包括:基于输入的文章主题和以下任意一项生成文章提纲:提纲模型,根据对应所述文章主题的用户行为数据建立的提纲数据库,以及人工设定的提纲;从预先建立的素材库中,提取与所述文章提纲的特征相关联的素材;向所述文章提纲中,插入提取的素材,得到生成的文章。
- 根据权利要求1所述的方法,其特征在于,所述根据对应所述文章主题的用户行为数据建立的提纲数据库包括:检索全网围绕所述文章主题的子主题,建立子主题数据库;根据用户对所述子主题数据库中的子主题的点击顺序和/或所述子主题数据库中的子主题的语义递进顺序,排序所述子主题数据库中的子主题;剔除所述子主题数据库中不符合预定逻辑规则的子主题,得到符合预定逻辑规则的子主题;将各符合预定逻辑规则的子主题作为提纲,得到提纲数据库。
- 根据权利要求1所述的方法,其特征在于,所述预先建立的素材库通过以下步骤建立:获取素材的特征,所述素材为将现有的文章的内容根据筛选规则筛选得到和/或变换现有的文章的内容得到;根据所述素材的特征建立索引结构,得到所述素材库。
- 根据权利要求1所述的方法,其特征在于,所述方法还包括:对所述生成的文章进行优化处理,得到优化后的所述生成的文章,所述优化处理包括以下一项或多项:润色处理、插入富媒体数据处理以及排版优化处理。
- 根据权利要求4所述的方法,其特征在于,所述润色处理包括以下一项或多项:统一所述生成的文章的文法风格;删除与前后语句不连贯的语句;以及替换与前后语句不连贯的语句。
- 根据权利要求4所述的方法,其特征在于,所述插入富媒体数据处理包括:从预先建立的资源库,提取与所述生成的文章的特征相关联的富媒体数据;向所述生成的文章中,插入提取的富媒体数据。
- 根据权利要求6所述的方法,其特征在于,所述从预先建立的资源库,提取与所述生成的文章的特征相关联的富媒体数据包括:根据以下一项或多项从预先建立的资源库中提取富媒体数据生成候选富媒体列表:所述文章主题、所述文章提纲、所述生成的文章的各段落的摘要以及所述生成的文章的各段落的关键词;采用质量筛选从所述候选富媒体列表中提取与所述生成的文章的特征相关联的富媒体数据。
- 根据权利要求6-7任意一项所述的方法,其特征在于,所述预先建立的资源库通过以下步骤建立:获取富媒体数据的特征;根据所述富媒体数据的特征建立索引结构,得到所述资源库。
- 根据权利要求7所述的方法,其特征在于,所述质量筛选根据以下一项或多项进行:图文相关性、图片分辨率、图片长宽比、图片来源权威度、广告过滤策略、反作弊过滤策略、反黄过滤策略和水印过滤策略。
- 根据权利要求1-9任意一项所述的方法,其特征在于,所述 方法还包括:将所述文章主题和所述文章提纲输入标题模型,得到所述生成的文章的标题。
- 根据权利要求10所述的方法,其特征在于,所述方法还包括:对所述标题中的核心词进行属性扩展;对属性扩展后的标题中的核心词进行替换和改写,得到更新后的标题。
- 一种用于生成文章的装置,其特征在于,所述装置包括:提纲生成单元,用于基于输入的文章主题和以下任意一项生成文章提纲:提纲模型,根据对应所述文章主题的用户行为数据建立的提纲数据库,以及人工设定的提纲;素材提取单元,用于从预先建立的素材库中,提取与所述文章提纲的特征相关联的素材;素材插入单元,用于向所述文章提纲中,插入提取的素材,得到生成的文章。
- 根据权利要求12所述的装置,其特征在于,所述提纲生成单元中的所述根据对应所述文章主题的用户行为数据建立的提纲数据库包括:检索全网围绕所述文章主题的子主题,建立子主题数据库;根据用户对所述子主题数据库中的子主题的点击顺序和/或所述子主题数据库中的子主题的语义递进顺序,排序所述子主题数据库中的子主题;剔除所述子主题数据库中不符合预定逻辑规则的子主题,得到符合预定逻辑规则的子主题;将各符合预定逻辑规则的子主题作为提纲,得到提纲数据库。
- 根据权利要求12所述的装置,其特征在于,所述素材提取单 元中的所述预先建立的素材库通过以下步骤建立:获取素材的特征,所述素材为将现有的文章的内容根据筛选规则筛选得到和/或变换现有的文章的内容得到;根据所述素材的特征建立索引结构,得到所述素材库。
- 根据权利要求12所述的装置,其特征在于,所述装置还包括:文章优化单元,用于对所述生成的文章进行优化处理,得到优化后的所述生成的文章,所述优化处理包括以下一项或多项:润色处理、插入富媒体数据处理以及排版优化处理。
- 根据权利要求15所述的装置,其特征在于,所述文章优化单元中的所述润色处理包括以下一项或多项:统一所述生成的文章的文法风格;删除与前后语句不连贯的语句;以及替换与前后语句不连贯的语句。
- 根据权利要求15所述的装置,其特征在于,所述文章优化单元中的所述插入富媒体数据处理包括:从预先建立的资源库,提取与所述生成的文章的特征相关联的富媒体数据;向所述生成的文章中,插入提取的富媒体数据。
- 根据权利要求17所述的装置,其特征在于,所述文章优化单元中的所述从预先建立的资源库,提取与所述生成的文章的特征相关联的富媒体数据包括:根据以下一项或多项从预先建立的资源库中提取富媒体数据生成候选富媒体列表:所述文章主题、所述文章提纲、所述生成的文章的各段落的摘要以及所述生成的文章的各段落的关键词;采用质量筛选从所述候选富媒体列表中提取与所述生成的文章的特征相关联的富媒体数据。
- 根据权利要求17-18任意一项所述的装置,其特征在于,所述文章优化单元中的所述预先建立的资源库通过以下步骤建立:获取富媒体数据的特征;根据所述富媒体数据的特征建立索引结构,得到所述资源库。
- 根据权利要求18所述的装置,其特征在于,所述文章优化单元中的所述质量筛选根据以下一项或多项进行:图文相关性、图片分辨率、图片长宽比、图片来源权威度、广告过滤策略、反作弊过滤策略、反黄过滤策略和水印过滤策略。
- 根据权利要求12-20任意一项所述的装置,其特征在于,所述装置还包括:标题生成单元,用于将所述文章主题和所述文章提纲输入标题模型,得到所述生成的文章的标题。
- 根据权利要求21所述的装置,其特征在于,所述装置还包括:属性扩展单元,用于对所述标题中的核心词进行属性扩展;标题更新单元,用于对属性扩展后的标题中的核心词进行替换和改写,得到更新后的标题。
- 一种设备,其特征在于,包括:一个或多个处理器;存储装置,用于存储一个或多个程序;当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现如权利要求1-11中任一所述的用于生成文章的方法。
- 一种计算机可读存储介质,其上存储有计算机程序,其特征在于,该程序被处理器执行时实现如权利要求1-11中任一所述的用于生成文章的方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US16/355,263 US20190213216A1 (en) | 2017-03-31 | 2019-03-15 | Method and device for generating article |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201710206961.1 | 2017-03-31 | ||
| CN201710206961.1A CN106970898A (zh) | 2017-03-31 | 2017-03-31 | 用于生成文章的方法和装置 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US16/355,263 Continuation US20190213216A1 (en) | 2017-03-31 | 2019-03-15 | Method and device for generating article |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018176758A1 true WO2018176758A1 (zh) | 2018-10-04 |
Family
ID=59335645
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2017/102620 Ceased WO2018176758A1 (zh) | 2017-03-31 | 2017-09-21 | 用于生成文章的方法和装置 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20190213216A1 (zh) |
| CN (1) | CN106970898A (zh) |
| WO (1) | WO2018176758A1 (zh) |
Families Citing this family (30)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106970898A (zh) * | 2017-03-31 | 2017-07-21 | 百度在线网络技术(北京)有限公司 | 用于生成文章的方法和装置 |
| CN108052494B (zh) * | 2017-12-28 | 2021-09-14 | 掌阅科技股份有限公司 | 漫画书生成方法、电子设备及计算机存储介质 |
| CN108694160B (zh) * | 2018-05-15 | 2021-01-22 | 北京三快在线科技有限公司 | 文章生成方法、设备及存储介质 |
| CN110555199B (zh) * | 2018-06-01 | 2023-07-04 | 北京百度网讯科技有限公司 | 基于热点素材的文章生成方法、装置、设备及存储介质 |
| CN108829854B (zh) * | 2018-06-21 | 2021-08-31 | 北京百度网讯科技有限公司 | 用于生成文章的方法、装置、设备和计算机可读存储介质 |
| CN109165379A (zh) * | 2018-07-03 | 2019-01-08 | 湖北今古传奇数字新媒体有限公司 | 一种网络杂志的版面制作方法 |
| CN108959234A (zh) * | 2018-08-08 | 2018-12-07 | 山东理工职业学院 | 一种基于群体决策的智能小说生成系统 |
| CN109446505A (zh) * | 2018-10-31 | 2019-03-08 | 广东小天才科技有限公司 | 一种范文生成方法及系统 |
| CN109948409A (zh) * | 2018-11-30 | 2019-06-28 | 北京百度网讯科技有限公司 | 用于生成文章的方法、装置、设备和计算机可读存储介质 |
| CN109657043B (zh) * | 2018-12-14 | 2022-01-04 | 北京百度网讯科技有限公司 | 自动生成文章的方法、装置、设备及存储介质 |
| CN109902305A (zh) * | 2019-03-04 | 2019-06-18 | 上海宝尊电子商务有限公司 | 基于命名实体识别的模板生成、搜索及文本生成设备与方法 |
| CN109885821B (zh) * | 2019-03-05 | 2023-07-18 | 中国联合网络通信集团有限公司 | 基于人工智能的文章撰写方法及装置、计算机存储介质 |
| CN109918516B (zh) * | 2019-03-13 | 2021-07-30 | 百度在线网络技术(北京)有限公司 | 一种数据处理方法、装置及终端 |
| CN110516227A (zh) * | 2019-03-28 | 2019-11-29 | 苏州八叉树智能科技有限公司 | 标题文本生成方法、装置、电子设备及计算机可读介质 |
| CN110059307B (zh) * | 2019-04-15 | 2021-05-14 | 百度在线网络技术(北京)有限公司 | 写作方法、装置和服务器 |
| CN110245339B (zh) * | 2019-06-20 | 2023-04-18 | 北京百度网讯科技有限公司 | 文章生成方法、装置、设备和存储介质 |
| US11194816B2 (en) * | 2019-10-16 | 2021-12-07 | International Business Machines Corporation | Structured article generation |
| CN111428472A (zh) * | 2020-03-13 | 2020-07-17 | 浙江华坤道威数据科技有限公司 | 一种基于自然语言处理及图像算法的文章自动生成系统和方法 |
| CN111859118A (zh) * | 2020-06-19 | 2020-10-30 | 京华信息科技股份有限公司 | 一种基于文档目录的智能信息推荐方法及装置 |
| CN112148857B (zh) * | 2020-09-23 | 2024-06-21 | 中国电子科技集团公司第十五研究所 | 一种公文自动生成系统和方法 |
| US12423507B2 (en) | 2021-07-12 | 2025-09-23 | International Business Machines Corporation | Elucidated natural language artifact recombination with contextual awareness |
| US11475211B1 (en) * | 2021-07-12 | 2022-10-18 | International Business Machines Corporation | Elucidated natural language artifact recombination with contextual awareness |
| CN113688633B (zh) * | 2021-08-02 | 2025-06-27 | 珠海金山办公软件有限公司 | 一种提纲确定方法及装置 |
| CN114970467B (zh) * | 2022-05-30 | 2023-09-01 | 平安科技(深圳)有限公司 | 基于人工智能的作文初稿生成方法、装置、设备及介质 |
| CN115170196A (zh) * | 2022-07-15 | 2022-10-11 | 珍岛信息技术(上海)股份有限公司 | 一种基于大数据智能写作的推广方法 |
| US11868313B1 (en) | 2023-03-28 | 2024-01-09 | Lede AI | Apparatus and method for generating an article |
| US12293149B2 (en) | 2023-03-28 | 2025-05-06 | Lede AI | Apparatus and method for generating an article |
| CN116628183A (zh) * | 2023-04-25 | 2023-08-22 | 科大讯飞股份有限公司 | 一种目标提纲生成方法、系统及相关装置 |
| TWI879461B (zh) * | 2024-02-29 | 2025-04-01 | 台灣大哥大股份有限公司 | 搜尋引擎優化與回饋關鍵字的系統及方法 |
| CN120087341B (zh) * | 2025-04-25 | 2025-08-15 | 江西师范大学 | 一种基于结构化元知识库的大模型写作方法及系统 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101441621A (zh) * | 2008-11-26 | 2009-05-27 | 北大方正集团有限公司 | 一种版式文件自动成文的方法及系统 |
| CN102566945A (zh) * | 2010-12-24 | 2012-07-11 | 北大方正集团有限公司 | 一种实现图书自动组稿按需印刷的方法和系统 |
| CN104123269A (zh) * | 2014-07-16 | 2014-10-29 | 华中科技大学 | 一种基于模板的出版物半自动生成方法及系统 |
| CN106407168A (zh) * | 2016-09-06 | 2017-02-15 | 首都师范大学 | 一种应用文自动生成方法 |
| CN106503255A (zh) * | 2016-11-15 | 2017-03-15 | 科大讯飞股份有限公司 | 基于描述文本自动生成文章的方法及系统 |
| CN106970898A (zh) * | 2017-03-31 | 2017-07-21 | 百度在线网络技术(北京)有限公司 | 用于生成文章的方法和装置 |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102207948B (zh) * | 2010-07-13 | 2013-07-24 | 天津海量信息技术有限公司 | 一种事件陈述句素材库的生成方法 |
| CN104933020A (zh) * | 2015-07-17 | 2015-09-23 | 北京奇虎科技有限公司 | 基于模板生成目标文档的方法及装置 |
-
2017
- 2017-03-31 CN CN201710206961.1A patent/CN106970898A/zh active Pending
- 2017-09-21 WO PCT/CN2017/102620 patent/WO2018176758A1/zh not_active Ceased
-
2019
- 2019-03-15 US US16/355,263 patent/US20190213216A1/en not_active Abandoned
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101441621A (zh) * | 2008-11-26 | 2009-05-27 | 北大方正集团有限公司 | 一种版式文件自动成文的方法及系统 |
| CN102566945A (zh) * | 2010-12-24 | 2012-07-11 | 北大方正集团有限公司 | 一种实现图书自动组稿按需印刷的方法和系统 |
| CN104123269A (zh) * | 2014-07-16 | 2014-10-29 | 华中科技大学 | 一种基于模板的出版物半自动生成方法及系统 |
| CN106407168A (zh) * | 2016-09-06 | 2017-02-15 | 首都师范大学 | 一种应用文自动生成方法 |
| CN106503255A (zh) * | 2016-11-15 | 2017-03-15 | 科大讯飞股份有限公司 | 基于描述文本自动生成文章的方法及系统 |
| CN106970898A (zh) * | 2017-03-31 | 2017-07-21 | 百度在线网络技术(北京)有限公司 | 用于生成文章的方法和装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN106970898A (zh) | 2017-07-21 |
| US20190213216A1 (en) | 2019-07-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20190213216A1 (en) | Method and device for generating article | |
| Zhao et al. | Iconate: Automatic compound icon generation and ideation | |
| CN115082602B (zh) | 生成数字人的方法、模型的训练方法、装置、设备和介质 | |
| US10664660B2 (en) | Method and device for extracting entity relation based on deep learning, and server | |
| CN110717017A (zh) | 一种处理语料的方法 | |
| US20200167529A1 (en) | Translation Review Workflow Systems and Methods | |
| CN109271594B (zh) | 电子书的推荐方法、电子设备及计算机存储介质 | |
| JP2010532897A (ja) | 知的なテキスト注釈の方法、システム及びコンピュータ・プログラム | |
| CN117436414A (zh) | 演示文稿生成方法、装置、电子设备和存储介质 | |
| KR102685135B1 (ko) | 영상 편집 자동화 시스템 | |
| US9298689B2 (en) | Multiple template based search function | |
| CN107885785A (zh) | 文本情感分析方法和装置 | |
| WO2025087394A1 (zh) | 素材展示方法、装置、设备及存储介质 | |
| CN119250067B (zh) | 一种文本评审方法、系统、终端及介质 | |
| CN110427519B (zh) | 视频的处理方法及装置 | |
| Bai et al. | The application of knowledge graphs in the Chinese cultural field: the ancient capital culture of Beijing | |
| Lindley | jQuery Cookbook: Solutions & Examples for jQuery Developers | |
| Collins | Pro HTML5 with CSS, JavaScript, and Multimedia | |
| CN117909564A (zh) | 使用生成式ai的内容速度和超个性化 | |
| CN110297965B (zh) | 课件页面的显示及页面集的构造方法、装置、设备和介质 | |
| CN109657043B (zh) | 自动生成文章的方法、装置、设备及存储介质 | |
| CN108664535B (zh) | 信息输出方法和装置 | |
| US10878005B2 (en) | Context aware document advising | |
| CN113535125A (zh) | 金融需求项生成方法及装置 | |
| US8875009B1 (en) | Analyzing links for NCX navigation |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17903516 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205N DATED 22/11/2019) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17903516 Country of ref document: EP Kind code of ref document: A1 |