WO2025258027A1 - 台本生成システム、台本生成方法、及びプログラム - Google Patents
台本生成システム、台本生成方法、及びプログラムInfo
- Publication number
- WO2025258027A1 WO2025258027A1 PCT/JP2024/021527 JP2024021527W WO2025258027A1 WO 2025258027 A1 WO2025258027 A1 WO 2025258027A1 JP 2024021527 W JP2024021527 W JP 2024021527W WO 2025258027 A1 WO2025258027 A1 WO 2025258027A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- script
- input
- language model
- information
- session
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/02—Marketing; Price estimation or determination; Fundraising
- G06Q30/0241—Advertisements
Definitions
- This disclosure relates to a script generation system, a script generation method, and a program.
- Patent Document 1 describes a script production device that receives an engine production engine containing production information for content with embedded advertising data, uses the content production engine to generate a script with embedded identification information that identifies the scriptwriter who created the script, and transmits the script to a viewing device used by a viewer.
- Patent Document 1 also describes templates related to scripts.
- One of the goals of this disclosure is to generate more flexible scripts.
- the script generation system disclosed herein includes an input information acquisition unit that acquires input information related to a video introducing a product or service, the input information being input to a trained language model capable of generating a product written in natural language, and a script generation unit that generates a script related to the video by inputting the input information into the language model.
- This disclosure allows for more flexible script generation.
- FIG. 1 is a diagram illustrating an example of a hardware configuration of a script generation system.
- FIG. 10 is a diagram illustrating an example of how live streaming is performed.
- FIG. 2 is a diagram illustrating an example of functions realized by the script generation system.
- FIG. 10 is a diagram illustrating an example of a script database.
- FIG. 2 is a diagram illustrating an example of input and output to a language model.
- FIG. 10 is a diagram showing an example of input information based on extracted information extracted from content.
- FIG. 2 is a diagram illustrating an example of processing executed by the script generation system.
- FIG. 10 is a diagram illustrating an example of a function realized in a modified example.
- a live streaming service is a service that distributes videos in real time to an unspecified number of people. Performers in the live stream proceed with the live stream according to a script prepared in advance. Viewers watch the videos that are distributed in real time or videos that are saved as archives.
- Figure 1 is a diagram showing an example of the hardware configuration of a script generation system.
- script generation system 1 includes a script generation device 10, a server 20, a performer device 30, and a viewer device 40.
- Each of the script generation device 10, the server 20, the performer device 30, and the viewer device 40 is connected to a network N such as the Internet or a LAN. While Figure 1 shows one each of the script generation device 10, the server 20, the performer device 30, and the viewer device 40, there may be multiple devices of at least one of these.
- the script generation device 10 is a device that generates a script.
- the script generation device 10 is a personal computer, a server computer, a tablet, or a smartphone.
- the script generation device 10 includes a control unit 11, a memory unit 12, a communication unit 13, an operation unit 14, and a display unit 15.
- the control unit 11 includes at least one processor.
- the memory unit 12 includes at least one of volatile memory such as RAM and non-volatile memory such as flash memory.
- the communication unit 13 includes at least one of a communication interface for wired communication and a communication interface for wireless communication.
- the operation unit 14 is an input device such as a touch panel.
- the display unit 15 is an LCD or organic EL display.
- the server 20 is a server computer for a live streaming service.
- the server 20 includes a control unit 21, a memory unit 22, and a communication unit 23.
- the hardware configurations of the control unit 21, the memory unit 22, and the communication unit 23 may be similar to those of the control unit 11, the memory unit 12, and the communication unit 13, respectively.
- the performer device 30 is a device of a performer.
- the performer device 30 is a personal computer, tablet, or smartphone.
- the performer device 30 includes a control unit 31, a memory unit 32, a communication unit 33, an operation unit 34, and a display unit 35.
- the hardware configurations of the control unit 31, the memory unit 32, the communication unit 33, the operation unit 34, and the display unit 35 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 14, and the display unit 15, respectively.
- a camera unit 36 is connected to the performer device 30.
- the camera unit 36 includes at least one camera.
- the camera unit 36 may be included inside the performer device 30.
- the viewer device 40 is a viewer's device.
- the viewer device 40 is a personal computer, a tablet, or a smartphone.
- the viewer device 40 includes a control unit 41, a memory unit 42, a communication unit 43, an operation unit 44, and a display unit 45.
- the hardware configurations of the control unit 41, the memory unit 42, the communication unit 43, the operation unit 44, and the display unit 45 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 14, and the display unit 15, respectively.
- the programs stored in the storage units 12, 22, 32, 42 may be supplied to the script generation device 10, the server 20, the performer device 30, or the viewer device 40 via the network N.
- the script generation device 10, the server 20, the performer device 30, or the viewer device 40 may include at least one of a reading unit (e.g., a memory card slot) that reads a computer-readable information storage medium and an input/output unit (e.g., a USB port) for inputting and outputting data to and from an external device.
- a program stored on an information storage medium may be supplied to the script generation device 10, the server 20, the performer device 30, or the viewer device 40 via at least one of the reading unit and the input/output unit.
- the script generation system 1 only needs to include at least one computer.
- the computers included in the script generation system 1 are not limited to the example shown in Figure 1.
- the script generation system 1 may include only the script generation device 10 and the server 20.
- the performer device 30 and the viewer device 40 exist outside the script generation system 1.
- the script generation system 1 may also include only the script generation device 10.
- the server 20, the performer device 30, and the viewer device 40 exist outside the script generation system 1.
- the script generation system 1 may include the script generation device 10 and other computers not shown in Figure 1.
- the live streaming service may be one of the services provided by the operator of the e-commerce service, or may be a separate service from the e-commerce service.
- a store affiliated with the e-commerce service may perform a live streaming to introduce the product or service it sells.
- the live streaming may be performed by any party.
- the live streaming may be performed by a party other than the store, such as a product manufacturer, a service provider, or an influencer.
- Figure 2 shows an example of how live streaming takes place.
- a performer introduces a product or service in front of the filming unit 36 according to a prepared script.
- the performer may be a store associate, or another person requested by the store to appear.
- the performer device 30 transmits data (video data) showing the results of filming by the filming unit 36 to the server 20.
- the server 20 distributes video in real time to the viewer device 40 based on the data received from the performer device 30.
- the viewer device 40 displays an introduction screen SC on the display unit 45, showing a video of the performer introducing the product or service.
- the operator of the live streaming service provides a script generation service that generates a script for a store that has applied for live streaming.
- the store staff may prepare the script themselves, or may use the script generation service to prepare a script. For example, if a store staff member uses the script generation service, the store staff member fills out the necessary information on an orientation sheet, described below, and requests the operator to use the script generation service. Based on the request from the store staff member, the operator generates a script that corresponds to the product or service to be introduced in the live streaming.
- the script generation system 1 of this embodiment uses a trained language model to generate a more flexible script. Details of the script generation system 1 will be explained below.
- FIG. 3 shows an example of functions realized by the script generation device 10.
- the script generation device 10 includes a data storage unit 100, an input information acquisition unit 101, and a script generation unit 102.
- the data storage unit 100 is realized by the storage unit 12.
- the input information acquisition unit 101 and the script generation unit 102 are each realized by the control unit 11.
- the data storage unit 100 stores data necessary for generating a script.
- the data storage unit 100 stores a trained language model M capable of generating a product written in a natural language, and a script database DB in which a script generated by the language model M is stored.
- the data stored in the data storage unit 100 is not limited to these examples.
- the data storage unit 100 can store any data.
- a natural language is a language that can be understood by humans.
- a natural language may be any language, such as Japanese, English, or Chinese.
- a product is data generated by a language model M.
- a language model M can generate any product.
- a product may be text, a table, a diagram, an image, a video, or a combination thereof.
- Text includes letters, numbers, symbols, or a combination thereof.
- Text may be in any format.
- text may be a sentence, a bulleted list, a list of words, program code, or code written in a markup language.
- Language model M is a model used in the field of natural language processing.
- language model M is a model that uses machine learning techniques.
- Language model M is sometimes called a large-scale language model or generative AI (Artificial Intelligence).
- Language model M may be a model called by other names.
- language model M includes a program that executes information processing to process input information and generate a product, and parameters referenced by the program. The parameters are adjusted through learning.
- the parameters of language model M may be publicly known parameters.
- the parameters of language model M may be weights or biases. In this embodiment, it is assumed that language model M has been previously trained using a large dataset.
- the language model M may be any of various types of known models.
- the language model M may be a Generative Pre-trained Transformer (GPT), a Transformer-based model other than the GPT (for example, BERT: Bidirectional Encoder Representations from Transformers, or T5: Text-To-Text Transfer Transformer), a neural network capable of natural language processing, or other models (for example, Pegasus, UniLM, or Electra).
- GPT Generative Pre-trained Transformer
- BERT Bidirectional Encoder Representations from Transformers
- T5 Text-To-Text Transfer Transformer
- the programs and parameters included in the language model M may be similar to those of these known models.
- the language model M may be exactly the same as a known model, or may be a model fine-tuned using training data specialized for script generation.
- the data storage unit 100 stores the language model M
- the language model M may also be stored in a device other than the script generation device 10.
- the other device may be a device managed by a company that provides the language model M functions online.
- the script generation device 10 sends input information to be input into the language model M to the other device.
- the other device inputs the input information received from the script generation device 10 into the language model M stored therein.
- the other device sends the product output by the language model M to the script generation device 10.
- the script generation device 10 receives the product from the other device.
- Figure 4 is a diagram showing an example of a script database DB.
- the script database DB stores a live streaming ID, an orientation sheet, and a script.
- the script database DB may store any information related to the script.
- the script database DB may store questions and answers as described in the modified example below, video footage of performers introducing products or services in accordance with the script (e.g., video for archive distribution), or input information used when generating the script.
- a live streaming ID is an ID that can identify an individual live stream. For example, when a store staff member applies for a live streaming service, a new live streaming ID is issued.
- An orientation sheet is data that shows basic information about the video to be streamed on the live streaming service.
- the orientation sheet includes information about the product or service to be introduced.
- the information about the product or service may be any information, such as the product name, service name, characteristics of the product or service itself, price, inventory, color variations, size variations, or other information.
- the orientation sheet may also include input characteristic information, input setting information, and input session information, which will be described below.
- the orientation sheet may be in any data format.
- the orientation sheet may be in a spreadsheet data format, CSV format, text format, document format, or other format.
- the orientation sheet may also include information other than the product or service.
- the orientation sheet may include information about the store requesting the live stream (e.g., the store's ID or name), the date and time of the live stream, the length (length) of the live stream, information about the performers (e.g., the performers' names, stage names, or profiles), whether a script has been requested, the level of detail of the script (described below), prohibited words, or other information.
- a store staff member wishing to live stream operates their own terminal to enter the necessary information into an orientation sheet. They may be required to enter all items on the orientation sheet, or only some of the items.
- the store staff member's terminal sends the orientation sheet to the server 20.
- the server 20 receives the orientation sheet from the store staff member's terminal.
- the server 20 issues a live streaming ID and records the live streaming ID and orientation sheet in the memory unit 22.
- the orientation sheet may be generated by any person. For example, the operator of the live streaming service may generate the orientation sheet based on the store's requests.
- the operator of the live streaming service will review the contents of the orientation sheet to determine whether to allow a store to perform live streaming. If the store passes the review and the orientation sheet indicates that the store will use the script generation service, the script generation device 10 will obtain the live streaming ID and orientation sheet from the server 20 and store them in the script database DB.
- the script is stored in the script database DB.
- the term "script" refers to data that indicates the script.
- the script may be in any data format.
- the script may be in text format, document format, spreadsheet data format, CSV format, or another format.
- the data format of the script may be specified by a default prompt, which will be described later.
- the script database DB may also store input information used when generating the script.
- the script may be divided into separate data for each session, which will be described later.
- the input information acquisition unit 101 acquires input information to be input to the language model M, which is input information related to a video introducing a product or service.
- the input information indicates text (e.g., a sentence) written in a natural language.
- the input information is also called a prompt.
- the input information may be in a format that can be processed by the language model M and includes at least one character.
- the input information may include information other than the text written in a natural language (e.g., an image or a video). Other information other than the input information (e.g., information that serves as a sample of a product) may be input to the language model M together with the input information.
- FIG. 5 is a diagram showing an example of input and output to language model M.
- a script for a video introducing a product or service is generated as a product of language model M, so the input information includes at least one piece of information related to the video for which the script is to be generated.
- the input information includes four pieces of information: a default prompt, input feature information, input setting information, and input session information.
- the input information may include any number of pieces of information.
- the input information may include one, two, three, five or more pieces of information.
- a default prompt is a prompt that is prepared in advance.
- a default prompt is a type of input information.
- a default prompt can contain any information.
- a default prompt includes text written in a natural language.
- a default prompt may include sentences, program code, code written in a markup language such as JSON, or other text.
- a default prompt may include information other than text (for example, an image or video).
- Default prompts are stored in the data storage unit 100. The person in charge of generating the script may edit the content of the default prompt.
- the language model M is not a model specialized for a specific purpose, but a general-purpose model capable of generating a variety of products.
- the input information includes a default prompt indicating the processing content to be performed by the language model M, so that the general-purpose language model M can recognize the processing content to be performed by itself.
- the processing content to be performed by the language model M can also be referred to as the role (task) of the language model M, or the type of product to be generated by the language model M.
- a generic language model M recognizes the processing content it should perform based on the default prompt included in the input information. That is, a generic language model M recognizes its own role or the type of product it should generate based on the default prompt included in the input information.
- the default prompt indicates that the language model M should generate a script based on the input information, such as "You are the scriptwriter for a live broadcast in which a product or service will be introduced. Please generate a script for the live broadcast based on the input information entered by you.”
- the default prompt is not limited to the example in Figure 5.
- the default prompt may indicate that language model M should generate a script using wording other than that in Figure 5.
- the default prompt may include information other than information indicating that language model M should generate a script.
- the default prompt may include the language of the script, the length of the script (e.g., the number of characters or pages), the layout of the script (e.g., a format in which lines are written after the names of the actors), the file format of the script, or other information.
- the default prompt may also include settings for language model M itself.
- the language model M may be a model specialized for generating a script.
- the language model M can generate a script even if the default prompt does not explicitly indicate that the language model M should generate a script, and therefore the input information does not need to include a default prompt.
- training data specialized for generating a script is assumed to have been learned by the language model M. For example, training data including pairs of training input information and a correct script is learned by the language model M. A large amount of training data may also be learned by the language model M.
- a language model M specialized for generating a script can generate a script based on input information input to it, based on parameters adjusted in advance, even without a default prompt.
- the input information includes input feature information, input setting information, and input session information in addition to the default prompt.
- the input information may include only some of the input feature information, input setting information, and input session information.
- the input information may include only input feature information, only input setting information, only input session information, only input feature information and input setting information, only input feature information and input session information, or input setting information and input session information.
- the input information may also include other information not shown in FIG. 5.
- the input information acquisition unit 101 acquires, from the script database DB, an orientation sheet associated with the live streaming ID of the live streaming for which a script is to be generated.
- the live streaming for which a script is to be generated is specified from the operation unit 14 of the script generation device 10, but may be identified by other methods.
- the input information acquisition unit 101 may refer to the script database DB and identify a live streaming for which a script has not yet been generated as the live streaming for which a script is to be generated.
- the input information acquisition unit 101 may also identify a live streaming for which the length of the period until the streaming date and time is less than a threshold as the live streaming for which a script is to be generated.
- the input information acquisition unit 101 acquires input feature information, which is input information related to the features of a product or service.
- Input feature information is a type of input information.
- the features of a product or service can also be referred to as a description of the product or service.
- the input feature information is written in text in a natural language.
- the input feature information indicates the features of the product or service using letters, numbers, symbols, or a combination of these.
- the input feature information may indicate the product or service's identification information (e.g., name, model number, manufacturer, or JAN code), classification (e.g., genre, category, attribute, or attribute value), appearance (e.g., design, size, color, or pattern), quality, function, material, price, discount rate, appealing points, store information, or other information.
- the input feature information may also be the title, description, search information, or other information of the product or service listed on the e-commerce service.
- the input information acquisition unit 101 acquires input feature information included in an orientation sheet.
- the input information acquisition unit 101 may generate input feature information based on information included in the orientation sheet, rather than acquiring the input feature information included in the orientation sheet.
- the input information acquisition unit 101 may generate input feature information based on information not included in the orientation sheet.
- the input information acquisition unit 101 may generate input feature information when the amount of input feature information included in the orientation sheet is insufficient, or when the orientation sheet does not include input feature information.
- the input information acquisition unit 101 may acquire input feature information generated by inputting extracted information extracted from content related to the features of a product or service into the language model M or another language model.
- the content is electronic information.
- the content may be a website, a web advertisement, a digital catalog, a digital pamphlet, a digital flyer, an image showing a scanned piece of paper, or other information.
- a website will be described as an example of content.
- FIG. 6 is a diagram showing an example of input information based on extracted information extracted from content.
- the input information acquisition unit 101 uses a known web search service to perform a search for the product or service to be introduced.
- the search query may be entered by the person operating the script generation device 10, or may be acquired based on input feature information included in an orientation sheet.
- the input information acquisition unit 101 acquires, as content, an image I showing the product or service from the website found in the search.
- the input information acquisition unit 101 performs optical character recognition on the content and extracts text from the content as extracted information.
- the extracted information may be information other than text (for example, a table or a diagram).
- the input information acquisition unit 101 may acquire the extracted information extracted from the content as input feature information as is. If the content includes text, the input information acquisition unit 101 may extract the extracted information by extracting text from the content without particularly performing optical character recognition. Extraction of the extracted information may be performed by a function other than the input information acquisition unit 101.
- the input information acquisition unit 101 may acquire first input information, which is input information prepared in advance, and second input information, which is input information generated by inputting the first input information into language model M or another language model.
- the first input information may be information included in an orientation sheet.
- input feature information included in an orientation sheet corresponds to the first input information.
- the other language model is a language model different from language model M that generates the script. Like language model M, the other language model may be any model. In the description of the other language model, the reference to "language model M" in the description of language model M can be replaced with "other language model.”
- the input information acquisition unit 101 inputs the extracted information into language model M or another language model, and acquires input feature information output from language model M or another language model.
- the input information acquisition unit 101 may input input feature information included in an orientation sheet together with the extracted information to language model M or another language model.
- the input feature information included in the orientation sheet corresponds to the first input information.
- the input information acquisition unit 101 may input a default prompt to language model M or another language model indicating that it should generate input feature information.
- the default prompt may indicate the processing content to be performed by language model M or another language model (the role of language model M or another language model, or the type of product to be generated by language model M or another language model), such as "You are AI that generates input feature information indicating the features of a product or service based on product or service information extracted from content.” Note that if language model M or another language model is a model specialized for generating input feature information from extracted information, such a default prompt may not need to be input.
- language model M or another language model divides information such as extracted information input to it into tokens and calculates embeddings for each token based on parameters adjusted during training.
- the token division method may be a known method.
- the embeddings indicate the features of the tokens.
- the embeddings may be multidimensional vectors or may be in other formats.
- Language model M or another language model predicts the next text as needed based on the order of the token embeddings, and then generates the input feature information as a product.
- language model M or another language model extracts information from the extracted information that does not overlap with the input feature information included in the orientation sheet, and outputs it as input feature information.
- the input information acquisition unit 101 acquires the input feature information output from language model M or another language model.
- the input feature information output from language model M or another language model corresponds to the second input information.
- characters extracted by optical character recognition from image I of a website which is an example of content, are formatted by language model M.
- Language model M outputs input feature information that indicates features not included in the orientation sheet.
- the input information acquisition unit 101 acquires input setting information, which is input information related to the settings of a video.
- Input setting information is a type of input information.
- Video settings can also be said to be characteristics of the video itself, rather than characteristics of the product or service itself.
- Video settings can also be said to be saved information associated with a video.
- Input setting information is written in text in a natural language.
- input setting information indicates the video settings using letters, numbers, symbols, or a combination of these.
- input setting information is the distribution date and time of the video, its length (length), target demographic (e.g., age group or gender), performer information (e.g., name, profile, or role), prohibited expressions, video title, video summary, or other information.
- the input information acquisition unit 101 acquires input setting information included in an orientation sheet.
- the input information acquisition unit 101 may generate input setting information based on information included in the orientation sheet, rather than acquiring the input setting information included in the orientation sheet.
- the input information acquisition unit 101 may generate input setting information based on information not included in the orientation sheet.
- the input information acquisition unit 101 may generate input setting information when the amount of input setting information included in the orientation sheet is insufficient, or when the orientation sheet does not include input setting information.
- the input information acquisition unit 101 may input information included in an orientation sheet (e.g., input feature information, input setting information, or input session information included in the orientation sheet) into language model M or another language model, and acquire input setting information output from language model M or another language model.
- the information included in the orientation sheet corresponds to the first input information.
- the information not included in the orientation sheet is input to language model M or another language model, the information not included in the orientation sheet corresponds to the first input information.
- the input information acquisition unit 101 may input a default prompt to language model M or another language model, indicating that input setting information should be generated.
- the default prompt may indicate the processing content to be performed by language model M or another language model (the role of language model M or another language model, or the type of product to be generated by language model M or another language model), such as "You are AI that generates input setting information indicating video settings based on information input to you.”
- language model M or another language model divides the information input to it (information used to generate the input setting information) into tokens and calculates the embedded representation of each token based on parameters adjusted through learning.
- Language model M or another language model predicts the next text as needed based on the order of the embedded representations of the tokens, and generates the input setting information as a product.
- language model M or another language model generates a video title, summary, etc. based on the information input to it based on a default prompt, and outputs this as input setting information.
- the input information acquisition unit 101 acquires the input setting information output from language model M or another language model. In this case, the input setting information output from language model M or another language model corresponds to the second input information.
- the input information acquisition unit 101 acquires input session information, which is input information related to each of multiple sessions in a video.
- a session is an individual part that makes up a video.
- a session may also be called a section or other name.
- a video may be divided into multiple sessions from any perspective.
- sessions may be divided by topic or by time.
- Input session information is a type of input information.
- Input session information is written in text in a natural language.
- input session information indicates a session using letters, numbers, symbols, or a combination of these.
- Input session information may indicate an outline introduced in each session.
- session information may be a number indicating the order of the session, a heading, text indicating an outline, a time length, or other information.
- the input information acquisition unit 101 acquires input session information included in an orientation sheet. Instead of acquiring input session information included in the orientation sheet, the input information acquisition unit 101 may generate input session information based on information included in the orientation sheet. The input information acquisition unit 101 may generate input session information based on information not included in the orientation sheet. The input information acquisition unit 101 may generate input session information when the amount of input session information included in the orientation sheet is insufficient, or when the orientation sheet does not include input session information.
- the input information acquisition unit 101 inputs information included in an orientation sheet (e.g., input feature information, input setting information, or input session information included in the orientation sheet) into language model M or another language model, and acquires input session information output from language model M or another language model.
- the information included in the orientation sheet corresponds to the first input information.
- the information not included in the orientation sheet is input to language model M or another language model, the information not included in the orientation sheet corresponds to the first input information.
- the input information acquisition unit 101 may input a default prompt to language model M or another language model, indicating that input session information should be generated.
- the default prompt may indicate the processing content to be performed by language model M or another language model (the role of language model M or another language model, or the type of product to be generated by language model M or another language model), such as "You are AI that generates input session information indicating a video session based on information input to you.”
- the number of sessions may also be specified in the default prompt.
- language model M or another language model divides the information input thereto (information used to generate input session information) into tokens and calculates the embedded representation of each token based on parameters adjusted through learning.
- Language model M or another language model predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and generates the input session information as a product.
- language model M or another language model generates a session heading or the like corresponding to the information input thereto based on a default prompt, and outputs this as input session information.
- the input information acquisition unit 101 acquires the input session information output from language model M or another language model. In this case, the input session information output from language model M or another language model corresponds to the second input information.
- the input information acquisition unit 101 acquires input session information according to the level of detail related to the script.
- the level of detail is the degree of detail in the script.
- the level of detail can also be said to be the volume of the script. For example, there may be two levels of detail, such as a simple version and a detailed version, or there may be three or more levels of detail.
- the level of detail can be specified by any person. For example, a store employee may specify the level of detail.
- the input information acquisition unit 101 acquires input session information such that the higher the level of detail, the more finely divided the video is.
- the level of detail is indicated by letters, numbers, symbols, or a combination of these.
- the input information acquisition unit 101 acquires the level of detail included in the orientation sheet. If the orientation sheet already includes input session information corresponding to the level of detail, the input information acquisition unit 101 acquires input session information corresponding to the level of detail included in the orientation sheet.
- the input information acquisition unit 101 may acquire input session information corresponding to the level of detail based on the language model M or another language model.
- the higher the level of detail the greater the number of sessions.
- the higher the level of detail the more sub-sessions are generated by dividing the session into smaller parts.
- Sub-sessions can also be considered lower levels of the session. There may be three or more levels of sessions, not just two. The higher the level of detail, the greater the number of levels of sessions.
- the input information acquisition unit 101 when the level of detail indicates a simplified version, the input information acquisition unit 101 generates input session information using the method described above, and terminates the process of acquiring input session information without dividing the session into smaller parts.
- the input information acquisition unit 101 causes language model M or another language model to generate sub-sessions by further dividing the session based on the input session information generated by the method described above.
- the input information acquisition unit 101 may input the input session information to language model M or another language model, and acquire the input session information output from language model M or another language model.
- the input information acquisition unit 101 may input a default prompt to language model M or another language model, indicating that it should generate input session information according to the level of detail.
- the default prompt may indicate the processing content to be performed by language model M or another language model (the role of language model M or another language model, or the type of product to be generated by language model M or another language model), such as "You are AI that generates input session information indicating sub-sessions into which a session is further divided, based on the input session information input to you.”
- language model M or another language model divides the input session information input thereto into tokens and calculates the embedding representation of each token based on parameters adjusted through learning.
- Language model M or another language model predicts the next text as needed based on the order of the embedded representations of the tokens, and generates input session information that is more detailed than the input session information input thereto as a product.
- language model M or another language model generates sub-session headings or the like according to the input session information input thereto based on a default prompt, and outputs this as input session information.
- the input information acquisition unit 101 acquires the input session information output from language model M or another language model.
- the input information acquisition unit 101 may input the level of detail to language model M or another language model.
- language model M or another language model may acquire input session information indicating sub-session headings, etc., based on the level of detail input thereto.
- the script generation unit 102 generates a script for a video by inputting input information into the language model M.
- the language model M divides the input information into tokens and calculates embedded representations of each token based on parameters adjusted by learning.
- the language model M predicts the subsequent text as needed based on the order of the embedded representations of the tokens, and then generates a script as a product. For example, the language model M generates and outputs a script corresponding to the input information based on a default prompt.
- the input information acquisition unit 101 acquires the script output from the language model M.
- the script generation unit 102 generates a script by inputting input feature information into the language model M.
- the language model M divides the input feature information into tokens and calculates embedded representations for each token based on parameters adjusted through learning.
- the language model M predicts the next text as necessary based on the order of the embedded representations of the tokens, and then generates a script as a product.
- the language model M generates and outputs a script corresponding to the input feature information based on a default prompt.
- the input feature information acquisition unit acquires the script output from the language model M.
- the script generation unit 102 generates a script by inputting input setting information into the language model M.
- the language model M divides the input setting information into tokens and calculates the embedded representation of each token based on parameters adjusted by learning.
- the language model M predicts the subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script as a product.
- the language model M generates and outputs a script corresponding to the input setting information based on a default prompt.
- the input setting information acquisition unit acquires the script output from the language model M.
- the script generation unit 102 generates a script by inputting input session information into the language model M.
- the language model M divides the input session information into tokens and calculates embedded representations for each token based on parameters adjusted through learning.
- the language model M predicts the next text as necessary based on the order of the embedded representations of the tokens, and then generates a script as a product.
- the language model M generates and outputs a script according to the input session information based on a default prompt.
- the input session information acquisition unit acquires the script output from the language model M.
- the script generation unit 102 generates a script according to the level of detail by inputting input session information according to the level of detail into the language model M.
- the language model M divides the input session information according to the level of detail into tokens and calculates embedded representations for each token based on parameters adjusted through learning.
- the language model M predicts the next text as necessary based on the order of the embedded representations of the tokens, and generates a script as a product.
- the language model M generates and outputs a script according to the input session information according to the level of detail based on a default prompt.
- the input session information acquisition unit acquires the script output from the language model M.
- the script generation unit 102 For each session, the script generation unit 102 generates a script portion that is part of the script, by inputting the input session information for that session into the language model M, and for the second or subsequent session among the multiple sessions, the script generation unit 102 generates a script portion for the second or subsequent session by also inputting the script portions of sessions before the second or subsequent session into the language model M, thereby generating a script based on the script portions of each of the multiple sessions.
- the default prompt may indicate that a script portion should be generated for each individual session.
- the script generation unit 102 generates a script by inputting first input information and second input information into the language model M.
- the language model M divides the first input information and second input information into tokens and calculates embedded representations for each token based on parameters adjusted through learning.
- the language model M predicts subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script as a product.
- the language model M generates and outputs a script corresponding to the first input information and second input information based on a default prompt.
- the input session information acquisition unit acquires the script output from the language model M.
- the script generation unit 102 may generate a script including all sessions at once, rather than generating script portions for each session. In this case, a default prompt may indicate that the script should be generated at once. Also, the script does not need to be divided into multiple sessions. In this case, input session information is not acquired.
- the script generation unit 102 may generate a script based only on input feature information.
- the script generation unit 102 may generate a script based only on input setting information.
- the script generation unit 102 may generate a script based only on input session information.
- the script generation unit 102 may generate a script based on input information.
- the script generation unit 102 may generate a script without any particular regard to the level of detail.
- Fig. 7 is a diagram showing an example of processing executed by the script generation system 1.
- Fig. 7 shows processing executed by the script generation device 10 among the processing executed by the script generation system 1.
- the processing of Fig. 7 is executed by the control unit 11 executing a program stored in the storage unit 12.
- Each step of Fig. 7 is an example of a step included in the script generation method according to the present disclosure. For example, when a person in charge of generating a script in a script generation service operates the script generation device 10 to specify a live broadcast for which a script is to be generated, the processing of Fig. 7 is executed.
- the script generation device 10 obtains, from the script database DB, an orientation sheet associated with the live streaming ID of the live streaming for which a script is to be generated (S1).
- the script generation device 10 obtains input feature information based on the orientation sheet (S2).
- the script generation device 10 obtains, as input feature information, the information entered in the item on the orientation sheet that indicates the features of the product or service.
- the script generation device 10 searches for content of products or services to be introduced in the live broadcast based on the input feature information acquired in S2 (S3).
- the script generation device 10 searches publicly known web search services using information such as the product name or service name indicated in the input feature information as a search query.
- the script generation device 10 acquires extracted information from the content searched in S3 (S4).
- the script generation device 10 acquires input feature information by inputting a default prompt indicating that input feature information will be generated and the extracted information extracted in S4 into the language model M (S5).
- the language model M divides these into tokens and outputs input feature information based on the sequence of the embedded expressions of the tokens.
- the script generation device 10 acquires the input feature information output from the language model M.
- steps S3 to S5 do not need to be performed.
- the person in charge of generating the script may specify whether or not steps S3 to S5 need to be performed by operating the operation unit 14. If information has been entered into all or a predetermined number of items on the orientation sheet that indicate the characteristics of the product or service, the script generation device 10 does not need to perform steps S3 to S5.
- the script generation device 10 acquires input setting information based on the orientation sheet (S6).
- the script generation device 10 acquires the input setting information contained in the orientation sheet. If the orientation sheet does not contain input setting information or does not contain a sufficient amount of input setting information, the script generation device 10 inputs a default prompt indicating that input setting information will be generated and the information contained in the orientation sheet into the language model M.
- the language model M divides these into tokens and outputs input setting information based on the sequence of the embedded expressions of the tokens.
- the script generation device 10 acquires the input setting information output from the language model M.
- the script generation device 10 may also acquire information such as a summary of the live broadcast as input setting information.
- the script generation device 10 acquires input session information based on the orientation sheet (S7).
- the script generation device 10 acquires the input session information contained in the orientation sheet. If the orientation sheet does not contain input session information, or if the orientation sheet does not contain a sufficient amount of input session information, the script generation device 10 inputs a default prompt indicating that input session information will be generated and the information contained in the orientation sheet into the language model M.
- the language model M divides these into tokens and outputs input session information based on the sequence of the embedded expressions of the tokens.
- the script generation device 10 acquires the input session information output from the language model M.
- the script generation device 10 determines whether to generate a simplified or detailed script based on the orientation sheet (S8). In S8, the script generation device 10 determines whether the level of detail of the script included in the orientation sheet indicates a simplified or detailed version. If the orientation sheet does not include the level of detail of the script, the person in charge of generating the script may specify the level of detail of the script by operating the operation unit 14. The script generation device 10 may make the determination in S8 based on the level of detail of the script specified by the person in charge.
- the script generation device 10 If it is determined in S8 that a simplified script is to be generated (S8: simplified version), the script generation device 10 generates a script portion for the first session by inputting input information including input feature information, input setting information, and input session information into the language model M (S9). In S9, the script generation device 10 also inputs a default prompt indicating that a script portion corresponding to the session will be generated into the language model M.
- the language model M divides the input information, etc. into tokens and outputs the script portion for the first session based on the sequence of the embedded representations of the tokens. Because there are no other sessions before the first session, the language model M outputs the script portion for the first session without basing it on the script portions of the other sessions.
- the script generation device 10 obtains the script portion for the first session output from the language model M.
- the script generation device 10 generates a script portion for the next session by inputting input information including input feature information, input setting information, and input session information, and a script portion for a session that has already been generated, into the language model M (S10).
- the script generation device 10 also inputs a default prompt indicating that a script portion for the next session will be generated into the language model M.
- the language model M divides the input information and script portion into tokens and outputs the script portion for the next session based on the sequence of the embedded representations of the tokens. For example, the script generation device 10 inputs the script portions of all sessions that have been generated up to that point into the language model M. The script generation device 10 obtains the script portion for the next session output from the language model M.
- the script generation device 10 determines whether or not the script portion for the final session has been generated (S11). If it is determined in S11 that the script portion for the final session has not been generated (S11: N), the process returns to S10, and the script portion for the next session is generated. If it is determined that the script portion for the final session has been generated (S11: Y), the script generation device 10 generates a final script based on the script portions of each session (S12), and this process ends. In S12, the script generation device 10 generates the final script by connecting the script portions of each session. The script generation device 10 stores the final script in the script database DB.
- the script generation device 10 acquires input session information indicating each sub-session of multiple sessions by generating a sub-session for each session indicated by the input session information (S13).
- the script generation device 10 inputs a default prompt indicating that a sub-session will be generated for each session and the input session information to the language model M.
- the language model M divides these into tokens and outputs input session information indicating the sub-sessions based on the sequence of the embedded expressions of the tokens.
- the script generation device 10 acquires the input session information output from the language model M.
- the script generation device 10 generates a script portion of the first sub-session of the session to be processed by inputting input information including input feature information, input setting information, and input session information for the session to be processed into the language model M (S14).
- the session to be processed is the session that is the target of the loop from S14 to S17. Sessions to be processed are selected in order, starting with the first session.
- the script generation device 10 also inputs a default prompt to the language model M indicating that a script portion will be generated according to the sub-session.
- the language model M divides the input information, etc. into tokens and outputs the script portion of the sub-session of the first session based on the sequence of the embedded representations of the tokens.
- the language model M outputs the script portion of the first sub-session without basing it on the script portions of the other sub-sessions.
- the script generation device 10 obtains the script portion of the first sub-session output from the language model M.
- the script generation device 10 generates a script portion for the next sub-session of the session being processed based on input information including input feature information, input setting information, and input session information for the session being processed, and the script portion of the sub-session that has already been created (S15).
- the script generation device 10 also inputs a default prompt to the language model M indicating that a script portion for the next sub-session will be generated.
- the language model M divides the input information and script portion into tokens, and outputs the script portion of the next sub-session based on the sequence of the embedded representations of the tokens. For example, the script generation device 10 inputs the script portions of all sub-sessions that have been generated up to that point into the language model M.
- the script generation device 10 obtains the script portion of the next sub-session output from the language model M.
- the script generation device 10 determines whether the script portion of the last sub-session of the session being processed has been generated (S16). If it is determined that the script portion of the last sub-session of the session being processed has not been generated (S16: N), the process returns to S15, and the script portion of the next sub-session of the session being processed is generated. If it is determined that the script portion of the last sub-session of the session being processed has been generated (S16: Y), the script generation device 10 determines based on the input session information whether the script portion of the last session has been generated (S17).
- the process returns to S14, and the script portion for the first sub-session of the next session is generated. In other words, the next session becomes the session to be processed. If it is determined that the script portion for the final session has been generated (S17: Y), the script generation device 10 generates a final script based on the script portions of each sub-session of each session (S18), and this process ends. In S18, the script generation device 10 generates the final script by piecing together the script portions of each sub-session of each session. The script generation device 10 stores the final script in the script database DB.
- the script generation system 1 of this embodiment acquires input information related to a video in which a product or service is introduced.
- the script generation system 1 generates a script for the video by inputting the input information into a language model M.
- the script generation system 1 can generate a flexible script according to the input information by using the language model M. For example, if a person providing a script generation service generates a script based on a template, the script can only be generated within the scope of the template, and therefore a flexible script cannot be generated. However, the language model M can generate a more flexible script by performing flexible natural language processing according to the input information.
- the script generation system 1 can also reduce the workload of the person providing the script, thereby reducing the costs associated with script generation. For example, if a person providing a script generation service outsources the generation of a script to an external company, costs will be incurred for the external company, but the script generation system 1 can avoid such costs.
- the script generation system 1 generates a script by inputting input feature information into the language model M.
- This enables the language model M to flexibly generate scripts that correspond to the features of the product or service, allowing the script generation system 1 to improve the accuracy of the script, which is the product of the language model M.
- the script generation system 1 can generate a script that corresponds to the appearance features.
- the script generation system 1 can generate a script that corresponds to the functional features.
- the script generation system 1 also acquires input feature information generated by inputting extracted information extracted from content related to the features of a product or service into the language model M or another language model. This allows the script generation system 1 to generate a script from more features, thereby improving the accuracy of the script, which is the product of the language model M. For example, even if there is no input feature information in the orientation sheet, or if there is not enough input feature information in the orientation sheet, the script generation system 1 can acquire input feature information from the content.
- the script generation system 1 generates a script by inputting input setting information into the language model M. Because the language model M is able to generate a flexible script that corresponds to the settings of the video, the script generation system 1 can improve the accuracy of the script that is the product of the language model M. For example, if the input setting information indicates the distribution date and time of the video, the script generation system 1 can generate a script of a length that corresponds to the distribution date and time of the video. If the input setting information indicates the length of the video, the script generation system 1 can generate a script of a length that corresponds to the length of the video.
- the script generation system 1 generates a script by inputting input session information into the language model M. Because the language model M is able to generate flexible scripts that correspond to video sessions, the script generation system 1 can improve the accuracy of the script, which is the product of the language model M. For example, if the language model M attempts to generate a large number of products at once, the accuracy of the products may decrease. In this regard, if the language model M generates a script for each session, the amount of products that the language model M generates at one time can be reduced, and the script generation system 1 can improve the accuracy of the script.
- the script generation system 1 generates a script according to the level of detail by inputting input session information according to the level of detail into the language model M. Because the language model M can generate flexible scripts according to the level of detail, the script generation system 1 can improve the accuracy of the script, which is the product of the language model M. For example, the script generation system 1 can generate a simplified script requested by a store for which a simplified script showing the overall progress of the live broadcast is sufficient. In this case, the amount of text generated by the language model M can be reduced, thereby reducing the processing load on the script generation device 10. If the language model M is stored in a computer other than the script generation device 10, the processing load on the other computer can be reduced. For example, the script generation system 1 can generate a detailed script requested by a store that requests a detailed script showing the progress of the live broadcast in detail.
- the script generation system 1 For each session, the script generation system 1 generates a script portion for that session by inputting input session information for that session into the language model M. For the second or subsequent session among multiple sessions, the script generation system 1 generates the script portion for the second or subsequent session by inputting the script portions of sessions before the second or subsequent session into the language model M. The script generation system 1 generates a script based on the script portions of each of the multiple sessions. This allows the script generation system 1 to generate a script that has connections between sessions. For example, if the language model M attempts to generate a large amount of text at once, the processing load on the script generation device 10 may increase. However, by having the language model M generate small script portions for each session, the processing load on the script generation device 10 can be reduced. Furthermore, if the language model M attempts to generate a large amount of text at once, the accuracy of the generated product may decrease. However, by having the language model M generate small script portions for each session, the script generation system 1 can increase the accuracy of the final script.
- the script generation system 1 also acquires first input information prepared in advance and second input information generated by inputting the first input information into the language model M or another language model.
- the script generation system 1 generates a script by inputting the first input information and the second input information into the language model M. This allows the script generation system 1 to input a wide variety of input information into the language model M, thereby improving the accuracy of the script.
- FIG. 8 is a diagram showing an example of functions realized in the modified example.
- the script generation device 10 in the modified example is realized by a question and answer generation unit 103, a first sample acquisition unit 104, and a second sample acquisition unit 105.
- Each of the question and answer generation unit 103, the first sample acquisition unit 104, and the second sample acquisition unit 105 is realized by the control unit 11.
- the live streaming service accepts comments from viewers.
- viewers may enter questions about the products or services introduced by the performers.
- the performers may answer questions from viewers during the live streaming.
- questions that are expected from viewers and answers to those questions are generated from the contents of the script generated by the script generation unit 102 so that the performers can prepare for questions that are expected from viewers.
- the script generation system 1 of variant example 1 includes a question and answer generation unit 103.
- the question and answer generation unit 103 generates questions and answers related to the video by inputting a script into language model M or another language model.
- the question and answer generation unit 103 may input information other than the script into language model M or another language model.
- the question and answer generation unit 103 may input at least one of input feature information, input setting information, and input session information into language model M or another language model.
- Questions and answers are data that include at least one question and at least one answer to that question.
- the questions and answers may be in any data format.
- the questions and answers may be in a spreadsheet data format, CSV format, text format, document format, or other format.
- the question and answer generation unit 103 can generate any number of questions and answers. The number of questions and answers that the question and answer generation unit 103 should generate may or may not be specified in advance.
- the question and answer generation unit 103 may input a default prompt to language model M or another language model indicating that it should generate questions and answers.
- the default prompt may indicate the processing content to be performed by language model M or another language model (the role of language model M or another language model, or the type of product to be generated by language model M or another language model), such as "You are an AI that generates questions and answers from a script entered into you.”
- language model M or another language model divides the script input to it into tokens and calculates the embedding representation (feature vector) of each token based on parameters adjusted through learning.
- Language model M or another language model predicts the next text as needed based on the order of the embedded representations of the tokens, and then generates questions and answers as products.
- language model M or another language model generates and outputs questions and answers in accordance with the script input to it based on a default prompt.
- the question and answer generation unit 103 obtains the questions and answers output from language model M or another language model.
- the question and answer generation unit 103 stores the questions and answers in the script database DB.
- the server 20 transmits the questions and answers stored in the script database DB to the performer device 30.
- the question and answer generation unit 103 may store the questions and answers in a database other than the script database DB.
- the question and answer generation unit 103 may record the questions and answers in a computer other than the script generation device 10 or in an external information storage medium.
- the questions and answers generated by the question and answer generation unit 103 may be transmitted to a computer other than the performer device 30.
- the script generation system 1 of variant 1 generates questions and answers related to a video by inputting a script into language model M or another language model.
- the script generation system 1 can support the progress of the performers by generating questions and answers based on the script.
- an example is given of a case where input information, one example of which is input feature information, is input to the language model M. Any information that can be used to generate a script may be input to the language model M.
- a script sample is input to the language model M.
- the sample is script data used as a reference for the language model M.
- the sample may be in any data format.
- the sample may be a script format or template.
- the sample may be generated manually or by the language model M.
- the sample may be a transcript of the audio of a video.
- the script generation system 1 of variant example 2 includes a first sample acquisition unit 104.
- the first sample acquisition unit 104 acquires samples of other videos introducing products different from the products introduced in the video for which a script is to be generated, or other services different from the services introduced in the video for which a script is to be generated, samples in which the features of the other products or services are omitted.
- the other videos are videos other than the video for which a script is to be generated.
- the other videos may be videos that have been distributed in the past, or videos for which a script has been generated but which have not yet been distributed.
- the other videos may be videos distributed by services other than the live streaming service.
- the other products or services are products or services introduced in the other videos.
- the data storage unit 100 stores samples.
- the first sample acquisition unit 104 acquires samples from the data storage unit 100.
- the samples may be stored in a computer other than the script generation device 10, or in an external information storage medium. In this case, the first sample acquisition unit 104 simply acquires the samples from the other computer or external information storage medium.
- the first sample acquisition unit 104 acquires a sample from a script previously generated by the script generation unit 102, in which portions of words included in the input feature information used when generating the script have been masked. Masking involves hiding specific portions (for example, filling those portions with spaces or specific symbols).
- the first sample acquisition unit 104 may acquire a sample by masking the script based on the input feature information.
- the first sample acquisition unit 104 may also acquire a sample in which masking has been completed by a function other than the first sample acquisition unit 104.
- the method of omitting features of other products or services is not limited to masking.
- the first sample acquisition unit 104 may acquire a sample by deleting, from a script previously generated by the script generation unit 102, a portion of a word contained in the input feature information used when generating the script.
- the first sample acquisition unit 104 may also acquire a sample from which the portion has been deleted by a function other than the first sample acquisition unit 104.
- the script generation unit 102 of variant example 2 generates a script by further inputting samples into the language model M. That is, the script generation unit 102 inputs the input information and samples described in the embodiment into the language model M.
- the language model M divides the input information and samples into tokens and calculates embedded representations for each token based on parameters adjusted by learning.
- the language model M predicts the subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script corresponding to the sample as a product.
- the language model M may generate and output a script corresponding to the sample based on a default prompt.
- the default prompt may include text indicating that the language model M should refer to the sample, such as "Please use this sample as a reference when generating a script.”
- the script generation unit 102 obtains the script output from the language model M.
- the script generation system 1 of variant example 2 generates a script by further inputting samples into the language model M from which the characteristics of other products or services have been omitted. This makes the language model M less susceptible to the influence of the characteristics of other products or services, allowing the script generation system 1 to generate a script that reflects the characteristics of the products or services introduced in the video for which the script is being generated. In other words, the script generation system 1 can improve the accuracy of the script.
- the script generation system 1 of variant example 3 includes a second sample acquisition unit 105.
- the second sample acquisition unit 105 acquires samples of other videos that feature the same performers as the performers in the video.
- the meaning of the sample is the same as in variant example 2.
- the second sample acquisition unit 105 identifies the performers of the video for which a script is to be generated based on an orientation sheet stored in the script database DB.
- the second sample acquisition unit 105 identifies other videos that feature the same performers as the identified performers based on the orientation sheet stored in the script database DB.
- the second sample acquisition unit 105 acquires the scripts of the other videos that feature the identified performers as samples.
- the second sample acquisition unit 105 may acquire as samples the scripts of all other videos featuring the same performer, or may acquire as samples the scripts of some other videos featuring the same performer. Furthermore, when combining variants 2 and 3, the second sample acquisition unit 105 may acquire as samples scripts of other videos featuring the same performer as the video for which a script is to be generated, from which information about other products or services has been omitted. The second sample acquisition unit 105 may also acquire as samples the scripts of videos featuring the same performer that have been distributed via a service other than a live distribution service.
- the script generation unit 102 of Modification 3 generates a script by further inputting samples into the language model M. That is, the script generation unit 102 inputs the input information and samples described in the embodiment into the language model M.
- the language model M divides the input information and samples into tokens and calculates embedded representations for each token based on parameters adjusted through learning.
- the language model M predicts the subsequent text as necessary based on the order of the embedded representations of the tokens, and then generates a script corresponding to the sample as a product.
- the language model M may generate and output a script corresponding to the sample based on a default prompt similar to that of Modification 2.
- the script generation unit 102 obtains the script output from the language model M.
- the script generation system 1 of variant example 3 generates a script by inputting further samples of other videos in which the same performer appears into the language model M. This allows the script generation system 1 to reflect the characteristics of the performer in the script. In other words, the script generation system 1 can improve the accuracy of the script.
- script generation system 1 can also be applied to other situations.
- Script generation system 1 may also generate scripts for videos other than live streaming services.
- script generation system 1 may generate scripts for videos that are not streamed in real time, scripts for videos that advertise products or services, scripts for videos in which AI (artificial voice) speaks instead of a human, scripts for videos that introduce products or services sold through services other than e-commerce services, or scripts for other videos.
- the script generation system 1 may generate a script for a video introducing products or services other than e-commerce services.
- the script generation system 1 may generate a script for a video introducing products or services sold in physical stores, services provided at accommodation facilities such as hotels, communication services, payment services, financial services, e-book services, services provided at beauty salons or restaurants, or products or services introduced on social media.
- the products introduced in the video are not limited to tangible objects, but may also be content such as music or movies.
- the functions described as being realized by the script generation device 10 may be realized by the server 20, the performer device 30, or another computer.
- the processing described as being realized by the script generation device 10 may be shared among multiple computers.
- script generation system can also be configured as follows.
- an input information acquisition unit that acquires input information to be input to a trained language model capable of generating a product described in a natural language, the input information being related to a video introducing a product or service; a script generation unit that generates a script for the video by inputting the input information into the language model;
- a script generation system including: (2) the input information acquisition unit acquires input feature information, which is the input information related to features of the product or the service; the script generation unit generates the script by inputting the input feature information into the language model.
- a script generation system according to (1) (3) the input information acquisition unit acquires the input feature information generated by inputting extracted information extracted from content related to features of the product or the service into the language model or another language model; A script generation system according to (2).
- the input information acquisition unit acquires input setting information, which is the input information related to settings of the video; the script generation unit generates the script by inputting the input setting information into the language model.
- a script generation system according to any one of (1) to (3).
- the input information acquisition unit acquires input session information, which is the input information related to each of a plurality of sessions in the video; the script generation unit generates the script by inputting the input session information into the language model.
- the input information acquisition unit acquires the input session information according to a level of detail related to the script; the script generation unit generates the script according to the level of detail by inputting the input session information according to the level of detail to the language model.
- a script generation system according to (5).
- the script generation unit for each session generating a script portion that is part of the script by inputting the input session information for that session into the language model, the script portion for that session; generating the script portion of the second or later session by inputting the script portion of the session before the second or later session into the language model for the second or later session among the plurality of sessions; generating the script based on the script portions of each of the plurality of sessions; A script generation system according to (5) or (6).
- the input information acquisition unit acquires first input information, which is the input information prepared in advance, and second input information, which is the input information generated by inputting the first input information into the language model or another language model; the script generation unit generates the script by inputting the first input information and the second input information to the language model.
- a script generation system according to any one of (1) to (7).
- the script generation system further includes a question and answer generation unit that generates questions and answers related to the video by inputting the script into the language model or another language model;
- a script generation system according to any one of (1) to (8).
- the script generation system further includes a first sample acquisition unit that acquires a sample of another video introducing a product different from the product or a service different from the service, the sample omitting features of the other product or service; the script generation unit generates the script by further inputting the samples into the language model.
- a script generation system according to any one of (1) to (9).
- the script generation system further includes a second sample acquisition unit that acquires samples of other videos in which the same performers as the performers of the video appear, the script generation unit generates the script by further inputting the samples into the language model.
- a script generation system according to any one of (1) to (10).
Landscapes
- Business, Economics & Management (AREA)
- Strategic Management (AREA)
- Engineering & Computer Science (AREA)
- Accounting & Taxation (AREA)
- Development Economics (AREA)
- Finance (AREA)
- Economics (AREA)
- Game Theory and Decision Science (AREA)
- Entrepreneurship & Innovation (AREA)
- Marketing (AREA)
- Physics & Mathematics (AREA)
- General Business, Economics & Management (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Stored Programmes (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
台本生成システム(1)の入力情報取得部(101)は、自然言語で記述された生成物を生成可能な学習済みの言語モデルに入力される入力情報であって、商品又はサービスが紹介される動画に関する前記入力情報を取得する。台本生成部(102)は、言語モデルに入力情報を入力することによって、動画に関する台本を生成する。
Description
本開示は、台本生成システム、台本生成方法、及びプログラムに関する。
従来、商品又はサービスの紹介のための動画に関する台本を生成する技術が知られている。例えば、特許文献1には、広告データが埋め込まれたコンテンツの演出情報を含むエンジン制作エンジンを受信し、コンテンツ制作エンジンを用いて台本を制作した台本制作者を識別する識別情報が埋め込まれた台本を生成し、視聴者が使用する視聴装置に台本を送信する台本制作装置が記載されている。特許文献1には、台本に関するテンプレートについても記載されている。
しかしながら、特許文献1の台本制作装置では、台本制作者がテンプレートを利用して台本を制作したとしても、テンプレートの内容が固定されているので、台本制作者は、テンプレートの範囲内でしか台本を制作することができない。このため、特許文献1の技術では、最終的に制作される台本の柔軟性を高めることができなかった。この点は、特許文献1のような台本の制作に限られず、商品又はサービスの紹介のための動画に関する台本を生成する従来の技術全般についても同様である。
本開示の目的の1つは、より柔軟な台本を生成することである。
本開示に係る台本生成システムは、自然言語で記述された生成物を生成可能な学習済みの言語モデルに入力される入力情報であって、商品又はサービスが紹介される動画に関する前記入力情報を取得する入力情報取得部と、前記言語モデルに前記入力情報を入力することによって、前記動画に関する台本を生成する台本生成部と、を含む。
本開示によれば、より柔軟な台本を生成できる。
[1.台本生成システムのハードウェア構成]
本開示に係る台本生成システム、台本生成方法、及びプログラムの実施形態の一例を説明する。本実施形態では、台本生成システム、台本生成方法、及びプログラムがライブ配信サービスに適用される場合を例に挙げる。ライブ配信サービスは、不特定多数の者に対してリアルタイムに動画を配信するサービスである。ライブ配信の出演者は、予め用意された台本に沿ってライブ配信の進行を行う。視聴者は、リアルタイムに配信される動画、又は、アーカイブとして保存された動画を視聴する。
本開示に係る台本生成システム、台本生成方法、及びプログラムの実施形態の一例を説明する。本実施形態では、台本生成システム、台本生成方法、及びプログラムがライブ配信サービスに適用される場合を例に挙げる。ライブ配信サービスは、不特定多数の者に対してリアルタイムに動画を配信するサービスである。ライブ配信の出演者は、予め用意された台本に沿ってライブ配信の進行を行う。視聴者は、リアルタイムに配信される動画、又は、アーカイブとして保存された動画を視聴する。
図1は、台本生成システムのハードウェア構成の一例を示す図である。例えば、台本生成システム1は、台本生成装置10、サーバ20、出演者装置30、及び視聴者装置40を含む。台本生成装置10、サーバ20、出演者装置30、及び視聴者装置40の各々は、インターネット又はLAN等のネットワークNに接続される。図1では、台本生成装置10、サーバ20、出演者装置30、及び視聴者装置40の各々を1台ずつ示しているが、これらのうちの少なくとも1つは、複数台存在してもよい。
台本生成装置10は、台本を生成する装置である。例えば、台本生成装置10は、パーソナルコンピュータ、サーバコンピュータ、タブレット、又はスマートフォンである。例えば、台本生成装置10は、制御部11、記憶部12、通信部13、操作部14、及び表示部15を含む。制御部11は、少なくとも1つのプロセッサを含む。記憶部12は、RAM等の揮発性メモリと、フラッシュメモリ等の不揮発性メモリと、の少なくとも一方を含む。通信部13は、有線通信用の通信インタフェースと、無線通信用の通信インタフェースと、の少なくとも一方を含む。操作部14は、タッチパネル等の入力デバイスである。表示部15は、液晶又は有機ELのディスプレイである。
サーバ20は、ライブ配信サービスのサーバコンピュータである。例えば、サーバ20は、制御部21、記憶部22、及び通信部23を含む。制御部21、記憶部22、及び通信部23のハードウェア構成は、それぞれ制御部11、記憶部12、及び通信部13と同様であってよい。
出演者装置30は、出演者の装置である。例えば、出演者装置30は、パーソナルコンピュータ、タブレット、又はスマートフォンである。例えば、出演者装置30は、制御部31、記憶部32、通信部33、操作部34、及び表示部35を含む。制御部31、記憶部32、通信部33、操作部34、及び表示部35のハードウェア構成は、それぞれ制御部11、記憶部12、通信部13、操作部14、及び表示部15と同様であってよい。出演者装置30には、撮影部36が接続される。撮影部36は、少なくとも1つのカメラを含む。撮影部36は、出演者装置30の内部に含まれてもよい。
視聴者装置40は、視聴者の装置である。例えば、視聴者装置40は、パーソナルコンピュータ、タブレット、又はスマートフォンである。例えば、視聴者装置40は、制御部41、記憶部42、通信部43、操作部44、及び表示部45を含む。制御部41、記憶部42、通信部43、操作部44、及び表示部45のハードウェア構成は、それぞれ制御部11、記憶部12、通信部13、操作部14、及び表示部15と同様であってよい。
なお、記憶部12,22,32,42に記憶されるプログラムは、ネットワークNを介して、台本生成装置10、サーバ20、出演者装置30、又は視聴者装置40に供給されてもよい。また、コンピュータ読み取り可能な情報記憶媒体を読み取る読取部(例えば、メモリカードスロット)と、外部機器とデータの入出力をするための入出力部(例えば、USBポート)と、の少なくとも一方が、台本生成装置10、サーバ20、出演者装置30、又は視聴者装置40に含まれてもよい。例えば、情報記憶媒体に記憶されたプログラムが、読取部及び入出力部の少なくとも一方を介して、台本生成装置10、サーバ20、出演者装置30、又は視聴者装置40に供給されてもよい。
また、台本生成システム1は、少なくとも1つのコンピュータを含めばよい。台本生成システム1に含まれるコンピュータは、図1の例に限られない。例えば、台本生成システム1は、台本生成装置10及びサーバ20だけを含んでもよい。この場合、出演者装置30及び視聴者装置40は、台本生成システム1の外部に存在する。台本生成システム1は、台本生成装置10だけを含んでもよい。この場合、サーバ20、出演者装置30、及び視聴者装置40は、台本生成システム1の外部に存在する。例えば、台本生成システム1は、台本生成装置10と、図1に示さない他のコンピュータと、を含んでもよい。
[2.本実施形態の概要]
本実施形態では、電子商取引サービスで販売される商品又はサービスがライブ配信サービスで紹介される場合を例に挙げる。ライブ配信サービスは、電子商取引サービスの運営者が提供するサービスの1つであってもよいし、電子商取引サービスからは切り離された他のサービスであってもよい。例えば、電子商取引サービスに加盟する店舗は、自身が販売する商品又はサービスを紹介するためにライブ配信を行う。ライブ配信は、任意の者が主体となって行われてよい。例えば、ライブ配信は、店舗ではなく、商品のメーカー、サービスの提供者、又はインフルエンサー等の他の者が主体となって行われてもよい。
本実施形態では、電子商取引サービスで販売される商品又はサービスがライブ配信サービスで紹介される場合を例に挙げる。ライブ配信サービスは、電子商取引サービスの運営者が提供するサービスの1つであってもよいし、電子商取引サービスからは切り離された他のサービスであってもよい。例えば、電子商取引サービスに加盟する店舗は、自身が販売する商品又はサービスを紹介するためにライブ配信を行う。ライブ配信は、任意の者が主体となって行われてよい。例えば、ライブ配信は、店舗ではなく、商品のメーカー、サービスの提供者、又はインフルエンサー等の他の者が主体となって行われてもよい。
図2は、ライブ配信が行われる様子の一例を示す図である。例えば、出演者は、予め用意された台本に従って、撮影部36の前で商品又はサービスを紹介する。出演者は、店舗の関係者であってもよいし、店舗から出演を依頼された他の者であってもよい。出演者装置30は、サーバ20に対し、撮影部36の撮影結果を示すデータ(動画のデータ)を送信する。サーバ20は、出演者装置30から受信したデータに基づいて、視聴者装置40に対し、リアルタイムに動画を配信する。視聴者装置40は、出演者が商品又はサービスを紹介する動画を示す紹介画面SCを、表示部45に表示させる。
本実施形態では、ライブ配信サービスの運営者は、ライブ配信の申込を行った店舗に対し、台本を生成する台本生成サービスを提供する。店舗の担当者は、自身で台本を用意してもよいし、台本生成サービスを利用して台本を用意してもよい。例えば、店舗の担当者が台本生成サービスを利用する場合、店舗の担当者は、後述のオリエンテーションシートに必要事項を記入して、運営者に対し、台本生成サービスの利用を依頼する。運営者は、店舗の担当者からの依頼に基づいて、ライブ配信で紹介される商品又はサービスに応じた台本を生成する。
例えば、運営者が、手作業で台本を書くことによって、台本を生成することも考えるが、運営者の手間がかかる。運営者が、台本をテンプレート化して、一から台本を書く必要がないようにすることも考えられるが、運営者は、テンプレートの範囲内でしか台本を制作することができないので、柔軟性に欠ける。そこで、本実施形態の台本生成システム1は、学習済みの言語モデルを利用して、より柔軟な台本を生成するようにしている。以降、台本生成システム1の詳細を説明する。
[3.台本生成システムで実現される機能]
図3は、台本生成システム1で実現される機能の一例を示す図である。図3では、台本生成装置10で実現される機能の一例が示されている。例えば、台本生成装置10は、データ記憶部100、入力情報取得部101、及び台本生成部102を含む。データ記憶部100は、記憶部12により実現される。入力情報取得部101及び台本生成部102の各々は、制御部11により実現される。
図3は、台本生成システム1で実現される機能の一例を示す図である。図3では、台本生成装置10で実現される機能の一例が示されている。例えば、台本生成装置10は、データ記憶部100、入力情報取得部101、及び台本生成部102を含む。データ記憶部100は、記憶部12により実現される。入力情報取得部101及び台本生成部102の各々は、制御部11により実現される。
[3-1.データ記憶部]
データ記憶部100は、台本の生成に必要なデータを記憶する。例えば、データ記憶部100は、自然言語で記述された生成物を生成可能な学習済みの言語モデルMと、言語モデルMにより生成された台本が格納される台本データベースDBと、を記憶する。なお、データ記憶部100に記憶されるデータは、これらの例に限られない。データ記憶部100は、任意のデータを記憶可能である。
データ記憶部100は、台本の生成に必要なデータを記憶する。例えば、データ記憶部100は、自然言語で記述された生成物を生成可能な学習済みの言語モデルMと、言語モデルMにより生成された台本が格納される台本データベースDBと、を記憶する。なお、データ記憶部100に記憶されるデータは、これらの例に限られない。データ記憶部100は、任意のデータを記憶可能である。
自然言語は、人間が認識可能な言語である。例えば、自然言語は、日本語、英語、又は中国語といった任意の言語であってよい。生成物は、言語モデルMが生成するデータである。言語モデルMは、任意の生成物を生成可能である。例えば、生成物は、テキスト、表、図、画像、動画、又はこれらの組み合わせであってもよい。テキストは、文字、数字、記号、又はこれらの組み合わせを含む。テキストは、任意の形式であってよい。例えば、テキストは、文章、箇条書き、単語の羅列、プログラムコード、又はマークアップ言語で記述されたコードあってもよい。
言語モデルMは、自然言語処理の分野で利用されるモデルである。例えば、言語モデルMは、機械学習の手法が利用されるモデルである。言語モデルMは、大規模言語モデル又は生成AI(Artificial Intelligence)と呼ばれることもある。言語モデルMは、他の名称で呼ばれるモデルであってもよい。例えば、言語モデルMは、入力情報を処理して生成物を生成するための情報処理を実行するプログラムと、当該プログラムにより参照されるパラメータと、を含む。学習によってパラメータが調整される。言語モデルMのパラメータは、公知のパラメータであってよい。例えば、言語モデルMのパラメータは、重み又はバイアスであってよい。本実施形態では、言語モデルMは、事前に大規模なデータセットによる学習が完了しているものとする。
なお、言語モデルMは、公知の種々のタイプのモデルであってよい。例えば、言語モデルMは、GPT(Generative Pre-trained Transformer)、GPT以外のトランスフォーマー系のモデル(例えば、BERT:Bidirectional Encoder Representations from Transformers、又は、T5:Text-To-Text Transfer Transformer)、自然言語処理が可能なニューラルネットワーク、又はその他のモデル(例えば、Pegasus、UniLM、又はElectra)であってもよい。言語モデルMに含まれるプログラム及びパラメータは、これらの公知のモデルと同様であってよい。例えば、言語モデルMは、公知のモデルと全く同じであってもよいし、台本の生成に特化した訓練データによってファインチューニングされたモデルであってもよい。
本実施形態では、データ記憶部100が言語モデルMを記憶する場合を例に挙げるが、言語モデルMは、台本生成装置10以外の他の装置に記憶されてもよい。例えば、他の装置は、オンラインで言語モデルMの機能を提供する会社が管理する装置であってもよい。この場合、台本生成装置10は、他の装置に対し、言語モデルMに入力する入力情報を送信する。他の装置は、台本生成装置10から受信した入力情報を、自身に記憶された言語モデルMに入力する。他の装置は、台本生成装置10に対し、言語モデルMが出力した生成物を送信する。台本生成装置10は、他の装置から、生成物を受信する。
図4は、台本データベースDBの一例を示す図である。例えば、台本データベースDBには、ライブ配信ID、オリエンテーションシート、及び台本が格納される。台本データベースDBには、台本に関する任意の情報が格納されてよい。例えば、台本データベースDBには、後述する変形例で説明する質問と回答、出演者が台本に沿って商品又はサービスを紹介する様子が撮影された動画(例えば、アーカイブ配信用の動画)、又は台本の生成時に利用された入力情報が格納されてもよい。
ライブ配信IDは、個々のライブ配信を識別可能なIDである。例えば、店舗の担当者がライブ配信サービスに申込を行うと、新たなライブ配信IDが発行される。オリエンテーションシートは、ライブ配信サービスで配信される動画の基本的な情報が示されたデータである。例えば、オリエンテーションシートは、紹介の対象となる商品又はサービスに関する情報を含む。商品又はサービスに関する情報は、任意の情報であってよく、例えば、商品名、サービス名、商品若しくはサービス自体の特徴、価格、在庫、色のバリエーション、サイズのバリエーション、又は他の情報であってもよい。オリエンテーションシートは、後述の入力特徴情報、入力設定情報、及び入力セッション情報を含んでもよい。
なお、オリエンテーションシートは、任意のデータ形式であってよい。例えば、オリエンテーションシートは、表計算ソフトのデータ形式、CSV形式、テキスト形式、文書形式、又は他の形式であってもよい。また、オリエンテーションシートは、商品又はサービス以外の他の情報を含んでもよい。例えば、オリエンテーションシートは、ライブ配信を希望する店舗の情報(例えば、店舗のID又は名前)、ライブ配信の日時、ライブ配信の長さ(尺)、出演者の情報(例えば、出演者の氏名、芸名、又はプロフィール)、台本の制作依頼の有無、後述する台本の詳細度、NGワード、又は他の情報を含んでもよい。
例えば、ライブ配信を希望する店舗の担当者は、自身の端末を操作してオリエンテーションシートに必要事項を入力する。オリエンテーションシートの全ての項目の入力が要求されてもよいし、一部の項目の入力だけが要求されてもよい。店舗の担当者の端末は、サーバ20に対し、オリエンテーションシートを送信する。サーバ20は、店舗の担当者の端末から、オリエンテーションシートを受信する。サーバ20は、ライブ配信IDを発行し、ライブ配信ID及びオリエンテーションシートを記憶部22に記録する。オリエンテーションシートは、任意の者によって生成されてよい。例えば、ライブ配信サービスの運営者が店舗からの希望を元にオリエンテーションシートを生成してもよい。
例えば、ライブ配信サービスの運営者は、オリエンテーションシートの内容に基づいて、店舗によるライブ配信を許可するか否かを審査する。店舗が審査に合格し、店舗が台本生成サービスを利用することがオリエンテーションシートに示されていた場合、台本生成装置10は、サーバ20からライブ配信ID及びオリエンテーションシートを取得し、台本データベースDBに格納する。
本実施形態では、台本データベースDBに台本が格納される。本実施形態で台本と記載した箇所は、台本を示すデータを意味する。台本は、任意のデータ形式であってよい。例えば、台本は、テキスト形式、文書形式、表計算ソフトのデータ形式、CSV形式、又は他の形式であってもよい。台本のデータ形式は、後述のデフォルトプロンプトで指定されてもよい。台本データベースDBには、台本の生成時に利用された入力情報が格納されてもよい。台本は、後述のセッションごとに別々のデータに分けられていてもよい。
[3―2.入力情報取得部]
入力情報取得部101は、言語モデルMに入力される入力情報であって、商品又はサービスが紹介される動画に関する入力情報を取得する。入力情報は、自然言語で記述されたテキスト(例えば、文章)を示す。入力情報は、プロンプトと呼ばれることもある。入力情報は、言語モデルMが処理可能な形式の情報であればよく、少なくとも1つの文字を含むものとする。入力情報は、自然言語で記述されたテキスト以外の他の情報(例えば、画像又は動画)を含んでもよい。言語モデルMには、入力情報とともに、入力情報以外の他の情報(例えば、生成物のサンプルとなる情報)が入力されてもよい。
入力情報取得部101は、言語モデルMに入力される入力情報であって、商品又はサービスが紹介される動画に関する入力情報を取得する。入力情報は、自然言語で記述されたテキスト(例えば、文章)を示す。入力情報は、プロンプトと呼ばれることもある。入力情報は、言語モデルMが処理可能な形式の情報であればよく、少なくとも1つの文字を含むものとする。入力情報は、自然言語で記述されたテキスト以外の他の情報(例えば、画像又は動画)を含んでもよい。言語モデルMには、入力情報とともに、入力情報以外の他の情報(例えば、生成物のサンプルとなる情報)が入力されてもよい。
図5は、言語モデルMに対する入力及び出力の一例を示す図である。本実施形態では、商品又はサービスが紹介される動画の台本が、言語モデルMの生成物として生成されるので、入力情報は、台本の生成対象となる動画に関する少なくとも1つの情報を含む。図5の例では、入力情報は、デフォルトプロンプト、入力特徴情報、入力設定情報、及び入力セッション情報といった4つの情報を含む。入力情報は、任意の数の情報を含んでよい。例えば、入力情報は、1つ、2つ、3つ、又は5つ以上の情報を含んでもよい。
デフォルトプロンプトは、予め用意されたプロンプトである。デフォルトプロンプトは、入力情報の一種である。デフォルトプロンプトは、任意の情報を含むことができる。例えば、デフォルトプロンプトは、自然言語で記述されたテキストを含む。デフォルトプロンプトは、文章、プログラムコード、JSON等のマークアップ言語で記述されたコード、又は他のテキストを含んでもよい。デフォルトプロンプトは、テキスト以外の他の情報(例えば、画像又は動画)を含んでもよい。デフォルトプロンプトは、データ記憶部100に記憶される。台本の生成を担当する担当者は、デフォルトプロンプトの内容を編集してもよい。
本実施形態では、言語モデルMが、特定の目的に特化したモデルではなく、種々の生成物を生成可能な汎用的なモデルである場合を例に挙げる。このため、汎用的な言語モデルMが、自身が実行すべき処理内容を認識できるように、入力情報は、言語モデルMが実行すべき処理内容を示すデフォルトプロンプトを含むものとする。言語モデルMが実行すべき処理内容は、言語モデルMの役割(タスク)、又は、言語モデルMが生成すべき生成物の種類ということもできる。
例えば、汎用的な言語モデルMは、入力情報に含まれるデフォルトプロンプトによって、自身が実行すべき処理内容を認識する。即ち、汎用的な言語モデルMは、入力情報に含まれるデフォルトプロンプトによって、自身の役割、又は、言語モデルMが生成すべき生成物の種類を認識する。図5の例では、デフォルトプロンプトは、「あなたは、商品又はサービスが紹介されるライブ配信の脚本家です。あなたに入力された入力情報に基づいて、ライブ配信の台本を生成して下さい。」といったように、言語モデルMが入力情報に基づいて台本を生成すべきであることを示す。
なお、デフォルトプロンプトは、図5の例に限られない。例えば、デフォルトプロンプトは、図5の文言以外の他の文言で、言語モデルMが台本を生成すべきことを示してもよい。デフォルトプロンプトは、言語モデルMが台本を生成すべきことを示す情報以外の他の情報を含んでもよい。例えば、デフォルトプロンプトは、台本の言語、台本の分量(例えば、文字数又はページ数)、台本のレイアウト(例えば、出演者の氏名の後にセリフを記載する等の形式)、台本のファイル形式、又は他の情報を含んでもよい。デフォルトプロンプトは、言語モデルM自体の設定を含んでもよい。
また、言語モデルMは、台本の生成に特化したモデルであってもよい。この場合には、言語モデルMが台本を生成すべきことがデフォルトプロンプトで明示されなくても、言語モデルMが台本を生成できるので、入力情報は、デフォルトプロンプトを含まなくてもよい。言語モデルMが台本の生成に特化したモデルである場合には、台本の生成に特化した訓練データが言語モデルMに学習されているものとする。例えば、訓練用の入力情報と、正解となる台本と、のペアを含む訓練データが言語モデルMに学習されている。言語モデルMには、多数の訓練データが学習されていてもよい。台本の生成に特化した言語モデルMは、デフォルトプロンプトがなくても、事前に調整されたパラメータに基づいて、自身に入力された入力情報に基づいて、台本を生成可能である。
図5の例では、入力情報は、デフォルトプロンプト以外にも、入力特徴情報、入力設定情報、及び入力セッション情報を含む。入力情報は、入力特徴情報、入力設定情報、及び入力セッション情報の一部だけを含んでもよい。例えば、入力情報は、入力特徴情報だけ、入力設定情報だけ、入力セッション情報だけ、入力特徴情報及び入力設定情報だけ、入力特徴情報及び入力セッション情報だけ、又は入力設定情報及び入力セッション情報を含んでもよい。入力情報は、図5に示さない他の情報を含んでもよい。
例えば、入力情報取得部101は、台本データベースDBから、台本の生成対象となるライブ配信のライブ配信IDに関連付けられたオリエンテーションシートを取得する。台本の生成対象となるライブ配信は、台本生成装置10の操作部14から指定されるものとするが、他の方法によって特定されてもよい。例えば、入力情報取得部101は、台本データベースDBを参照し、まだ台本が生成されていないライブ配信を、台本の生成対象となるライブ配信として特定してもよい。入力情報取得部101は、配信日時までの期間の長さが閾値未満になったライブ配信を、台本の生成対象となるライブ配信として特定してもよい。
例えば、入力情報取得部101は、商品又はサービスの特徴に関する入力情報である入力特徴情報を取得する。入力特徴情報は、入力情報の一種である。商品又はサービスの特徴は、商品又はサービスの説明ということもできる。入力特徴情報は、自然言語のテキストで記述される。例えば、入力特徴情報は、文字、数字、記号、又はこれらの組み合わせによって、商品又はサービスの特徴を示す。例えば、入力特徴情報は、商品又はサービスの識別情報(例えば、名前、型番、メーカー、又はJANコード)、分類(例えば、ジャンル、カテゴリ、属性、又は属性値)、見た目(例えば、デザイン、サイズ、色、又は模様)、品質、機能、素材、価格、割引率、アピールポイント、店舗の情報、又は他の情報を示してもよい。入力特徴情報は、電子商取引サービスに掲載されている商品又はサービスのタイトル、説明文、検索用の情報、又は他の情報であってもよい。
例えば、入力情報取得部101は、オリエンテーションシートに含まれる入力特徴情報を取得する。入力情報取得部101は、オリエンテーションシートに含まれる入力特徴情報を取得するのではなく、オリエンテーションシートに含まれる情報に基づいて、入力特徴情報を生成してもよい。入力情報取得部101は、オリエンテーションシートには含まれない情報に基づいて、入力特徴情報を生成してもよい。入力情報取得部101は、オリエンテーションシートに含まれる入力特徴情報の情報量が十分ではない場合、又は、オリエンテーションシートに入力特徴情報が含まれない場合に、入力特徴情報を生成してもよい。
例えば、入力情報取得部101は、商品又はサービスの特徴に関するコンテンツから抽出された抽出情報が言語モデルM又は他の言語モデルに入力されることによって生成された入力特徴情報を取得してもよい。コンテンツは、電子的な情報である。例えば、コンテンツは、ウェブサイト、ウェブ広告、デジタルカタログ、デジタルパンフレット、デジタルチラシ、スキャナで取り込まれた紙を示す画像、又は他の情報であってもよい。本実施形態では、コンテンツの一例としてウェブサイトを説明する。
図6は、コンテンツから抽出された抽出情報に基づく入力情報の一例を示す図である。例えば、入力情報取得部101は、公知のウェブ検索サービスを利用して、紹介対象となる商品又はサービスの検索を実行する。検索時の検索クエリは、台本生成装置10を操作する者によって入力されてもよいし、オリエンテーションシートに含まれる入力特徴情報に基づいて取得されてもよい。入力情報取得部101は、検索でヒットしたウェブサイトから、商品又はサービスが示された画像Iをコンテンツとして取得する。
例えば、入力情報取得部101は、コンテンツに対して光学文字認識を実行し、コンテンツからテキストを、抽出情報として抽出する。抽出情報は、テキスト以外の他の情報(例えば、表又は図)であってもよい。入力情報取得部101は、コンテンツから抽出された抽出情報を、そのまま入力特徴情報として取得してもよい。コンテンツがテキストを含む場合には、入力情報取得部101は、特に光学文字認識を実行せずに、コンテンツからテキストを抽出することによって、抽出情報を抽出してもよい。抽出情報の抽出は、入力情報取得部101以外の他の機能によって実行されてもよい。
例えば、入力情報取得部101は、予め用意された入力情報である第1入力情報と、言語モデルM又は他の言語モデルに第1入力情報が入力されることによって生成された入力情報である第2入力情報と、を取得してもよい。第1入力情報は、オリエンテーションシートに含まれる情報であってもよい。例えば、オリエンテーションシートに含まれる入力特徴情報は、第1入力情報に相当する。他の言語モデルは、台本を生成する言語モデルMとは異なる言語モデルである。他の言語モデルは、言語モデルMと同様、任意のモデルであってよい。他の言語モデルの説明は、言語モデルMの説明における「言語モデルM」の記載を「他の言語モデル」と読み替えればよい。
本実施形態では、コンテンツから抽出された抽出情報には、商品又はサービスの特徴とは関係のない情報が含まれる可能性があるので、入力情報取得部101は、言語モデルM又は他の言語モデルに抽出情報を入力し、言語モデルM又は他の言語モデルから出力された入力特徴情報を取得する。例えば、入力情報取得部101は、言語モデルM又は他の言語モデルに対し、抽出情報とともに、オリエンテーションシートに含まれる入力特徴情報を入力してもよい。この場合、オリエンテーションシートに含まれる入力特徴情報が第1入力情報に相当する。
例えば、入力情報取得部101は、言語モデルM又は他の言語モデルに対し、入力特徴情報を生成すべきことを示すデフォルトプロンプトを入力してもよい。例えば、デフォルトプロンプトは、「あなたは、コンテンツから抽出された商品又はサービスの情報に基づいて、商品又はサービスの特徴を示す入力特徴情報を生成するAIです。」といったように、言語モデルM又は他の言語モデルが実行すべき処理内容(言語モデルM若しくは他の言語モデルの役割、又は、言語モデルM若しくは他の言語モデルが生成すべき生成物の種類)を示してもよい。なお、言語モデルM又は他の言語モデルが、抽出情報から入力特徴情報を生成することに特化したモデルである場合には、このようなデフォルトプロンプトが入力されなくてもよい。
例えば、言語モデルM又は他の言語モデルは、自身に入力された抽出情報等の情報をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。トークンの分割方法は、公知の方法であってよい。埋め込み表現は、トークンの特徴を示す。例えば、埋め込み表現は、多次元ベクトルであってもよいし、他の形式であってもよい。言語モデルM又は他の言語モデルは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、入力特徴情報を生成物として生成する。
例えば、言語モデルM又は他の言語モデルは、デフォルトプロンプトに基づいて、抽出情報の中から、オリエンテーションシートに含まれる入力特徴情報とは重複しない情報を抽出し、入力特徴情報として出力する。入力情報取得部101は、言語モデルM又は他の言語モデルから出力された入力特徴情報を取得する。この場合、言語モデルM又は他の言語モデルから出力された入力特徴情報が、第2入力情報に相当する。図6の例では、コンテンツの一例であるウェブサイトの画像Iから光学文字認識によって抽出された文字が言語モデルMによって体裁が整えられる。言語モデルMは、オリエンテーションシートに含まれない特徴を示す入力特徴情報を出力する。
例えば、入力情報取得部101は、動画の設定に関する入力情報である入力設定情報を取得する。入力設定情報は、入力情報の一種である。動画の設定は、商品又はサービス自体の特徴ではなく、動画自体の特徴ということもできる。動画の設定は、動画に関連付けられた保存される情報ということもできる。入力設定情報は、自然言語のテキストで記述される。例えば、入力設定情報は、文字、数字、記号、又はこれらの組み合わせによって、動画の設定を示す。例えば、入力設定情報は、動画の配信日時、長さ(尺)、ターゲット層(例えば、年齢層又は性別)、出演者の情報(例えば、氏名、プロフィール、又は役割)、NG表現、動画のタイトル、動画の要約、又は他の情報である。
例えば、入力情報取得部101は、オリエンテーションシートに含まれる入力設定情報を取得する。入力情報取得部101は、オリエンテーションシートに含まれる入力設定情報を取得するのではなく、オリエンテーションシートに含まれる情報に基づいて、入力設定情報を生成してもよい。入力情報取得部101は、オリエンテーションシートには含まれない情報に基づいて、入力設定情報を生成してもよい。入力情報取得部101は、オリエンテーションシートに含まれる入力設定情報の情報量が十分ではない場合、又は、オリエンテーションシートに入力設定情報が含まれない場合に、入力設定情報を生成してもよい。
例えば、入力情報取得部101は、言語モデルM又は他の言語モデルに、オリエンテーションシートに含まれる情報(例えば、オリエンテーションシートに含まれる入力特徴情報、入力設定情報、又は入力セッション情報)を入力し、言語モデルM又は他の言語モデルから出力された入力設定情報を取得してもよい。この場合、オリエンテーションシートに含まれる情報が第1入力情報に相当する。オリエンテーションシートに含まれない情報が言語モデルM又は他の言語モデルに入力される場合には、オリエンテーションシートに含まれない情報が第1入力情報に相当する。
例えば、入力情報取得部101は、言語モデルM又は他の言語モデルに対し、入力設定情報を生成すべきことを示すデフォルトプロンプトを入力してもよい。例えば、デフォルトプロンプトは、「あなたは、あなたに入力された情報に基づいて、動画の設定を示す入力設定情報を生成するAIです。」といったように、言語モデルM又は他の言語モデルが実行すべき処理内容(言語モデルM若しくは他の言語モデルの役割、又は、言語モデルM若しくは他の言語モデルが生成すべき生成物の種類)を示してもよい。
例えば、言語モデルM又は他の言語モデルは、自身に入力された情報(入力設定情報の生成で利用される情報)をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルM又は他の言語モデルは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、入力設定情報を生成物として生成する。例えば、言語モデルM又は他の言語モデルは、デフォルトプロンプトに基づいて、自身に入力された情報に応じた動画のタイトル及び要約等を生成し、入力設定情報として出力する。入力情報取得部101は、言語モデルM又は他の言語モデルから出力された入力設定情報を取得する。この場合、言語モデルM又は他の言語モデルから出力された入力設定情報が、第2入力情報に相当する。
例えば、入力情報取得部101は、動画における複数のセッションの各々に関する入力情報である入力セッション情報を取得する。セッションは、動画を構成する個々の部分である。セッションは、セクション等の他の名前で呼ばれることもある。動画は、任意の観点で複数のセッションに分けられてよい。例えば、話題ごとにセッションが分けられてもよいし、時間でセッションが分けられてもよい。入力セッション情報は、入力情報の一種である。入力セッション情報は、自然言語のテキストで記述される。例えば、入力セッション情報は、文字、数字、記号、又はこれらの組み合わせによって、セッションを示す。入力セッション情報は、個々のセッションで紹介される概要を示してもよい。例えば、セッション情報は、セッションの順序を示す番号、見出し、概要を示すテキスト、時間的な長さ、又は他の情報であってもよい。
例えば、入力情報取得部101は、オリエンテーションシートに含まれる入力セッション情報を取得する。入力情報取得部101は、オリエンテーションシートに含まれる入力セッション情報を取得するのではなく、オリエンテーションシートに含まれる情報に基づいて、入力セッション情報を生成してもよい。入力情報取得部101は、オリエンテーションシートには含まれない情報に基づいて、入力セッション情報を生成してもよい。入力情報取得部101は、オリエンテーションシートに含まれる入力セッション情報の情報量が十分ではない場合、又は、オリエンテーションシートに入力セッション情報が含まれない場合に、入力セッション情報を生成してもよい。
例えば、入力情報取得部101は、言語モデルM又は他の言語モデルに、オリエンテーションシートに含まれる情報(例えば、オリエンテーションシートに含まれる入力特徴情報、入力設定情報、又は入力セッション情報)を入力し、言語モデルM又は他の言語モデルから出力された入力セッション情報を取得する。この場合、オリエンテーションシートに含まれる情報が第1入力情報に相当する。オリエンテーションシートに含まれない情報が言語モデルM又は他の言語モデルに入力される場合には、オリエンテーションシートに含まれない情報が第1入力情報に相当する。
例えば、入力情報取得部101は、言語モデルM又は他の言語モデルに対し、入力セッション情報を生成すべきことを示すデフォルトプロンプトを入力してもよい。例えば、デフォルトプロンプトは、「あなたは、あなたに入力された情報に基づいて、動画のセッションを示す入力セッション情報を生成するAIです。」といったように、言語モデルM又は他の言語モデルが実行すべき処理内容(言語モデルM若しくは他の言語モデルの役割、又は、言語モデルM若しくは他の言語モデルが生成すべき生成物の種類)を示してもよい。デフォルトプロンプトには、セッションの数が指定されていてもよい。
例えば、言語モデルM又は他の言語モデルは、自身に入力された情報(入力セッション情報の生成で利用される情報)をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルM又は他の言語モデルは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、入力セッション情報を生成物として生成する。例えば、言語モデルM又は他の言語モデルは、デフォルトプロンプトに基づいて、自身に入力された情報に応じたセッションの見出し等を生成し、入力セッション情報として出力する。入力情報取得部101は、言語モデルM又は他の言語モデルから出力された入力セッション情報を取得する。この場合、言語モデルM又は他の言語モデルから出力された入力セッション情報が、第2入力情報に相当する。
本実施形態では、入力情報取得部101は、台本に関する詳細度に応じた入力セッション情報を取得する。詳細度は、台本の詳しさの程度である。詳細度は、台本のボリュームということもできる。例えば、簡易版又は詳細版といった2段階の詳細度であってもよいし、3段階以上の詳細度であってもよい。詳細度は、任意の者によって指定可能である。例えば、店舗の担当者が詳細度を指定してもよい。入力情報取得部101は、詳細度が高いほど、動画が細かく区切られるように、入力セッション情報を取得する。詳細度は、文字、数字、記号、又はこれらの組み合わせによって示される。
例えば、オリエンテーションシートに詳細度が含まれる場合、入力情報取得部101は、オリエンテーションシートに含まれる詳細度を取得する。詳細度に応じた入力セッション情報がオリエンテーションシートに既に含まれている場合には、入力情報取得部101は、オリエンテーションシートに含まれる詳細度に応じた入力セッション情報を取得する。入力情報取得部101は、言語モデルM又は他の言語モデルに基づいて、詳細度に応じた入力セッション情報を取得してもよい。
例えば、詳細度が高いほど、セッションの数が多くなる。詳細度が高いほど、セッションが更に細かく分割されたサブセッションが生成される。サブセッションは、セッションの下位の階層ということもできる。セッションは、2段階ではなく、3段階以上の階層が存在してもよい。詳細度が高いほど、セッションの階層数が多くなってもよい。例えば、入力情報取得部101は、詳細度が簡易版を示す場合には、先述した方法で入力セッション情報を生成し、それ以上細かくはセッションを分割せずに、入力セッション情報を取得する処理を終了する。
例えば、入力情報取得部101は、詳細度が詳細版を示す場合には、先述した方法で生成した入力セッション情報に基づいて、更にセッションを細かく分割したサブセッションを、言語モデルM又は他の言語モデルに生成させる。例えば、入力情報取得部101は、言語モデルM又は他の言語モデルに入力セッション情報を入力し、言語モデルM又は他の言語モデルから出力された入力セッション情報を取得してもよい。
例えば、入力情報取得部101は、言語モデルM又は他の言語モデルに対し、詳細度に応じた入力セッション情報を生成すべきことを示すデフォルトプロンプトを入力してもよい。例えば、デフォルトプロンプトは、「あなたは、あなたに入力された入力セッション情報に基づいて、セッションを更に細かく分割したサブセッションを示す入力セッション情報を生成するAIです。」といったように、言語モデルM又は他の言語モデルが実行すべき処理内容(言語モデルM若しくは他の言語モデルの役割、又は、言語モデルM若しくは他の言語モデルが生成すべき生成物の種類)を示してもよい。
例えば、言語モデルM又は他の言語モデルは、自身に入力された入力セッション情報をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルM又は他の言語モデルは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、自身に入力された入力セッション情報よりも詳細な入力セッション情報を生成物として生成する。例えば、言語モデルM又は他の言語モデルは、デフォルトプロンプトに基づいて、自身に入力された入力セッション情報に応じたサブセッションの見出し等を生成し、入力セッション情報として出力する。入力情報取得部101は、言語モデルM又は他の言語モデルから出力された入力セッション情報を取得する。
なお、入力情報取得部101は、言語モデルM又は他の言語モデルに詳細度を入力してもよい。この場合、言語モデルM又は他の言語モデルは、自身に入力された詳細度に基づいて、サブセッションの見出し等を示す入力セッション情報を取得してもよい。
[3―3.台本生成部]
台本生成部102は、言語モデルMに入力情報を入力することによって、動画に関する台本を生成する。言語モデルMは、入力情報をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルMは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、台本を生成物として生成する。例えば、言語モデルMは、デフォルトプロンプトに基づいて、入力情報に応じた台本を生成して出力する。入力情報取得部101は、言語モデルMから出力された台本を取得する。
台本生成部102は、言語モデルMに入力情報を入力することによって、動画に関する台本を生成する。言語モデルMは、入力情報をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルMは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、台本を生成物として生成する。例えば、言語モデルMは、デフォルトプロンプトに基づいて、入力情報に応じた台本を生成して出力する。入力情報取得部101は、言語モデルMから出力された台本を取得する。
例えば、台本生成部102は、言語モデルMに入力特徴情報を入力することによって、台本を生成する。言語モデルMは、入力特徴情報をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルMは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、台本を生成物として生成する。例えば、言語モデルMは、デフォルトプロンプトに基づいて、入力特徴情報に応じた台本を生成して出力する。入力特徴情報取得部は、言語モデルMから出力された台本を取得する。
例えば、台本生成部102は、言語モデルMに入力設定情報を入力することによって、台本を生成する。言語モデルMは、入力設定情報をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルMは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、台本を生成物として生成する。例えば、言語モデルMは、デフォルトプロンプトに基づいて、入力設定情報に応じた台本を生成して出力する。入力設定情報取得部は、言語モデルMから出力された台本を取得する。
例えば、台本生成部102は、言語モデルMに入力セッション情報を入力することによって、台本を生成する。言語モデルMは、入力セッション情報をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルMは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、台本を生成物として生成する。例えば、言語モデルMは、デフォルトプロンプトに基づいて、入力セッション情報に応じた台本を生成して出力する。入力セッション情報取得部は、言語モデルMから出力された台本を取得する。
例えば、台本生成部102は、言語モデルMに、詳細度に応じた入力セッション情報を入力することによって、詳細度に応じた台本を生成する。言語モデルMは、詳細度に応じた入力セッション情報をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルMは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、台本を生成物として生成する。例えば、言語モデルMは、デフォルトプロンプトに基づいて、詳細度に応じた入力セッション情報に応じた台本を生成して出力する。入力セッション情報取得部は、言語モデルMから出力された台本を取得する。
例えば、台本生成部102は、セッションごとに、当該セッションの入力セッション情報を言語モデルMに入力することによって、台本の一部である台本部分であって、当該セッションの台本部分を生成し、複数のセッションのうち、2番目以降のセッションについては、当該2番目以降のセッションよりも前のセッションの台本部分も言語モデルMに入力することによって、当該2番目以降のセッションの台本部分を生成し、複数のセッションの各々の台本部分に基づいて、台本を生成する。このように、個々のセッションごとに台本部分を生成すべきことがデフォルトプロンプトに示されていてもよい。
例えば、台本生成部102は、言語モデルMに第1入力情報及び第2入力情報を入力することによって、台本を生成する。言語モデルMは、第1入力情報及び第2入力情報をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルMは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、台本を生成物として生成する。例えば、言語モデルMは、デフォルトプロンプトに基づいて、第1入力情報及び第2入力情報に応じた台本を生成して出力する。入力セッション情報取得部は、言語モデルMから出力された台本を取得する。
なお、台本生成部102は、セッションごとに台本部分を生成するのではなく、全てのセッションを含む台本を一度に生成してもよい。この場合、台本を一度に生成すべきことがデフォルトプロンプトに示されていてもよい。また、台本は、特に複数のセッションに分けられなくてもよい。この場合、入力セッション情報が取得されない。台本生成部102は、入力特徴情報だけに基づいて、台本を生成してもよい。台本生成部102は、入力設定情報だけに基づいて、台本を生成してもよい。台本生成部102は、入力セッション情報だけに基づいて、台本を生成してもよい。台本生成部102は、入力情報に基づいて、台本を生成すればよい。台本生成部102は、特に詳細度に関係なく、台本を生成してもよい。
[4.台本生成システムで実行される処理]
図7は、台本生成システム1で実行される処理の一例を示す図である。図7では、台本生成システム1で実行される処理のうち、台本生成装置10の処理が示されている。制御部11が、記憶部12に記憶されたプログラムを実行することによって、図7の処理が実行される。図7の各ステップは、本開示に係る台本生成方法に含まれるステップの一例である。例えば、台本生成サービスで台本の生成を担当する担当者が、台本生成装置10を操作して、台本の生成対象となるライブ配信を指定すると、図7の処理が実行される。
図7は、台本生成システム1で実行される処理の一例を示す図である。図7では、台本生成システム1で実行される処理のうち、台本生成装置10の処理が示されている。制御部11が、記憶部12に記憶されたプログラムを実行することによって、図7の処理が実行される。図7の各ステップは、本開示に係る台本生成方法に含まれるステップの一例である。例えば、台本生成サービスで台本の生成を担当する担当者が、台本生成装置10を操作して、台本の生成対象となるライブ配信を指定すると、図7の処理が実行される。
図7のように、台本生成装置10は、台本データベースDBから、台本の生成対象となるライブ配信のライブ配信IDに関連付けられたオリエンテーションシートを取得する(S1)。台本生成装置10は、オリエンテーションシートに基づいて、入力特徴情報を取得する(S2)。S2では、台本生成装置10は、オリエンテーションシートのうち、商品又はサービスの特徴を示す項目に入力された情報を、入力特徴情報として取得する。
台本生成装置10は、S2で取得された入力特徴情報に基づいて、ライブ配信で紹介される商品又はサービスのコンテンツを検索する(S3)。S3では、台本生成装置10は、入力特徴情報に示された商品名又はサービス名等の情報を検索クエリにして、公知のウェブ検索サービスを検索する。台本生成装置10は、S3で検索されたコンテンツから抽出情報を取得する(S4)。台本生成装置10は、入力特徴情報を生成することを示すデフォルトプロンプトと、S4で抽出された抽出情報と、を言語モデルMに入力することによって、入力特徴情報を取得する(S5)。S5では、言語モデルMは、これらをトークンに分割し、トークンの埋め込み表現の並びに基づいて、入力特徴情報を出力する。台本生成装置10は、言語モデルMから出力された入力特徴情報を取得する。
なお、十分な量の入力特徴情報がオリエンテーションシートに含まれる場合には、S3~S5の処理は実行されなくてもよい。例えば、台本の生成を担当する担当者が、操作部14を操作することによって、S3~S5の処理の実行要否を指定してもよい。台本生成装置10は、オリエンテーションシートのうち、商品又はサービスの特徴を示す全ての項目又は所定数以上の項目に情報が入力されている場合には、S3~S5の処理を実行しなくてもよい。
台本生成装置10は、オリエンテーションシートに基づいて、入力設定情報を取得する(S6)。S6では、オリエンテーションシートに入力設定情報が含まれている場合には、台本生成装置10は、オリエンテーションシートに含まれる入力設定情報を取得する。オリエンテーションシートに入力設定情報が含まれていない、又は、オリエンテーションシートに十分な量の入力設定情報が含まれていない場合には、台本生成装置10は、入力設定情報を生成することを示すデフォルトプロンプトと、オリエンテーションシートに含まれる情報と、を言語モデルMに入力する。言語モデルMは、これらをトークンに分割し、トークンの埋め込み表現の並びに基づいて、入力設定情報を出力する。台本生成装置10は、言語モデルMから出力された入力設定情報を取得する。S6では、台本生成装置10は、ライブ配信の要約等の情報も、入力設定情報として取得してもよい。
台本生成装置10は、オリエンテーションシートに基づいて、入力セッション情報を取得する(S7)。S7では、オリエンテーションシートに入力セッション情報が含まれている場合には、台本生成装置10は、オリエンテーションシートに含まれる入力セッション情報を取得する。オリエンテーションシートに入力セッション情報が含まれていない、又は、オリエンテーションシートに十分な量の入力セッション情報が含まれていない場合には、台本生成装置10は、入力セッション情報を生成することを示すデフォルトプロンプトと、オリエンテーションシートに含まれる情報と、を言語モデルMに入力する。言語モデルMは、これらをトークンに分割し、トークンの埋め込み表現の並びに基づいて、入力セッション情報を出力する。台本生成装置10は、言語モデルMから出力された入力セッション情報を取得する。
台本生成装置10は、オリエンテーションシートに基づいて、簡易版又は詳細版の何れの台本を生成するかを判定する(S8)。S8では、台本生成装置10は、オリエンテーションシートに含まれる台本の詳細度が簡易版又は詳細版の何れを示すかを判定する。オリエンテーションシートに台本の詳細度が含まれない場合には、台本の生成を担当する担当者が、操作部14を操作することによって、台本の詳細度を指定してもよい。台本生成装置10は、担当者が指定した台本の詳細度に基づいて、S8の判定を行ってもよい。
S8において、簡易版の台本を生成すると判定された場合(S8:簡易版)、台本生成装置10は、入力特徴情報、入力設定情報、及び入力セッション情報を含む入力情報を言語モデルMに入力することによって、最初のセッションの台本部分を生成する(S9)。S9では、台本生成装置10は、セッションに応じた台本部分を生成することを示すデフォルトプロンプトも言語モデルMに入力する。言語モデルMは、入力情報等をトークンに分割し、トークンの埋め込み表現の並びに基づいて、最初のセッションの台本部分を出力する。最初のセッションよりも前には他のセッションが存在しないので、言語モデルMは、他のセッションの台本部分には基づかずに、最初のセッションの台本部分を出力する。台本生成装置10は、言語モデルMから出力された最初のセッションの台本部分を取得する。
台本生成装置10は、入力特徴情報、入力設定情報、及び入力セッション情報を含む入力情報と、生成済みのセッションの台本部分と、を言語モデルMに入力することによって、次のセッションの台本部分を生成する(S10)。S10では、台本生成装置10は、次のセッションに応じた台本部分を生成することを示すデフォルトプロンプトも言語モデルMに入力する。言語モデルMは、入力情報及び台本部分をトークンに分割し、トークンの埋め込み表現の並びに基づいて、次のセッションの台本部分を出力する。例えば、台本生成装置10は、それまでに生成した全てのセッションの台本部分を言語モデルMに入力する。台本生成装置10は、言語モデルMから出力された次のセッションの台本部分を取得する。
台本生成装置10は、入力セッション情報に基づいて、最後のセッションの台本部分まで生成したか否かを判定する(S11)。S11において、最後のセッションの台本部分まで生成していないと判定された場合(S11:N)、S10の処理に戻り、次のセッションの台本部分が生成される。最後のセッションの台本部分まで生成したと判定された場合(S11:Y)、台本生成装置10は、各セッションの台本部分に基づいて、最終的な台本を生成し(S12)、本処理は、終了する。S12では、台本生成装置10は、各セッションの台本部分をつなぎ合わせることによって、最終的な台本を生成する。台本生成装置10は、台本データベースDBに最終的な台本を格納する。
S8において、詳細版の台本が必要であると判定された場合(S8:詳細版)、台本生成装置10は、入力セッション情報が示すセッションごとにサブセッションを生成することによって、複数のセッションの各々のサブセッションを示す入力セッション情報を取得する(S13)。S13では、台本生成装置10は、セッションごとにサブセッションを生成することを示すデフォルトプロンプトと、入力セッション情報と、を言語モデルMに入力する。言語モデルMは、これらをトークンに分割し、トークンの埋め込み表現の並びに基づいて、サブセッションを示す入力セッション情報を出力する。台本生成装置10は、言語モデルMから出力された入力セッション情報を取得する。
台本生成装置10は、処理対象のセッションについて、入力特徴情報、入力設定情報、及び入力セッション情報を含む入力情報を言語モデルMに入力することによって、処理対象のセッションの最初のサブセッションの台本部分を生成する(S14)。処理対象のセッションは、S14~S17のループの対象となるセッションである。最初のセッションから順番に、処理対象のセッションが選択される。S14では、台本生成装置10は、サブセッションに応じた台本部分を生成することを示すデフォルトプロンプトも言語モデルMに入力する。言語モデルMは、入力情報等をトークンに分割し、トークンの埋め込み表現の並びに基づいて、最初のセッションのサブセッションの台本部分を出力する。最初のセッションのサブセッションよりも前には他のサブセッションが存在しないので、言語モデルMは、他のサブセッションの台本部分には基づかずに、最初のサブセッションの台本部分を出力する。台本生成装置10は、言語モデルMから出力された最初のサブセッションの台本部分を取得する。
台本生成装置10は、処理対象のセッションについて、入力特徴情報、入力設定情報、及び入力セッション情報を含む入力情報と、作成済みのサブセッションの台本部分と、に基づいて、処理対象のセッションの次のサブセッションの台本部分を生成する(S15)。S15では、台本生成装置10は、次のサブセッションに応じた台本部分を生成することを示すデフォルトプロンプトも言語モデルMに入力する。言語モデルMは、入力情報及び台本部分をトークンに分割し、トークンの埋め込み表現の並びに基づいて、次のサブセッションの台本部分を出力する。例えば、台本生成装置10は、それまでに生成した全てのサブセッションの台本部分を言語モデルMに入力する。台本生成装置10は、言語モデルMから出力された次のサブセッションの台本部分を取得する。
台本生成装置10は、入力セッション情報に基づいて、処理対象のセッションの最後のサブセッションの台本部分まで生成したか否かを判定する(S16)。処理対象のセッションの最後のサブセッションの台本部分まで生成したと判定されない場合(S16:N)、S15の処理に戻り、処理対象のセッションの次のサブセッションの台本部分が生成される。処理対象のセッションの最後のサブセッションの台本部分まで生成したと判定された場合(S16:Y)、台本生成装置10は、入力セッション情報に基づいて、最後のセッションの台本部分まで生成したか否かを判定する(S17)。
S17において、最後のセッションの台本部分まで生成していないと判定された場合(S17:N)、S14の処理に戻り、次のセッションの最初のサブセッションの台本部分が生成される。即ち、次のセッションが処理対象のセッションになる。最後のセッションの台本部分まで生成したと判定された場合(S17:Y)、台本生成装置10は、各セッションの各サブセッションの台本部分に基づいて、最終的な台本を生成し(S18)、本処理は、終了する。S18では、台本生成装置10は、各セッションの各サブセッションの台本部分をつなぎ合わせることによって、最終的な台本を生成する。台本生成装置10は、台本データベースDBに最終的な台本を格納する。
[5.台本生成システムのまとめ]
本実施形態の台本生成システム1は、商品又はサービスが紹介される動画に関する入力情報を取得する。台本生成システム1は、言語モデルMに入力情報を入力することによって、動画に関する台本を生成する。台本生成システム1は、言語モデルMを利用して、入力情報に応じた柔軟な台本を生成できる。例えば、台本生成サービスの担当者がテンプレートに基づいて台本を生成すると、テンプレートの範囲内でしか台本を生成できないので、柔軟な台本を生成できないが、言語モデルMは、入力情報に応じた柔軟な自然言語処理を実行することによって、より柔軟な台本を生成できる。台本生成システム1は、担当者の手間を軽減することもできるので、台本の生成にかかるコストを抑制できる。例えば、台本生成サービスの担当者が、外部の業者に台本の生成を委託すると、外部の業者へのコストが発生するが、台本生成システム1は、このようなコストの発生を回避できる。
本実施形態の台本生成システム1は、商品又はサービスが紹介される動画に関する入力情報を取得する。台本生成システム1は、言語モデルMに入力情報を入力することによって、動画に関する台本を生成する。台本生成システム1は、言語モデルMを利用して、入力情報に応じた柔軟な台本を生成できる。例えば、台本生成サービスの担当者がテンプレートに基づいて台本を生成すると、テンプレートの範囲内でしか台本を生成できないので、柔軟な台本を生成できないが、言語モデルMは、入力情報に応じた柔軟な自然言語処理を実行することによって、より柔軟な台本を生成できる。台本生成システム1は、担当者の手間を軽減することもできるので、台本の生成にかかるコストを抑制できる。例えば、台本生成サービスの担当者が、外部の業者に台本の生成を委託すると、外部の業者へのコストが発生するが、台本生成システム1は、このようなコストの発生を回避できる。
また、台本生成システム1は、言語モデルMに入力特徴情報を入力することによって、台本を生成する。これにより、言語モデルMが商品又はサービスの特徴に応じた柔軟な台本を生成できるようになるので、台本生成システム1は、言語モデルMの生成物である台本の精度を高めることができる。例えば、入力特徴情報が見た目の特徴を示す場合、台本生成システム1は、見た目の特徴に応じた台本を生成できる。入力特徴情報が機能的な特徴を示す場合、台本生成システム1は、機能的な特徴に応じた台本を生成できる。
また、台本生成システム1は、商品又はサービスの特徴に関するコンテンツから抽出された抽出情報が言語モデルM又は他の言語モデルに入力されることによって生成された入力特徴情報を取得する。これにより、台本生成システム1は、より多くの特徴から台本を生成できるので、言語モデルMの生成物である台本の精度を高めることができる。例えば、オリエンテーションシートに入力特徴情報がない、又は、オリエンテーションシートに十分な量の入力特徴情報がなかったとしても、台本生成システム1は、コンテンツから入力特徴情報を取得できる。
また、台本生成システム1は、言語モデルMに入力設定情報を入力することによって、台本を生成する。言語モデルMが動画の設定に応じた柔軟な台本を生成できるようになるので、台本生成システム1は、言語モデルMの生成物である台本の精度を高めることができる。例えば、入力設定情報が動画の配信日時を示す場合、台本生成システム1は、動画の配信日時に応じた分量の台本を生成できる。入力設定情報が動画の長さを示す場合、台本生成システム1は、動画の長さに応じた分量の台本を生成できる。
また、台本生成システム1は、言語モデルMに入力セッション情報を入力することによって、台本を生成する。言語モデルMが動画のセッションに応じた柔軟な台本を生成できるようになるので、台本生成システム1は、言語モデルMの生成物である台本の精度を高めることができる。例えば、言語モデルMが一度に大量の生成物をしようとすると、生成物の精度が落ちることがある。この点、言語モデルMがセッションごとに台本を生成すると、言語モデルMが一度に生成する生成物の量を抑えることができるので、台本生成システム1は、台本の精度を高めることができる。
また、台本生成システム1は、言語モデルMに、詳細度に応じた入力セッション情報を入力することによって、詳細度に応じた台本を生成する。言語モデルMが詳細度に応じた柔軟な台本を生成できるようになるので、台本生成システム1は、言語モデルMの生成物である台本の精度を高めることができる。例えば、台本生成システム1は、ライブ配信の全体的な進行を示す簡易版の台本で十分な店舗のために、店舗が要求する簡易版の台本を生成できる。この場合、言語モデルMが生成するテキストの量を抑えることができるので、台本生成装置10の処理負荷を軽減することができる。台本生成装置10以外の他のコンピュータに言語モデルMが記憶される場合には、他のコンピュータの処理負荷を軽減することができる。例えば、台本生成システム1は、ライブ配信の詳細な進行を示す詳細版の台本を希望する店舗のために、店舗が要求する詳細版の台本を生成することもできる。
また、台本生成システム1は、セッションごとに、当該セッションの入力セッション情報を言語モデルMに入力することによって、当該セッションの台本部分を生成する。台本生成システム1は、複数のセッションのうち、2番目以降のセッションについては、当該2番目以降のセッションよりも前のセッションの台本部分も言語モデルMに入力することによって、当該2番目以降のセッションの台本部分を生成する。台本生成システム1は、複数のセッションの各々の台本部分に基づいて、台本を生成する。これにより、台本生成システム1は、セッション間の繋がりがある台本を生成できる。例えば、言語モデルMが一度の大量のテキストを生成しようとすると、台本生成装置10の処理負荷が高まる可能性があるが、言語モデルMが細切れの台本部分をセッションごとに生成することによって、台本生成装置10の処理負荷を軽減することができる。更に、言語モデルMが一度の大量のテキストを生成しようとすると、生成物の精度が低下する可能性があるが、言語モデルMが細切れの台本部分をセッションごとに生成することによって、台本生成システム1は、最終的な生成物である台本の精度を高めることができる。
また、台本生成システム1は、予め用意された第1入力情報と、言語モデルM又は他の言語モデルに第1入力情報が入力されることによって生成された第2入力情報と、を取得する。台本生成システム1は、言語モデルMに第1入力情報及び第2入力情報を入力することによって、台本を生成する。これにより、台本生成システム1は、バリエーション豊かな入力情報を言語モデルMに入力できるので、台本の精度を高めることができる。
[6.変形例]
本開示は、以上に説明した実施形態に限定されない。本開示は、本開示の趣旨を逸脱しない範囲で、適宜変更可能である。
本開示は、以上に説明した実施形態に限定されない。本開示は、本開示の趣旨を逸脱しない範囲で、適宜変更可能である。
図8は、変形例で実現される機能の一例を示す図である。例えば、変形例の台本生成装置10は、質問回答生成部103、第1サンプル取得部104、及び第2サンプル取得部105が実現される。質問回答生成部103、第1サンプル取得部104、及び第2サンプル取得部105の各々は、制御部11により実現される。
[6-1.変形例1]
例えば、図2の紹介画面SCの例では、ライブ配信サービスは、視聴者からのコメントを受け付ける。視聴者は、コメントとして、出演者が紹介する商品又はサービスに関する質問を入力することがある。出演者は、ライブ配信の中で、視聴者からの質問に回答することがある。変形例1では、出演者が、視聴者から想定される質問に備えることができるように、台本生成部102が生成した台本の内容から、視聴者から想定される質問とそれに対する回答が生成されるようになっている。
例えば、図2の紹介画面SCの例では、ライブ配信サービスは、視聴者からのコメントを受け付ける。視聴者は、コメントとして、出演者が紹介する商品又はサービスに関する質問を入力することがある。出演者は、ライブ配信の中で、視聴者からの質問に回答することがある。変形例1では、出演者が、視聴者から想定される質問に備えることができるように、台本生成部102が生成した台本の内容から、視聴者から想定される質問とそれに対する回答が生成されるようになっている。
変形例1の台本生成システム1は、質問回答生成部103を含む。質問回答生成部103は、言語モデルM又は他の言語モデルに台本を入力することによって、動画に関する質問と回答を生成する。質問回答生成部103は、台本以外の他の情報を、言語モデルM又は他の言語モデルに入力してもよい。例えば、質問回答生成部103は、入力特徴情報、入力設定情報、及び入力セッション情報のうちの少なくとも1つを、言語モデルM又は他の言語モデルに入力してもよい。
質問と回答は、少なくとも1つの質問と、当該質問に対する少なくとも1つの回答と、を含むデータである。質問と回答は、任意のデータ形式であってよい。例えば、質問と回答は、表計算ソフトのデータ形式、CSV形式、テキスト形式、文書形式、又は他の形式であってもよい。質問回答生成部103は、任意の数の質問と回答を生成可能である。質問回答生成部103が生成すべき質問と回答の数は、予め指定されていてもよいし、特に指定されていなくてもよい。
例えば、言語モデルM又は他の言語モデルが、質問と回答の生成に特化したモデルではない場合には、質問回答生成部103は、言語モデルM又は他の言語モデルに対し、質問と回答を生成すべきことを示すデフォルトプロンプトを入力してもよい。デフォルトプロンプトは、「あなたは、あなたに入力された台本から質問と回答を生成するAIです。」といったように、言語モデルM又は他の言語モデルが実行すべき処理内容(言語モデルM若しくは他の言語モデルの役割、又は、言語モデルM若しくは他の言語モデルが生成すべき生成物の種類)を示してもよい。
例えば、言語モデルM又は他の言語モデルは、自身に入力された台本をトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現(特徴ベクトル)を計算する。言語モデルM又は他の言語モデルは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、質問と回答を生成物として生成する。例えば、言語モデルM又は他の言語モデルは、デフォルトプロンプトに基づいて、自身に入力された台本に応じた質問と回答を生成して出力する。質問回答生成部103は、言語モデルM又は他の言語モデルから出力された質問と回答を取得する。
例えば、質問回答生成部103は、台本データベースDBに、質問と回答を格納する。サーバ20は、出演者装置30に対し、台本データベースDBに格納された質問と回答を送信する。質問回答生成部103は、台本データベースDB以外の他のデータベースに質問と回答を格納してもよい。質問回答生成部103は、台本生成装置10以外の他のコンピュータ、又は、外部情報記憶媒体に、質問と回答を記録してもよい。質問回答生成部103が生成した質問と回答は、出演者装置30以外の他のコンピュータに送信されてもよい。
変形例1の台本生成システム1は、言語モデルM又は他の言語モデルに台本を入力することによって、動画に関する質問と回答を生成する。台本生成システム1は、台本に応じた質問と回答を生成することによって、出演者の進行を支援できる。
[6-2.変形例2]
例えば、実施形態では、言語モデルMに対し、入力特徴情報を一例とする入力情報が入力される場合を例に挙げた。言語モデルMには、台本の生成に利用可能な任意の情報が入力されてよい。変形例2では、言語モデルMに対し、台本のサンプルが入力される場合を例に挙げる。サンプルは、言語モデルMの参考用の台本のデータである。台本生成部102により生成される台本と同様に、サンプルは、任意のデータ形式であってよい。例えば、サンプルは、台本のフォーマット又はテンプレートであってもよい。サンプルは、人手で生成されてもよいし、言語モデルMによって生成されてもよい。サンプルは、動画の音声が文字起こしされたものであってもよい。
例えば、実施形態では、言語モデルMに対し、入力特徴情報を一例とする入力情報が入力される場合を例に挙げた。言語モデルMには、台本の生成に利用可能な任意の情報が入力されてよい。変形例2では、言語モデルMに対し、台本のサンプルが入力される場合を例に挙げる。サンプルは、言語モデルMの参考用の台本のデータである。台本生成部102により生成される台本と同様に、サンプルは、任意のデータ形式であってよい。例えば、サンプルは、台本のフォーマット又はテンプレートであってもよい。サンプルは、人手で生成されてもよいし、言語モデルMによって生成されてもよい。サンプルは、動画の音声が文字起こしされたものであってもよい。
変形例2の台本生成システム1は、第1サンプル取得部104を含む。第1サンプル取得部104は、台本の生成対象となる動画で紹介される商品とは異なる他の商品、又は、台本の生成対象となる動画で紹介されるサービスとは異なる他のサービスが紹介される他の動画に関するサンプルであって、他の商品又は他のサービスの特徴が省かれたサンプルを取得する。他の動画は、台本の生成対象となる動画以外の動画である。他の動画は、過去に配信済みの動画であってもよいし、台本が生成されているがまだ配信されていない動画であってもよい。他の動画は、ライブ配信サービス以外の他のサービスで配信された動画であってもよい。他の商品又は他のサービスは、他の動画で紹介される商品又はサービスである。
変形例2では、データ記憶部100がサンプルを記憶する場合を例に挙げる。第1サンプル取得部104は、データ記憶部100からサンプルを取得する。サンプルは、台本生成装置10以外の他のコンピュータ、又は、外部情報記憶媒体に記憶されていてもよい。この場合、第1サンプル取得部104は、他のコンピュータ又は外部情報記憶媒体から、サンプルを取得すればよい。
例えば、第1サンプル取得部104は、台本生成部102により過去に生成された台本のうち、当該台本の生成時に利用された入力特徴情報に含まれる単語の部分がマスクされたサンプルを取得する。マスクは、特定の部分を隠す(例えば、当該部分をスペース又は特定の記号で埋める)ことである。第1サンプル取得部104は、当該入力特徴情報に基づいて、当該台本に対してマスクを実行することによってサンプルを取得してもよい。第1サンプル取得部104は、第1サンプル取得部104以外の他の機能によりマスクが完了したサンプルを取得してもよい。
なお、他の商品又は他のサービスの特徴を省く方法は、マスクに限られない。例えば、第1サンプル取得部104は、台本生成部102により過去に生成された台本のうち、当該台本の生成時に利用された入力特徴情報に含まれる単語の部分を削除することによって、サンプルを取得してもよい。第1サンプル取得部104は、第1サンプル取得部104以外の他の機能により当該部分の削除が行われたサンプルを取得してもよい。
変形例2の台本生成部102は、言語モデルMにサンプルを更に入力することによって、台本を生成する。即ち、台本生成部102は、実施形態で説明した入力情報と、サンプルと、を言語モデルMに入力する。言語モデルMは、入力情報及びサンプルをトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルMは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、サンプルに応じた台本を生成物として生成する。例えば、言語モデルMは、デフォルトプロンプトに基づいて、サンプルに応じた台本を生成して出力してもよい。デフォルトプロンプトは、「このサンプルを参考にして台本を生成して下さい。」といったように、言語モデルMがサンプルを参考にすべきことを示すテキストを含んでもよい。台本生成部102は、言語モデルMから出力された台本を取得する。
変形例2の台本生成システム1は、言語モデルMに、他の商品又は他のサービスの特徴が省かれたサンプルを更に入力することによって、台本を生成する。これにより、言語モデルMが、他の商品又は他のサービスの特徴の影響を受けにくくなるので、台本生成システム1は、台本の生成対象となる動画で紹介される商品又はサービスの特徴が反映された台本を生成できる。即ち、台本生成システム1は、台本の精度を高めることができる。
[6-3.変形例3]
例えば、出演者の話の流れや話し方には、出演者に特有の特徴が存在することがある。自身の特徴に沿った台本が生成されると、出演者は、ライブ配信を進行しやすいと考えられる。このため、台本の生成対象となる動画の出演者と同じ出演者が担当する他の動画の台本がサンプルとして取得されてもよい。変形例3では、台本データベースDBに、他の動画の台本と、当該他の動画の出演者と、が格納されているものとする。
例えば、出演者の話の流れや話し方には、出演者に特有の特徴が存在することがある。自身の特徴に沿った台本が生成されると、出演者は、ライブ配信を進行しやすいと考えられる。このため、台本の生成対象となる動画の出演者と同じ出演者が担当する他の動画の台本がサンプルとして取得されてもよい。変形例3では、台本データベースDBに、他の動画の台本と、当該他の動画の出演者と、が格納されているものとする。
変形例3の台本生成システム1は、第2サンプル取得部105を含む。第2サンプル取得部105は、動画の出演者と同じ出演者が出演した他の動画に関するサンプルを取得する。サンプルの意味は、変形例2と同様である。例えば、第2サンプル取得部105は、台本データベースDBに格納されたオリエンテーションシートに基づいて、台本の生成対象となる動画の出演者を特定する。第2サンプル取得部105は、台本データベースDBに格納されたオリエンテーションシートに基づいて、当該特定された出演者と同じ出演者の他の動画を特定する。第2サンプル取得部105は、当該特定された同じ出演者の他の動画の台本を、サンプルとして取得する。
なお、同じ出演者の他の動画が複数存在する場合には、第2サンプル取得部105は、同じ出演者の全ての他の動画の台本をサンプルとして取得してもよいし、同じ出演者の一部の他の動画の台本をサンプルとして取得してもよい。また、変形例2,3を組み合わせる場合には、第2サンプル取得部105は、台本の生成対象となる動画と同じ出演者の他の動画の台本のうち、他の商品又は他のサービスの情報が省かれた台本をサンプルとして取得してもよい。第2サンプル取得部105は、ライブ配信サービス以外の他のサービスで配信された同じ出演者の動画の台本をサンプルとして取得してもよい。
変形例3の台本生成部102は、言語モデルMにサンプルを更に入力することによって、台本を生成する。即ち、台本生成部102は、実施形態で説明した入力情報と、サンプルと、を言語モデルMに入力する。言語モデルMは、入力情報及びサンプルをトークンに分割し、学習で調整されたパラメータに基づいて、個々のトークンの埋め込み表現を計算する。言語モデルMは、トークンの埋め込み表現の並び順に基づいて、必要に応じて続きのテキストを予測したうえで、サンプルに応じた台本を生成物として生成する。例えば、言語モデルMは、変形例2と同様のデフォルトプロンプトに基づいて、サンプルに応じた台本を生成して出力してもよい。台本生成部102は、言語モデルMから出力された台本を取得する。
変形例3の台本生成システム1は、言語モデルMに、同じ出演者が出演した他の動画に関するサンプルを更に入力することによって、台本を生成する。これにより、台本生成システム1は、出演者の特徴を台本に反映させることができる。即ち、台本生成システム1は、台本の精度を高めることができる。
[6-4.その他の変形例]
例えば、上記変形例を組み合わせてもよい。
例えば、上記変形例を組み合わせてもよい。
例えば、ライブ配信サービスの中で提供される台本生成サービスに台本生成システム1が適用される場合を例に挙げたが、台本生成システム1は、他の場面にも適用可能である。台本生成システム1は、ライブ配信サービス以外の他の動画の台本を生成してもよい。例えば、台本生成システム1は、リアルタイムでは配信されない動画の台本、商品又はサービスの広告を示す動画の台本、人間ではなくAI(人工音声)が話をする動画の台本、電子商取引サービス以外の他のサービスで販売される商品若しくはサービスが紹介される動画の台本、又は他の動画の台本を生成してもよい。
例えば、台本生成システム1は、電子商取引サービス以外の他のサービスの商品又はサービスが紹介される動画の台本を生成してもよい。台本生成システム1は、実店舗で販売される商品若しくはサービス、ホテル等の宿泊施設で提供されるサービス、通信サービス、決済サービス、金融サービス、電子書籍サービス、美容院若しくはレストランで提供されるサービス、又はSNSで紹介される商品若しくはサービスが紹介される動画の台本を生成してもよい。動画で紹介される商品は、有体物に限られず、楽曲又は映画等のコンテンツであってもよい。
例えば、台本生成装置10で実現されるものとして説明した機能は、サーバ20、出演者装置30、又は他のコンピュータで実現されてもよい。台本生成装置10で実現されるものとして説明した処理は、複数のコンピュータで分担されてもよい。
[7.付記]
例えば、本開示に係る台本生成システムは、下記のような構成も可能である。
例えば、本開示に係る台本生成システムは、下記のような構成も可能である。
(1)
自然言語で記述された生成物を生成可能な学習済みの言語モデルに入力される入力情報であって、商品又はサービスが紹介される動画に関する前記入力情報を取得する入力情報取得部と、
前記言語モデルに前記入力情報を入力することによって、前記動画に関する台本を生成する台本生成部と、
を含む台本生成システム。
(2)
前記入力情報取得部は、前記商品又は前記サービスの特徴に関する前記入力情報である入力特徴情報を取得し、
前記台本生成部は、前記言語モデルに前記入力特徴情報を入力することによって、前記台本を生成する、
(1)に記載の台本生成システム。
(3)
前記入力情報取得部は、前記商品又は前記サービスの特徴に関するコンテンツから抽出された抽出情報が前記言語モデル又は他の言語モデルに入力されることによって生成された前記入力特徴情報を取得する、
(2)に記載の台本生成システム。
(4)
前記入力情報取得部は、前記動画の設定に関する前記入力情報である入力設定情報を取得し、
前記台本生成部は、前記言語モデルに前記入力設定情報を入力することによって、前記台本を生成する、
(1)~(3)の何れかに記載の台本生成システム。
(5)
前記入力情報取得部は、前記動画における複数のセッションの各々に関する前記入力情報である入力セッション情報を取得し、
前記台本生成部は、前記言語モデルに前記入力セッション情報を入力することによって、前記台本を生成する、
(1)~(4)の何れかに記載の台本生成システム。
(6)
前記入力情報取得部は、前記台本に関する詳細度に応じた前記入力セッション情報を取得し、
前記台本生成部は、前記言語モデルに、前記詳細度に応じた前記入力セッション情報を入力することによって、前記詳細度に応じた前記台本を生成する、
(5)に記載の台本生成システム。
(7)
前記台本生成部は、
前記セッションごとに、当該セッションの前記入力セッション情報を前記言語モデルに入力することによって、前記台本の一部である台本部分であって、当該セッションの前記台本部分を生成し、
前記複数のセッションのうち、2番目以降の前記セッションについては、当該2番目以降のセッションよりも前の前記セッションの前記台本部分も前記言語モデルに入力することによって、当該2番目以降のセッションの前記台本部分を生成し、
前記複数のセッションの各々の前記台本部分に基づいて、前記台本を生成する、
(5)又は(6)に記載の台本生成システム。
(8)
前記入力情報取得部は、予め用意された前記入力情報である第1入力情報と、前記言語モデル又は他の言語モデルに前記第1入力情報が入力されることによって生成された前記入力情報である第2入力情報と、を取得し、
前記台本生成部は、前記言語モデルに前記第1入力情報及び前記第2入力情報を入力することによって、前記台本を生成する、
(1)~(7)の何れかに記載の台本生成システム。
(9)
前記台本生成システムは、前記言語モデル又は他の言語モデルに前記台本を入力することによって、前記動画に関する質問と回答を生成する質問回答生成部を更に含む、
(1)~(8)の何れかに記載の台本生成システム。
(10)
前記台本生成システムは、前記商品とは異なる他の商品、又は、前記サービスとは異なる他のサービスが紹介される他の動画に関するサンプルであって、前記他の商品又は前記他のサービスの特徴が省かれた前記サンプルを取得する第1サンプル取得部を更に含み、
前記台本生成部は、前記言語モデルに前記サンプルを更に入力することによって、前記台本を生成する、
(1)~(9)の何れかに記載の台本生成システム。
(11)
前記台本生成システムは、前記動画の出演者と同じ出演者が出演した他の動画に関するサンプルを取得する第2サンプル取得部を更に含み、
前記台本生成部は、前記言語モデルに前記サンプルを更に入力することによって、前記台本を生成する、
(1)~(10)の何れかに記載の台本生成システム。
自然言語で記述された生成物を生成可能な学習済みの言語モデルに入力される入力情報であって、商品又はサービスが紹介される動画に関する前記入力情報を取得する入力情報取得部と、
前記言語モデルに前記入力情報を入力することによって、前記動画に関する台本を生成する台本生成部と、
を含む台本生成システム。
(2)
前記入力情報取得部は、前記商品又は前記サービスの特徴に関する前記入力情報である入力特徴情報を取得し、
前記台本生成部は、前記言語モデルに前記入力特徴情報を入力することによって、前記台本を生成する、
(1)に記載の台本生成システム。
(3)
前記入力情報取得部は、前記商品又は前記サービスの特徴に関するコンテンツから抽出された抽出情報が前記言語モデル又は他の言語モデルに入力されることによって生成された前記入力特徴情報を取得する、
(2)に記載の台本生成システム。
(4)
前記入力情報取得部は、前記動画の設定に関する前記入力情報である入力設定情報を取得し、
前記台本生成部は、前記言語モデルに前記入力設定情報を入力することによって、前記台本を生成する、
(1)~(3)の何れかに記載の台本生成システム。
(5)
前記入力情報取得部は、前記動画における複数のセッションの各々に関する前記入力情報である入力セッション情報を取得し、
前記台本生成部は、前記言語モデルに前記入力セッション情報を入力することによって、前記台本を生成する、
(1)~(4)の何れかに記載の台本生成システム。
(6)
前記入力情報取得部は、前記台本に関する詳細度に応じた前記入力セッション情報を取得し、
前記台本生成部は、前記言語モデルに、前記詳細度に応じた前記入力セッション情報を入力することによって、前記詳細度に応じた前記台本を生成する、
(5)に記載の台本生成システム。
(7)
前記台本生成部は、
前記セッションごとに、当該セッションの前記入力セッション情報を前記言語モデルに入力することによって、前記台本の一部である台本部分であって、当該セッションの前記台本部分を生成し、
前記複数のセッションのうち、2番目以降の前記セッションについては、当該2番目以降のセッションよりも前の前記セッションの前記台本部分も前記言語モデルに入力することによって、当該2番目以降のセッションの前記台本部分を生成し、
前記複数のセッションの各々の前記台本部分に基づいて、前記台本を生成する、
(5)又は(6)に記載の台本生成システム。
(8)
前記入力情報取得部は、予め用意された前記入力情報である第1入力情報と、前記言語モデル又は他の言語モデルに前記第1入力情報が入力されることによって生成された前記入力情報である第2入力情報と、を取得し、
前記台本生成部は、前記言語モデルに前記第1入力情報及び前記第2入力情報を入力することによって、前記台本を生成する、
(1)~(7)の何れかに記載の台本生成システム。
(9)
前記台本生成システムは、前記言語モデル又は他の言語モデルに前記台本を入力することによって、前記動画に関する質問と回答を生成する質問回答生成部を更に含む、
(1)~(8)の何れかに記載の台本生成システム。
(10)
前記台本生成システムは、前記商品とは異なる他の商品、又は、前記サービスとは異なる他のサービスが紹介される他の動画に関するサンプルであって、前記他の商品又は前記他のサービスの特徴が省かれた前記サンプルを取得する第1サンプル取得部を更に含み、
前記台本生成部は、前記言語モデルに前記サンプルを更に入力することによって、前記台本を生成する、
(1)~(9)の何れかに記載の台本生成システム。
(11)
前記台本生成システムは、前記動画の出演者と同じ出演者が出演した他の動画に関するサンプルを取得する第2サンプル取得部を更に含み、
前記台本生成部は、前記言語モデルに前記サンプルを更に入力することによって、前記台本を生成する、
(1)~(10)の何れかに記載の台本生成システム。
Claims (13)
- 自然言語で記述された生成物を生成可能な学習済みの言語モデルに入力される入力情報であって、商品又はサービスが紹介される動画に関する前記入力情報を取得する入力情報取得部と、
前記言語モデルに前記入力情報を入力することによって、前記動画に関する台本を生成する台本生成部と、
を含む台本生成システム。 - 前記入力情報取得部は、前記商品又は前記サービスの特徴に関する前記入力情報である入力特徴情報を取得し、
前記台本生成部は、前記言語モデルに前記入力特徴情報を入力することによって、前記台本を生成する、
請求項1に記載の台本生成システム。 - 前記入力情報取得部は、前記商品又は前記サービスの特徴に関するコンテンツから抽出された抽出情報が前記言語モデル又は他の言語モデルに入力されることによって生成された前記入力特徴情報を取得する、
請求項2に記載の台本生成システム。 - 前記入力情報取得部は、前記動画の設定に関する前記入力情報である入力設定情報を取得し、
前記台本生成部は、前記言語モデルに前記入力設定情報を入力することによって、前記台本を生成する、
請求項1~3の何れかに記載の台本生成システム。 - 前記入力情報取得部は、前記動画における複数のセッションの各々に関する前記入力情報である入力セッション情報を取得し、
前記台本生成部は、前記言語モデルに前記入力セッション情報を入力することによって、前記台本を生成する、
請求項1~3の何れかに記載の台本生成システム。 - 前記入力情報取得部は、前記台本に関する詳細度に応じた前記入力セッション情報を取得し、
前記台本生成部は、前記言語モデルに、前記詳細度に応じた前記入力セッション情報を入力することによって、前記詳細度に応じた前記台本を生成する、
請求項5に記載の台本生成システム。 - 前記台本生成部は、
前記セッションごとに、当該セッションの前記入力セッション情報を前記言語モデルに入力することによって、前記台本の一部である台本部分であって、当該セッションの前記台本部分を生成し、
前記複数のセッションのうち、2番目以降の前記セッションについては、当該2番目以降のセッションよりも前の前記セッションの前記台本部分も前記言語モデルに入力することによって、当該2番目以降のセッションの前記台本部分を生成し、
前記複数のセッションの各々の前記台本部分に基づいて、前記台本を生成する、
請求項5に記載の台本生成システム。 - 前記入力情報取得部は、予め用意された前記入力情報である第1入力情報と、前記言語モデル又は他の言語モデルに前記第1入力情報が入力されることによって生成された前記入力情報である第2入力情報と、を取得し、
前記台本生成部は、前記言語モデルに前記第1入力情報及び前記第2入力情報を入力することによって、前記台本を生成する、
請求項1~3の何れかに記載の台本生成システム。 - 前記台本生成システムは、前記言語モデル又は他の言語モデルに前記台本を入力することによって、前記動画に関する質問と回答を生成する質問回答生成部を更に含む、
請求項1~3の何れかに記載の台本生成システム。 - 前記台本生成システムは、前記商品とは異なる他の商品、又は、前記サービスとは異なる他のサービスが紹介される他の動画に関するサンプルであって、前記他の商品又は前記他のサービスの特徴が省かれた前記サンプルを取得する第1サンプル取得部を更に含み、
前記台本生成部は、前記言語モデルに前記サンプルを更に入力することによって、前記台本を生成する、
請求項1~3の何れかに記載の台本生成システム。 - 前記台本生成システムは、前記動画の出演者と同じ出演者が出演した他の動画に関するサンプルを取得する第2サンプル取得部を更に含み、
前記台本生成部は、前記言語モデルに前記サンプルを更に入力することによって、前記台本を生成する、
請求項1~3の何れかに記載の台本生成システム。 - 自然言語で記述された生成物を生成可能な学習済みの言語モデルに入力される入力情報であって、商品又はサービスが紹介される動画に関する前記入力情報を取得する入力情報取得ステップと、
前記言語モデルに前記入力情報を入力することによって、前記動画に関する台本を生成する台本生成ステップと、
を含む台本生成方法。 - 自然言語で記述された生成物を生成可能な学習済みの言語モデルに入力される入力情報であって、商品又はサービスが紹介される動画に関する前記入力情報を取得する入力情報取得部、
前記言語モデルに前記入力情報を入力することによって、前記動画に関する台本を生成する台本生成部、
としてコンピュータを機能させるためのプログラム。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202480044429.2A CN121532789A (zh) | 2024-06-13 | 2024-06-13 | 剧本生成系统、剧本生成方法和程序 |
| PCT/JP2024/021527 WO2025258027A1 (ja) | 2024-06-13 | 2024-06-13 | 台本生成システム、台本生成方法、及びプログラム |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2024/021527 WO2025258027A1 (ja) | 2024-06-13 | 2024-06-13 | 台本生成システム、台本生成方法、及びプログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025258027A1 true WO2025258027A1 (ja) | 2025-12-18 |
Family
ID=98050257
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2024/021527 Pending WO2025258027A1 (ja) | 2024-06-13 | 2024-06-13 | 台本生成システム、台本生成方法、及びプログラム |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN121532789A (ja) |
| WO (1) | WO2025258027A1 (ja) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006285879A (ja) * | 2005-04-05 | 2006-10-19 | Keitaro Oketa | 広告コンテンツの配信方法、及び、プロフィールムービーを利用した広告方法 |
| JP2019023782A (ja) * | 2017-07-24 | 2019-02-14 | カシオ計算機株式会社 | 広告管理装置及びプログラム |
-
2024
- 2024-06-13 WO PCT/JP2024/021527 patent/WO2025258027A1/ja active Pending
- 2024-06-13 CN CN202480044429.2A patent/CN121532789A/zh active Pending
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006285879A (ja) * | 2005-04-05 | 2006-10-19 | Keitaro Oketa | 広告コンテンツの配信方法、及び、プロフィールムービーを利用した広告方法 |
| JP2019023782A (ja) * | 2017-07-24 | 2019-02-14 | カシオ計算機株式会社 | 広告管理装置及びプログラム |
Non-Patent Citations (1)
| Title |
|---|
| BITO MICHIKO: "Revolutionize scriptwriting with AI that creates SNS scripts! Streamline SNS marketing for individual entrepreneurs using AI", ACTIVE-NOTE.JP, 18 May 2024 (2024-05-18), XP093381677, Retrieved from the Internet <URL:https://www.active-note.jp/chatgpt/ai-that-creates-sns-scripts/> * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121532789A (zh) | 2026-02-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12346387B2 (en) | Systems and methods for automatically generating a website and related marketing assets using generative artificial intelligence | |
| US11158349B2 (en) | Methods and systems of automatically generating video content from scripts/text | |
| US12008064B1 (en) | Systems and methods for a website generator that utilizes artificial intelligence | |
| US12124524B1 (en) | Generating prompts for user link notes | |
| US20240289411A1 (en) | Systems and methods for automatically generating a website and providing customer service using generative artificial intelligence | |
| US20240289410A1 (en) | Systems and methods for automatically generating a website and related sales campaigns using generative artificial intelligence | |
| US20240362265A1 (en) | Hyper-personalized prompt based content generation | |
| US20240289408A1 (en) | Systems and methods for automatically generating a website and selecting a related domain name using generative artificial intelligence | |
| EP3547245A1 (en) | System and method for producing a customized video file | |
| CN118870146B (zh) | 基于大模型的视频生成方法、装置、电子设备、存储介质及程序产品 | |
| US20250322173A1 (en) | Systems and methods for inserting excerpts into a query response | |
| US20240292070A1 (en) | Iterative ai prompt optimization for video generation | |
| US12346384B2 (en) | Systems and methods for automatically generating a website and suggesting a related business entity type using generative artificial intelligence | |
| CN117436417A (zh) | 演示文稿生成方法、装置、电子设备和存储介质 | |
| CN119848293B (zh) | 基于新媒体数据分析的视频内容创作方法及装置 | |
| US20250201234A1 (en) | System for generating conversational content by utilizing generative ai and method thereof | |
| CN116843408A (zh) | 商品评价内容处理方法及电子设备 | |
| WO2025096850A1 (en) | Systems, methods, and media for automated creation of analytics-driven audio-visual interactive episodes | |
| Shao et al. | Evaluation on algorithms and models for multi-modal information fusion and evaluation in new media art and film and television cultural creation | |
| US20250193477A1 (en) | Streaming a segmented artificial intelligence virtual assistant with probabilistic buffering | |
| US20240281866A1 (en) | Systems and methods for automatically generating a website and suggesting a related insurance product using generative artificial intelligence | |
| WO2025258027A1 (ja) | 台本生成システム、台本生成方法、及びプログラム | |
| US20240289546A1 (en) | Synthesized responses to predictive livestream questions | |
| KR20240050294A (ko) | 개인화된 메타데이터 기반 콘텐츠 제공 장치 및 그 방법 | |
| KR20240020787A (ko) | 선호도 기반의 사용자 맞춤형 콘텐츠 서비스 시스템 및 방법 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24943441 Country of ref document: EP Kind code of ref document: A1 |