WO2025109448A1 - 情報処理システム、情報処理方法 - Google Patents

情報処理システム、情報処理方法 Download PDF

Info

Publication number
WO2025109448A1
WO2025109448A1 PCT/IB2024/061478 IB2024061478W WO2025109448A1 WO 2025109448 A1 WO2025109448 A1 WO 2025109448A1 IB 2024061478 W IB2024061478 W IB 2024061478W WO 2025109448 A1 WO2025109448 A1 WO 2025109448A1
Authority
WO
WIPO (PCT)
Prior art keywords
document
information processing
translation
processing device
function
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/IB2024/061478
Other languages
English (en)
French (fr)
Inventor
桃純平
原彩実
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Semiconductor Energy Laboratory Co Ltd
Original Assignee
Semiconductor Energy Laboratory Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Semiconductor Energy Laboratory Co Ltd filed Critical Semiconductor Energy Laboratory Co Ltd
Publication of WO2025109448A1 publication Critical patent/WO2025109448A1/ja
Anticipated expiration legal-status Critical
Pending legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/40Processing or translation of natural language
    • G06F40/51Translation evaluation

Definitions

  • One aspect of the present invention relates to an information processing system and an information processing method.
  • one aspect of the present invention is not limited to the above technical field.
  • the technical field of one aspect of the invention disclosed in this specification relates to an object, a method, or a manufacturing method.
  • one aspect of the present invention relates to a process, a machine, a manufacture, or a composition of matter. Therefore, more specifically, examples of the technical field of one aspect of the present invention disclosed in this specification include a semiconductor device, a display device, a light-emitting device, a power storage device, a memory device, a driving method thereof, or a manufacturing method thereof.
  • machine translation which uses computers to translate one natural language into another.
  • machine translation examples include rule-based machine translation, which translates based on rules, statistical machine translation, which translates using language models and translation models, and neural machine translation, which translates using an artificial neural network (ANN, hereafter simply referred to as neural network).
  • ANN artificial neural network
  • Literal translation refers to a translation that faithfully replaces every word of the original.
  • Free translation refers to a translation that focuses on the overall meaning, taking into account the context and nuances surrounding the original, rather than focusing on every single word. For this reason, free translations may add words that are not in the original.
  • Non-Patent Document 1 discloses GPT-4 (registered trademark) (Generative Pre-trained Transformer 4) as a large-scale language model, and ChatGPT as a chat service.
  • machine translation is improving year by year, and in cases where it is sufficient for the reader to understand the general content, or for the general content to be conveyed, machine translated text can be used as is.
  • machine-translated texts often contain mistranslations, such as "information omissions" (where elements in the original text are not reflected in the translation) and “information overload” (where elements not in the original text are included in the translation).
  • Checking and correcting a translation requires the translator's specialized knowledge and ability, while also being required to carry out the work within limited time and cost. Furthermore, when checking and correcting a translation, the content of the check differs depending on whether the translation is a literal or a free translation. In other words, the content of the check differs between a translation that is required to be a literal translation and a translation that is allowed to include a free translation. Therefore, if an information processing system with a checking function that can flexibly handle both literal and free translations were to be realized, it would be possible to efficiently check and correct translations.
  • one aspect of the present invention has an objective to provide an information processing system capable of efficiently checking and correcting a translation. Another aspect of the present invention has an objective to provide an information processing system capable of determining whether a translation includes a free translation. Another objective is to provide an information processing system capable of determining whether an original text and a translation include a free translation. Another objective is to provide an information processing system capable of outputting a correspondence between a word included in an original text and a word in a translation that corresponds to the word included in the original text when the translation is a direct translation. Another objective is to provide an information processing method used in any one or more of the above information processing systems. Another objective is to provide an information processing device including any one or more of the above information processing systems.
  • One aspect of the present invention has as its objective the provision of a new information processing system with excellent convenience, usefulness, or reliability. Alternatively, it has as its objective the provision of a new information processing method with excellent convenience, usefulness, or reliability. Alternatively, it has as its objective the provision of a new information processing system, a new information processing method, or a new semiconductor device.
  • One aspect of the present invention is an information processing system having a first information processing device and a second information processing device, the first information processing device having a function of accepting input of a first document in a first language and a second document in a second language which is a translation of the first document, a function of acquiring a result of the judgment by transmitting to the second information processing device an instruction statement for judging whether the second document contains a free translation of the first document and a prompt including the first document and the second document, a function of acquiring the correspondence between the words and phrases contained in the second document which have been judged not to contain a free translation and the words and phrases contained in the first document by transmitting to the second information processing device an instruction statement for outputting the correspondence between the words and phrases contained in the second document which have been judged not to contain a free translation and the first document, and a function of outputting the correspondence.
  • the first information processing device has a function of accepting a selection of whether to allow the second document to include a free translation, and a function of acquiring the result of the synonymy determination by sending to the second information processing device an instruction statement for performing a synonymy determination between the second document selected to allow the second document to include a free translation and the first document, and a prompt including the first document and the second document. It is also preferable that the first information processing device has a function of accepting a correction of the second document when there is an error in the synonymity determination between the second document and the first document.
  • the first information processing device has a function of accepting corrections to the second document determined to contain a free translation, and repeating the acceptance of corrections, the sending of prompts for determination, and the acquisition of the results of the determination until it is determined that the second document does not contain a free translation.
  • the first information processing device has a function of accepting corrections to the second document when there is an error in the correspondence between words contained in the second document and words contained in the first document.
  • one aspect of the present invention is an information processing system having a first information processing device and a second information processing device, the first information processing device having a function of accepting input of a first document in a first language and a second document in a second language that is a translation of the first document, a function of generating a literal translation of the first document in the second language, a function of generating a free translation of the first document in the second language, a function of calculating a similarity between the second document and the literal translation, a function of calculating a similarity between the second document and the free translation, a function of comparing the similarity of the literal translation and the similarity of the free translation, and determining that the second document does not contain a free translation of the first document if the similarity of the literal translation is higher, and a function of acquiring the correspondence by sending a prompt including an instruction for outputting a correspondence between a phrase included in the second document determined not to contain a free translation and a phrase included in the first document, and the first document and the first document
  • one aspect of the present invention is an information processing method having the steps of accepting input of a first document in a first language and a second document in a second language that is a translation of the first document, creating a first prompt including an instruction sentence for determining whether the second document contains a free translation of the first document and the first document, making a determination in accordance with the first prompt, creating a second prompt including an instruction sentence for outputting a correspondence between words contained in the second document and words contained in the first document when it is determined that the second document does not contain a free translation, the first document, and the second document, and outputting the correspondence in accordance with the second prompt.
  • the above information processing method preferably includes a step of accepting a selection of whether to allow the second document to include a free translation, a step of creating an instruction sentence for performing a synonymity determination between the second document selected in the selection to allow the second document to include a free translation and the first document, a step of creating a third prompt including the first document and the second document, and a step of performing a synonymity determination in accordance with the third prompt.
  • the determination is performed by a large-scale language model, and the output of the correspondence is performed by the large-scale language model. It is also preferable that the determination is performed by a large-scale language model, the output of the correspondence is performed by the large-scale language model, and the synonymity determination is performed by the large-scale language model.
  • an information processing system capable of efficiently checking and correcting a translation.
  • an information processing system capable of determining whether a translation contains a free translation.
  • an information processing system capable of determining whether an original text and a translation are synonymous when a translation contains a free translation.
  • it is possible to provide an information processing system capable of outputting the correspondence between a word or phrase included in an original text and a corresponding word or phrase in a translation when the translation is a literal translation.
  • FIG. 1 is a diagram illustrating the configuration of an information processing system.
  • FIG. 2 is a diagram illustrating the information processing method.
  • FIG. 3 is a diagram illustrating the information processing method.
  • FIG. 4 is a diagram illustrating an information processing method.
  • FIG. 5 is a diagram illustrating the configuration of an information processing system.
  • FIG. 6 is a diagram illustrating the configuration of an information processing device.
  • Fig. 7A is a diagram illustrating an information processing method
  • Fig. 7B is a diagram illustrating a configuration of an information processing system.
  • the information processing system of one embodiment of the present invention is an information processing system that can be used to compare an original text in one natural language with a translation into another natural language to check whether the translation has been performed correctly.
  • the information processing system of one embodiment of the present invention is an information processing system that can assist a translator in checking a translation.
  • the above translation is not limited to a translation obtained by machine translation, but also includes a translation obtained by manual translation (also called human translation).
  • the information processing system of one embodiment of the present invention has a check function that can flexibly handle both literal translations and free translations. In other words, it has a function of performing different information processing in cases where a literal translation is required and cases where free translation is allowed.
  • the information processing system of one embodiment of the present invention is preferably an information processing system that has a function that can handle both cases where a literal translation is required and cases where free translation is allowed, but it may also be an information processing system that has a function that can handle either one of them.
  • a translation contains a free translation
  • the free translation can be detected and the user of the information processing system can be prompted to correct it.
  • a translation determined to not contain a free translation or a translation corrected until it is determined to not contain a free translation
  • an information processing system it is possible to determine whether a translated sentence is synonymous with the original sentence in cases where free translation is permitted. Furthermore, by presenting the results of the synonymity determination to a user of the information processing system, it is possible to prevent non-synonymous translated sentences (mistranslations) from being overlooked, and to efficiently correct such mistranslations.
  • an information processing system has a function of accepting input of a document in a first language (original text) and a document in a second language (translation) in which the original text is translated into a second language when a literal translation is required.
  • the information processing system also has a function of determining whether the input translation contains a free translation of the original text, and if this determination does not result in a free translation (if the translation is a literal translation), has a function of outputting the correspondence between the words in the original text and the words in the translation.
  • a user of the information processing system can check whether the translation has been performed correctly by confirming the above correspondence. If the input translation is determined to contain a free translation, the user of the information processing system can correct the translation. If an error or difference is included in the above correspondence, the user of the information processing system can correct the translation.
  • an information processing system has a function of accepting input of a document in a first language (original text) and a document in a second language (translated text) in which the original text is translated into a second language, in cases where it is permitted to include a free translation.
  • the information processing system also has a function of determining whether the translated text is synonymous with the original text (synonymous determination).
  • the information processing system also has a function of outputting the results of the synonymous determination. By checking the results of the synonymous determination, a user of the information processing system can check whether the translation has been performed correctly.
  • a user of the information processing system can correct the translated text if the inputted translated text contains a portion that is not synonymous with the original text.
  • the information processing system of one embodiment of the present invention has a function of repeating the above series of processes from input to output in units of one sentence, two or more sentences, one paragraph, two or more paragraphs, etc.
  • the user can check the original text, the translation, and the word correspondence or the results of the synonym determination as a set and make corrections as necessary, but this checking and selection may be done for each of the above repetition units, for multiple repetition units, or after the entire processing is completed.
  • FIG. 1 is a diagram illustrating the configuration of an information processing system according to one embodiment of the present invention.
  • the information processing system described in this embodiment includes an information processing device 10 and an information processing device 40.
  • the information processing device 10 and the information processing device 40 are connected via a network 30, and can transmit and receive document data and the like.
  • the information processing system described in this embodiment may be configured so that a user can directly operate the information processing device 10 to input documents, etc., or as shown in FIG. 1, may be configured so that a user can input documents, etc., using an information terminal 20 connected to the information processing device 10 via a network 31.
  • [Information processing method] 2 to 4 are flowcharts illustrating an example of an information processing method according to one embodiment of the present invention.
  • the information processing method is "started," and in step S101 of FIG. 2, a first document (original text) expressed in a first language and a second document (translation) in which the first document has been translated into a second language are input.
  • step S111 of FIG. 2 a selection is made as to whether or not the translation (second document) is allowed to include a free translation. Note that this selection may be made each time by the user of this information processing system, or may be input in advance in step S101 together with the original text (first document) and the translation (second document). In this way, the selection of whether or not to allow the translation (second document) to include a free translation is accepted.
  • step S111 of FIG. 2 if it is selected that the translation (second document) may contain a free translation (if the inclusion of a free translation is permitted), the process proceeds to connector A, which connects from step S111 to the flowchart shown in FIG. 4.
  • step S111 of FIG. 2 if it is selected that the translation (second document) is not allowed to include a free translation (if a literal translation is required), the process proceeds to step S121.
  • step S121 of FIG. 2 it is determined whether or not the translation (second document) includes a free translation.
  • step S121A which includes the processes of steps S201 to S204, as an example of the process in step S121 in FIG. 2.
  • the first process in step S121A is to create a prompt in step S201 of FIG. 3 to determine whether the translation (second document) contains a free translation of the original text (first document).
  • the prompt for determining whether a free translation is included includes the original text (first document), the translation (second document), and an instruction such as "Please point out any parts of the original text and translation that are free translated."
  • a prompt can be thought of as an input sentence that causes the language model to perform a desired operation.
  • a prompt When a prompt is given to the language model, it generates a response content based on the prompt.
  • step S201 The prompt created in step S201 is sent to the large-scale language model in step S202, which allows the large-scale language model to determine in step S203 whether the translation (second document) contains a free translation of the original text (first document).
  • step S204 the determination result output from the large-scale language model is received, and the process can proceed to step S122 in FIG. 2.
  • the information processing system can determine whether the translation (second document) contains a free translation.
  • step S122 of FIG. 2 the judgment result obtained in step S121 is presented to the user of the information processing system. If it is judged that the translation (second document) contains a free translation, the portion judged to be a free translation can be presented to the user. In step S123, the user corrects the portion of the translation (second document) judged to be a free translation. The information processing system accepts the corrected translation (second document). Note that the above corrections can be repeated until it is judged that the translation (second document or the second document corrected in step S123) does not contain a free translation.
  • step S121 If it is determined in step S121 that the translation does not include a free translation, the process proceeds to step S131.
  • step S131 of FIG. 2 a prompt is created to determine the correspondence (correspondence determination) between the original text (first document) and the translation (second document or the second document corrected in step S123).
  • the prompt for determining correspondence includes the original text (first document), the translation (second document or the second document corrected in step S123), and an instruction such as "The original text and translation below are direct translations. Please create a correspondence table for each word.” By sending this prompt to the large-scale language model, the correspondence between the original text (first document) and the translation (second document or the second document corrected in step S123) can be output from the large-scale language model.
  • the large-scale language model determines the correspondence between the original text (first document) and the translation (second document or the second document corrected in step S123).
  • the large-scale language model also outputs, for example, a correspondence table as a result of the determination of the correspondence.
  • an instruction may be created to explain the content of that case.
  • the instruction included in the prompt may be something like, "The original text and translation below are literal translations. Please create a correspondence table for each word. Then, please report any differences in the correspondence from the correspondence table.”
  • Tables 1 to 9 show examples of the above prompts and their results.
  • Tables 1 to 3 are examples of cases where there is no difference in the correspondence between the original text and the translation
  • Tables 4 to 6 are a first example of cases where there is a difference in the correspondence between the original text and the translation
  • Tables 7 to 9 are a second example of cases where there is a difference in the correspondence between the original text and the translation. Note that in the first example above, there is a omission in the translation (a mouse), and in the second example above, there is an error in the numbers (symbols) in the translation (correct: 101, incorrect: 102).
  • Tables 1, 4, and 7 are examples of prompts to be input to the large-scale language model
  • Tables 2, 5, and 8 are examples of tables of correspondences output from the large-scale language model
  • Tables 3, 6, and 9 are examples of explanations of correspondences output from the large-scale language model.
  • step S133 of FIG. 2 a choice is made as to whether to revise the translation (the second document or the second document revised in step S123).
  • step S133 As a selection in step S133, as in the examples of Tables 1 to 3, if the correspondence output from the large-scale language model in step S132 does not include any differences, the information processing method of one embodiment of the present invention can be "ended.” Alternatively, as in the examples of Tables 4 to 9, as a selection in step S133, if the correspondence output from the large-scale language model in step S132 includes any differences, the process proceeds to step S134, where the translation is corrected. Note that the above corrections can be repeated until it is determined that the correspondence between the original text and the translation does not include any differences.
  • the selection in step S133 can also be made by the user of this information processing system by referring to the correspondence output from the large-scale language model in step S132.
  • step S111 A case where it is selected in step S111 that the translation (second document) may include a free translation (a case where the inclusion of a free translation is permitted) will be described with reference to FIG.
  • step S141 which is connected from connector A shown in Figure 4, a prompt is created to determine whether the original text (first document) and the translation (second document) are synonymous.
  • the synonymy determination prompt includes an original text (first document), a translation (second document), and an instruction such as, "Please indicate in a table the parts of the original text and translation below that have been loosely translated. Then, please indicate whether the original text and the translation have the same meaning by choosing "synonymous" or "not synonymous.”
  • This prompt By sending this prompt to a large-scale language model, it is possible to have the large-scale language model determine whether the original text (first document) and the translation (second document) are synonymous and output the determination result.
  • a machine learning model can be used for synonym determination.
  • a neural network model BERT (Bidirectional Encoder Representation from Transformers), can be fine-tuned and used.
  • an entailment recognition task can be used to replace entailment and non-entailment with synonymous and non-synonymous. For example, two sentences can be compared, and those with equivalent information can be labeled as synonymous, and those with unequal information can be labeled as non-synonymous, and then learned. Also, for example, a first document can be compared with a second document, and the information in the second document can be labeled as insufficient, excessive, equivalent (synonymous), or contradictory to the first document, and then learned.
  • sentence vectors may be compared and evaluated using the cosine similarity of the vectors. Similarity values above a certain value may be considered synonymous, and similarity values below a certain value may be considered not synonymous.
  • LLM can be used for synonymy determination.
  • a method may be used in which a binary choice is output, whether the first document and the second document are synonymous or not synonymous, rather than a numerical value of the degree of synonymy.
  • synonymy determination using LLM if the information contained in the second document is excessive or missing compared to the information contained in the first document, this information may be added and output together with the determination. Note that in synonymy determination using LLM, the first document and the second document are compared, and learning may be performed by labeling the information in the second document as being insufficient, excessive, equivalent (synonymous), or contradictory to the first document.
  • the cosine similarity of vectors and the synonymy judgment label may be used to judge synonymity.
  • the cosine similarity may be converted to a cosine distance
  • the label evaluation value may be calculated by multiplying the cosine distance by the label evaluation value, with over-abundance labels being 1, under-abundance labels being -1, synonymous labels being 0, and inconsistent labels being nan.
  • over-abundance labels being 1, under-abundance labels being -1
  • synonymous labels being 0, and inconsistent labels being nan.
  • inconsistent labels being nan.
  • step S143 of FIG. 4 the result of the synonymity determination obtained in step S142 is presented to the user of the information processing system, and the user selects whether to revise the translation (second document). The information processing system then accepts the selection.
  • step S143 If the selection in step S143 is that the synonymy determination results output from the large-scale language model in step S142 do not include a determination result of "not synonymous," the information processing method of one aspect of the present invention can be "ended.” Alternatively, if the selection in step S143 is that the synonymy determination results output from the large-scale language model in step S142 include a determination result of "not synonymous," the parts determined to be not synonymous can be presented to the user, and in step S144, the user corrects the parts of the translation (second document) determined to be not synonymous. The information processing system accepts the corrected translation (second document). Note that the above corrections can be repeated until the results of the synonymy determination between the original text and the translation no longer include results determined to be not synonymous.
  • the selection in step S143 can also be made by the user of this information processing system by referring to the results of the synonymity determination output from the large-scale language model in step S142.
  • Fig. 5 will be used to explain which device or terminal, information processing device 10, information processing device 40, or information terminal 20, performs the processing of each step.
  • FIG. 5 is a relationship diagram showing the processes performed by information processing device 10, information processing device 40, and information terminal 20, using the symbols for each step used in FIGS. 2 to 4.
  • the block arrow with network 30 attached indicates the transmission and reception of data between information processing device 10 and information processing device 40
  • the block arrow with network 31 attached indicates the transmission and reception of data between information processing device 10 and information terminal 20.
  • the information processing device 10 has a function for performing processing of step S201, a function for performing processing of step S202, a function for performing processing of step S204, a function for performing processing of step S122, a function for performing processing of step S131, and a function for performing processing of step S141.
  • the information processing device 40 also has a function for performing the process of step S203, a function for performing the process of step S132, and a function for performing the process of step S142.
  • the information terminal 20 has a function for performing processing of step S101, a function for performing processing of step S111, a function for performing processing of step S123, a function for performing processing of step S133, a function for performing processing of step S134, a function for performing processing of step S143, and a function for performing processing of step S144.
  • the following describes examples of the configurations of the information processing device 10, the information processing device 40, the information terminal 20, the network 30, and the network 31.
  • Configuration example of information processing device 10 A configuration example of an information processing device 10 included in an information processing system of one embodiment of the present invention will be described with reference to FIG.
  • the information processing device 10 has an input unit 110, a memory unit 120, a processing unit 130, an output unit 140, and a transmission path 150.
  • the input unit 110 can accept data from outside the information processing device 10.
  • the input unit 110 can accept data from the information terminal 20.
  • the input unit 110 can also accept data from the information processing device 40.
  • a device such as a wired communication port, a wireless communication port, or an optical communication port can be used.
  • the input unit 110 can supply the received data to one or both of the memory unit 120 and the processing unit 130 via the transmission path 150.
  • the storage unit 120 has a function of storing a program executed by the processing unit 130.
  • the storage unit 120 may also have a function of storing data generated by the processing unit 130 (e.g., a calculation result, an analysis result, an inference result), data accepted by the input unit 110, and the like.
  • the storage unit 120 may have a database. Furthermore, the information processing device 10 may have a database separate from the storage unit 120. The information processing device 10 may have a function to extract data from a database that exists outside the storage unit 120, outside the information processing device 10, or outside the information processing system. Furthermore, the information processing device 10 may have a function to extract data from both its own database and an external database.
  • One or both of the storage and the file server can be used for the memory unit 120. Also, a database that records the paths of files stored in the file server can be used for the memory unit 120.
  • the memory unit 120 has at least one of a volatile memory and a non-volatile memory.
  • volatile memory include DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory).
  • non-volatile memory include ReRAM (Resistive Random Access Memory, also called resistive memory), PRAM (Phase change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory, also called magnetoresistive memory), and flash memory.
  • the storage unit 120 may also have at least one of NOSRAM (registered trademark) and DOSRAM (registered trademark).
  • the storage unit 120 may also have a recording media drive. Examples of recording media drives include a hard disk drive (HDD) and a solid state drive (SSD).
  • NOSRAM is an abbreviation for "Nonvolatile Oxide Semiconductor Random Access Memory”.
  • NOSRAM is a memory in which the memory cell is a two-transistor (2T) or three-transistor (3T) gain cell, and the transistor is a transistor (also called an OS transistor) that uses metal oxide in the channel formation region.
  • the current that flows between the source and drain in the off state, that is, the leakage current is extremely small.
  • NOSRAM can be used as a nonvolatile memory by retaining a charge according to data in the memory cell using the characteristic of extremely small leakage current.
  • NOSRAM can read the retained data without destroying it (nondestructive read), so it is suitable for arithmetic processing in which only data read operations are repeated in large quantities. Since NOSRAM can increase the data capacity by stacking, it can be used as a large-scale cache memory, main memory, or storage memory to improve the performance of semiconductor devices.
  • DOSRAM is an abbreviation for "Dynamic Oxide Semiconductor RAM” and refers to a RAM with 1T (transistor) 1C (capacitor) type memory cells.
  • DOSRAM is a DRAM formed using OS transistors, and is a memory that temporarily stores information sent from the outside.
  • DOSRAM is a memory that takes advantage of the small off-current of OS transistors.
  • metal oxide is a metal oxide in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also called oxide semiconductors or simply OS), and the like. For example, when a metal oxide is used in the semiconductor layer of a transistor, the metal oxide may be called an oxide semiconductor.
  • the metal oxide in the channel formation region preferably contains indium (In).
  • the carrier mobility (electron mobility) of the OS transistor is increased.
  • the metal oxide in the channel formation region is preferably an oxide semiconductor containing element M.
  • the element M is preferably at least one of aluminum (Al), gallium (Ga), and tin (Sn).
  • element M Other elements that can be used for element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W).
  • element M a combination of multiple elements described above may be used.
  • the element M is, for example, an element that has a high binding energy with oxygen. For example, it is an element that has a higher bond energy with oxygen than indium.
  • the metal oxide in the channel formation region is a metal oxide that contains zinc (Zn). Metal oxides that contain zinc may be more likely to crystallize.
  • the metal oxide of the channel-forming region is not limited to metal oxides containing indium.
  • the metal oxide of the channel-forming region may be, for example, a metal oxide containing zinc, a metal oxide containing gallium, or a metal oxide containing tin, which does not contain indium, such as zinc tin oxide or gallium tin oxide.
  • the processing unit 130 has a function of performing processes such as calculation, analysis, and inference using data supplied from one or both of the input unit 110 and the storage unit 120.
  • the processing unit 130 can supply generated data (e.g., calculation results, analysis results, inference results) to one or both of the storage unit 120 and the output unit 140.
  • the processing unit 130 has a function of acquiring data from the storage unit 120.
  • the processing unit 130 may also have a function of recording or registering data in the storage unit 120.
  • the processing unit 130 may have, for example, an arithmetic circuit.
  • the processing unit 130 may have, for example, a central processing unit (CPU: Central Processing Unit).
  • the processing unit 130 may also have a GPU (Graphics Processing Unit).
  • the processing unit 130 may have a microprocessor such as a DSP (Digital Signal Processor).
  • the microprocessor may be realized by a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array) or an FPAA (Field Programmable Analog Array).
  • the processing unit 130 may also have a quantum processor.
  • the processing unit 130 can perform various data processing and program control by interpreting and executing commands from various programs using the processor. Programs that can be executed by the processor are stored in at least one of the memory area of the processor and the storage unit 120.
  • the processing unit 130 may have a main memory.
  • the main memory may have at least one of a volatile memory such as a RAM (Random Access Memory) and a non-volatile memory such as a ROM (Read Only Memory).
  • the main memory may also have at least one of the above-mentioned NOSRAM and DOSRAM.
  • RAM for example, DRAM, SRAM, etc.
  • a virtual memory space is allocated and used as a working space for the processing unit 130.
  • the operating system, application programs, program modules, program data, lookup tables, etc. stored in the storage unit 120 are loaded into the RAM for execution. These data, programs, and program modules loaded into the RAM are each directly accessed and operated by the processing unit 130.
  • ROM can store BIOS (Basic Input/Output System) and firmware that do not require rewriting.
  • BIOS Basic Input/Output System
  • Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), and EPROM (Erasable Programmable Read Only Memory).
  • Examples of EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory), which allows stored data to be erased by exposure to ultraviolet light, EEPROM (Electrically Erasable Programmable Read Only Memory), and flash memory.
  • the processing unit 130 can have one or both of an OS transistor and a transistor having silicon in the channel formation region (Si transistor).
  • the processing unit 130 preferably includes an OS transistor. Since the off-state current of an OS transistor is extremely small, by using the OS transistor as a switch for retaining charge (data) that has flowed into a capacitive element that functions as a memory element, it is possible to ensure a long data retention period. By using this characteristic in at least one of the register and cache memory of the processing unit, it is possible to operate the processing unit only when necessary, and to turn off the processing unit at other times by saving the information of the previous processing in the memory element. In other words, normally-off computing becomes possible, and the power consumption of the information processing system can be reduced.
  • the information processing device 10 uses AI for at least some of its processing.
  • the information processing device 10 uses an artificial neural network.
  • the neural network is realized by a circuit (hardware) or a program (software).
  • a neural network refers to a general model that mimics the neural circuit network of a living organism, determines the strength of connections between neurons through learning, and has problem-solving capabilities.
  • a neural network has an input layer, an intermediate layer (hidden layer), and an output layer.
  • connection strengths also called weighting coefficients
  • the information processing device 10 can perform processing using a natural language processing model that uses AI.
  • machine translation processing can be performed using a natural language processing model that uses AI.
  • processing can be performed using a natural language processing model that uses AI, such as Sequence-to-Sequence (seq2seq), Transformer, BERT (Bidirectional Encoder Representations from Transformers), and T5 (Text-to-Text Transformer Transformer).
  • the output unit 140 can output at least one of the calculation result, the analysis result, and the inference result in the processing unit 130 to the outside of the information processing device 10.
  • a device such as a wired communication port, a wireless communication port, or an optical communication port can be used.
  • the output unit 140 can transmit data to the information processing device 40.
  • the output unit 140 can also transmit data to the information terminal 20.
  • the transmission path 150 has a function of transmitting data. Data can be transmitted and received between the input unit 110, the storage unit 120, the processing unit 130, and the output unit 140 via the transmission path 150. Specifically, a bus line on a motherboard, a wired communication cable, or an optical communication cable can be used.
  • the information processing device 40 can process the received data and transmit the results of the processing.
  • the information processing device 40 can perform processing such as calculation using the data received from the information processing device 10.
  • the information processing device 40 can transmit the results of the processing to the information processing device 10. This can reduce the burden of calculation on the information processing device 10.
  • the information processing device 40 can perform processing using a natural language processing model that uses AI. For example, it can execute processing using a natural language processing model that uses AI, such as BERT (Bidirectional Encoder Representations from Transformers) and T5 (Text-to-Text Transformer Transformer).
  • a natural language processing model that uses AI such as BERT (Bidirectional Encoder Representations from Transformers) and T5 (Text-to-Text Transformer Transformer).
  • the information processing device 40 can also perform processing using a model that utilizes a large-scale language model (such as a sentence generation model or a dialogue model).
  • a model that utilizes a large-scale language model such as a sentence generation model or a dialogue model.
  • step S132 described in FIG. 2 it is preferable to use a model that utilizes a large-scale language model to determine the correspondence between the words contained in the original text and the words contained in the translation.
  • step S203 described in FIG. 3 it is preferable to use a model that utilizes a large-scale language model to determine whether the translation contains a free translation of the original text.
  • step S142 described in FIG. 4 it is preferable to use a model that utilizes a large-scale language model to determine whether the original text and the translation are synonymous.
  • processing can be performed using large-scale language models such as GPT-3, GPT-3.5, GPT-4 (registered trademark), LaMDA (Language Model for Dialogue Applications), PaLM (Pathways Language Model), Llama2, and Llama3.
  • GPT-4 registered trademark
  • the information processing device 40 can translate a document written in language A into a document written in language B.
  • the information processing device 40 can also translate a document written in language A into a document written in language B according to a given instruction statement.
  • the instruction statement can include a constraint. This makes it possible to control the degree of freedom in the translation by giving a constraint.
  • the information processing device 40 can execute processing using a general-purpose language processing model that can perform a variety of natural language processing tasks.
  • the information processing device 40 is a large computer such as a server computer or a supercomputer. It is also preferable that the information processing device 40 has a function as a parallel computer. By using the information processing device 40 as a parallel computer, it is possible to perform large-scale calculations necessary for AI learning and inference, for example.
  • the information processing device 40 is a computer with higher processing power than the information processing device 10.
  • the information processing device 40 has higher processing power than the information processing device 10 and can perform large-scale calculations.
  • the information processing device 40 can execute processing using a large-scale AI model compared to the information processing device 10.
  • the service provider does not necessarily need to own the information processing device 40.
  • the service provider can use part of the service provided by another business operator using the information processing device 40.
  • the information terminal 20 can receive data input by a user of the information processing system according to an aspect of the present invention.
  • the information terminal 20 can also provide the user with data output by the information processing system according to an aspect of the present invention.
  • the information terminal 20 can transmit data received from the user to the information processing device 10. In addition, the information terminal 20 can provide data received from the information processing device 10 to the user.
  • the information terminal 20 can transmit data generated based on data received from the user to the information processing device 10. In addition, the information terminal 20 can provide the user with data generated based on data received from the information processing device 10.
  • dedicated application software or a web browser is installed on the information terminal 20.
  • a user can access the information processing device 10 via either of them. This allows the user to enjoy services using an information processing system according to one embodiment of the present invention, for example, using a computer with lower processing power than the information processing device 10.
  • the information terminal 20 can also be called a client computer, etc.
  • Each information terminal 20 is an information terminal device used by a user of the information processing system according to one embodiment of the present invention.
  • a desktop computer 20a a notebook computer 20b, a smartphone 20c, and a tablet computer 20d can be used as the information terminal 20.
  • the tablet computer 20d can also be used as a notebook computer by connecting it to a housing 21 that has a keyboard.
  • a user of the information processing system can use a computer, such as an information terminal 20, that has lower processing power than the information processing device 10 or the information processing device 40 to control the information processing system of one embodiment of the present invention, give instructions to the information processing device 40, and enjoy the service.
  • a computer such as an information terminal 20, that has lower processing power than the information processing device 10 or the information processing device 40 to control the information processing system of one embodiment of the present invention, give instructions to the information processing device 40, and enjoy the service.
  • Network 30 The network 30 connects the information processing device 10 and the information processing device 40. This enables input data and processed data to be transmitted and received between the two devices. In addition, the load related to information processing can be distributed.
  • the network 30 is mainly a computer network that is larger than the network 31.
  • a global network can be used for the network 30.
  • the Internet which is the foundation of the World Wide Web (WWW), can be used.
  • WWW World Wide Web
  • Network 31 The network 31 connects a plurality of information terminals 20 and the information processing device 10. This enables data to be transmitted and received between the two. Also, the load related to information processing can be distributed. Also, a service provider can provide a user with a service using an information processing method according to one aspect of the present invention via the network 31, for example.
  • a local network can be used for network 31.
  • An intranet or an extranet can also be used for network 31.
  • a PAN Personal Area Network
  • a LAN Local Area Network
  • a CAN Campus Area Network
  • a MAN Metropolitan Area Network
  • a WAN Wide Area Network
  • GAN Global Area Network
  • a provider of a service using an information processing method according to one embodiment of the present invention and a user who receives the service belong to the same organization such as a company
  • data transmission between the information terminal 20 and the information processing device 10 is performed, for example, using a network 31 constructed within the organization. This allows data to be transmitted and received between the information terminal 20 and the information processing device 40 more securely than when data is transmitted over the Internet. It also makes it possible to prevent confidential information within the organization from leaking to the outside.
  • communication standards such as the fourth generation mobile communication system (4G), fifth generation mobile communication system (5G), and sixth generation mobile communication system (6G), or specifications standardized by the IEEE such as Wi-Fi (registered trademark) and Bluetooth (registered trademark), can be used as communication protocols or communication technologies.
  • Embodiment 2 In this embodiment, an information processing system and an information processing method according to one embodiment of the present invention will be described.
  • the information processing system has a configuration example having a partly different configuration from the first configuration example of the information processing system described in Embodiment 1.
  • step S121 in FIG. 2 differs from the flow described in FIG. 3 of information processing system configuration example 1. Specifically, the processing content of step S121 in FIG. 2 is performed using step S121B having the processing of steps S301 to S304 shown in FIG. 7A. The processing of steps S301 to S304 will be described using FIG. 7A.
  • step S301 of FIG. 7A the first document (original text) is directly translated into the second language to generate a third document (direct translation).
  • step S302 of FIG. 7A the first document (original text) is translated into the second language to include a free translation, generating a fourth document (free translation).
  • steps S301 and S302 may be performed in such a way that step S302 is performed first and then step S301 is performed, or steps S301 and S302 may be performed in parallel.
  • a dedicated translation model trained on a dataset consisting of literal translations can be used, and for generating free translations (step S302), a dedicated translation model trained on a dataset consisting of free translations can be used. This reduces the amount of calculations without the need to create prompts, and because the generation is done using different models, different sentences can be generated for literal translations and free translations.
  • the natural language processing model described in embodiment 1 can be used as the translation model.
  • step S303 of FIG. 7A the similarity between the second document (translated text) input in step S101 of FIG. 2 and each of the third document (literal translation) and the fourth document (free translation) is calculated.
  • the sentences can be vectorized and the cosine similarity can be used.
  • SentenceBERT or similar can be used to vectorize sentences.
  • Cosine similarity is expressed as a real number between -1 and 1, and the closer the value is to 1, the more similar the two documents being compared are (the greater the similarity), and the closer the cosine similarity value is to -1, the more dissimilar they are (the lower the similarity).
  • an edit distance such as the Levenshtein distance or the Jaro-Winkler distance may be used instead of the similarity.
  • an edit distance is used instead of the similarity, the closer the value is to 0, the more similar the two documents being compared are (the higher the similarity), and the larger the edit distance value, the more dissimilar the documents are (the lower the similarity).
  • the BLEU (Bilingual Evaluation Understudy) score may be used instead of the similarity.
  • the BLEU score is expressed as a real number between 0 and 1, with a value closer to 1 indicating that the two documents being compared are more similar (high similarity), and a value closer to 0 indicating that the documents are less similar (low similarity).
  • step S304 of FIG. 7A the similarity between the second document (translation) and the third document (literal translation) is compared with the similarity between the second document (translation) and the fourth document (free translation). If the result of the comparison shows that the similarity of the fourth document (free translation) is higher than that of the third document (literal translation), it is determined that the second document (translation) contains a free translation. Alternatively, if the result of the comparison shows that the similarity of the third document (literal translation) is higher than that of the fourth document (free translation), it is determined that the second document (translation) does not contain a free translation. After this determination in step S304, the process can proceed to step S122 in FIG. 2.
  • FIG. 7B is a relationship diagram showing the processes performed by information processing device 10, information processing device 40, and information terminal 20 using the symbols for each step used in FIGS. 2, 4, and 7A, and is a different configuration example from FIG. 5.
  • the block arrow with network 30 attached indicates the transmission and reception of data between information processing device 10 and information processing device 40
  • the block arrow with network 31 attached indicates the transmission and reception of data between information processing device 10 and information terminal 20.
  • the information processing device 10 has a function for performing processing of step S301, a function for performing processing of step S302, a function for performing processing of step S303, a function for performing processing of step S304, a function for performing processing of step S122, a function for performing processing of step S131, and a function for performing processing of step S141.
  • the information processing device 40 has a function for performing the processing of step S132 and a function for performing the processing of step S142.
  • the information terminal 20 has a function for performing processing of step S101, a function for performing processing of step S111, a function for performing processing of step S123, a function for performing processing of step S133, a function for performing processing of step S134, a function for performing processing of step S143, and a function for performing processing of step S144.
  • step S121 determine whether the translation contains a free translation
  • step S121 determine whether the translation contains a free translation
  • the information processing system, information processing method, and information processing device described above make it possible to efficiently check and correct translations.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computational Linguistics (AREA)
  • General Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Machine Translation (AREA)

Abstract

利便性、有用性または信頼性に優れた新規な情報処理システムを提供する。 原文と翻訳文を比較し、翻訳が正しく行われているかをチェックする際に用いることができ、直訳であることが求められる場合と、意訳を含むことが許容される場合と、で異なる情報処理を行う機能を有する情報処理システムである。情報処理システムは、直訳であることが求められる場合において、翻訳文に意訳された箇所が含まれる場合に、当該意訳箇所を検出し、情報処理システムの使用者に修正を促す機能を有し、さらに、原文と翻訳文のそれぞれの語句の対応関係を判定する機能を有する。また、意訳を含むことが許容される場合において、翻訳文が、原文と同義であるかの判定を行う機能を有する。

Description

情報処理システム、情報処理方法
 本発明の一態様は、情報処理システム、情報処理方法に関する。
 なお、本発明の一態様は、上記の技術分野に限定されない。本明細書等で開示する発明の一態様の技術分野は、物、方法、または、製造方法に関するものである。または、本発明の一態様は、プロセス、マシン、マニュファクチャ、または、組成物(コンポジション・オブ・マター)に関するものである。そのため、より具体的に本明細書で開示する本発明の一態様の技術分野としては、半導体装置、表示装置、発光装置、蓄電装置、記憶装置、それらの駆動方法、または、それらの製造方法、を一例として挙げることができる。
 コンピュータを用いて、ある自然言語を別の自然言語に翻訳する、機械翻訳の研究開発が盛んに行われている。機械翻訳としては、ルールに基づいて翻訳を行うルールベース機械翻訳、言語モデル及び翻訳モデルなどを用いて翻訳を行う統計的機械翻訳、並びに、人工ニューラルネットワーク(ANN:Artificial Neural Network、以下、単にニューラルネットワークとも記す)を用いて翻訳を行うニューラル機械翻訳などが挙げられる。
 翻訳には、直訳と意訳がある。直訳は、原文の一語一語を忠実に置き換えるように訳すことをいう。また、意訳とは、原文の一語一語にとらわれず、前後の文脈、ニュアンスをくみ取って、全体の意味に重点をおいて訳すことをいう。そのため、意訳では、原文にはない語句を追加する場合がある。
 直訳と意訳は、どちらかが正しいというものではなく、文書の目的、状況等に応じて、適宜選択されるものであり、直訳が好ましい場合の翻訳、意訳が好ましい場合の翻訳がある。機械翻訳においても、意訳を行うことが可能な翻訳装置が検討されている(特許文献1)。
 近年、ニューラルネットワークを用いた言語モデルの開発が盛んに行われており、特に大規模言語モデル(LLM:Large Language Model)が注目されている。大規模言語モデルは、大量のデータを用いて学習された自然言語処理モデルである。大規模言語モデルにより、例えばユーザの指示に対して回答を行う対話モデルを実現できる。非特許文献1では、大規模言語モデルとしてGPT−4(登録商標)(Generative Pre−trained Transformer 4)が開示されており、また、チャットサービスとしてChatGPTが開示されている。
 大規模言語モデルを利用することで、自然言語処理モデルの能力が大幅に上昇している。一方で、言語モデルの巨大化により、自前で言語モデルを組み込んで運用することは設備及び費用の面から難しい。そのため、言語モデルを提供する外部サービスを利用することが、言語モデルの利用形態の一つとなっている。
特開2009−205518
Summary of ChatGPT/GPT−4 Research and Perspective Towards the Future of Large Language Models,Yiheng Liu et al.(Submitted on 4 Apr 2023、[online]、インターネット<URL:https://arxiv.org/abs/2304.01852>
 機械翻訳の精度は年々向上しており、おおよその内容が読み手に伝わればよい場合、おおよその内容を理解できればよい場合、などには、機械翻訳による翻訳文を、そのまま用いることができる。
 しかしながら、機械翻訳された文章には、原文に含まれている要素が翻訳文に反映されていない「情報抜け(訳抜け)」、原文に含まれていない要素が翻訳文に盛り込まれている「情報過多(湧き出し)」、などの誤訳が含まれる場合も多い。
 契約書、説明書、公文書、特許文書、論文など、誤訳が許されない文書も多く存在する。このような文書の翻訳においては、人手による翻訳、機械による翻訳、に関わらず、翻訳後のチェック及び修正する工程が重要である。なお、機械翻訳による翻訳文を、人手によりチェック及び修正することをポストエディットと呼ぶ場合があり、国際規格ISO18587では、ポストエディットについての要求事項が定められている。
 翻訳文のチェック及び修正には、翻訳者としての専門的な知識及び力量が求められる一方で、限られた時間、コストのなかで、これらを遂行することが求められる。また、翻訳文のチェック及び修正において、その翻訳が直訳か意訳であるかによってチェックの内容は異なる。つまり、直訳であることが求められる場合の翻訳文と、意訳が含まれることが許容される場合の翻訳文と、ではチェックの内容が異なる。そのため、直訳と意訳に柔軟に対応することが可能なチェック機能を備えた情報処理システムが実現されれば、翻訳文のチェック及び修正する作業を効率的に行うことが可能になる。
 そこで、本発明の一態様は、翻訳文のチェック及び修正する作業を効率的に行うことが可能な情報処理システムを提供することを課題の一とする。または、本発明の一態様は、翻訳文に意訳が含まれているかどうかを判定することが可能な情報処理システムを提供することを課題の一とする。または、翻訳文に意訳が含まれる場合に、原文と翻訳文が同義であるかを判定することが可能な情報処理システムを提供することを課題の一とする。または、翻訳文が直訳である場合に、原文に含まれる語句と、当該語句に対応する翻訳文の語句と、の対応関係を出力することが可能な情報処理システムを提供することを課題の一とする。または、上記の何れか一又は複数の情報処理システムに用いられる情報処理方法を提供することを課題の一とする。または、上記の何れか一又は複数の情報処理システムを備える情報処理装置を提供することを課題の一とする。
 本発明の一態様は、利便性、有用性または信頼性に優れた新規な情報処理システムを提供することを課題の一とする。または、利便性、有用性または信頼性に優れた新規な情報処理方法を提供することを課題の一とする。または、新規な情報処理システム、新規な情報処理方法、または、新規な半導体装置を提供することを課題の一とする。
 なお、これらの課題の記載は、他の課題の存在を妨げるものではない。なお、本発明の一態様は、これらの課題の全てを解決する必要はないものとする。なお、これら以外の課題は、明細書、図面、請求項などの記載から、自ずと明らかとなるものであり、明細書、図面、請求項などの記載から、これら以外の課題を抽出することが可能である。
 本発明の一態様は、第1の情報処理装置及び第2の情報処理装置を有し、第1の情報処理装置は、第1の言語の第1の文書、及び、第1の文書の翻訳文である第2の言語の第2の文書の入力を受け付ける機能と、第2の文書に第1の文書の意訳が含まれるか、を判定させるための指示文と、第1の文書と、第2の文書と、を含むプロンプトを第2の情報処理装置に送信することで、判定の結果を取得する機能と、判定で意訳が含まれていないと判定された第2の文書に含まれる語句と第1の文書に含まれる語句との、対応関係を出力させるための指示文と、第1の文書と、第2の文書と、を含むプロンプトを第2の情報処理装置に送信することで、対応関係を取得する機能と、対応関係を出力する機能と、を有する、情報処理システムである。
 上記において、第1の情報処理装置は、第2の文書に意訳を含むことを許容するか、の選択を受け付ける機能と、選択で意訳を含むことを許容すると選択された第2の文書と、第1の文書と、の同義判定を行なわせるための指示文と、第1の文書と、第2の文書と、を含むプロンプトを第2の情報処理装置に送信することで、同義判定の結果を取得する機能を有することが好ましい。また、第1の情報処理装置は、第2の文書と第1の文書の同義判定に誤りがある場合に、第2の文書の修正を受け付ける機能を有することが好ましい。
 または、上記において、第1の情報処理装置は、意訳が含まれていると判定された第2の文書の修正を受け付け、意訳が含まれていないと判定されるまで、修正の受付と、判定のためのプロンプトの送信と、判定の結果の取得と、を繰り返す機能を有することが好ましい。
 または、上記において、第1の情報処理装置は、第2の文書に含まれる語句と第1の文書に含まれる語句の対応関係に誤りがある場合に、第2の文書の修正を受け付ける機能を有することが好ましい。
 または、本発明の一態様は、第1の情報処理装置及び第2の情報処理装置を有し、第1の情報処理装置は、第1の言語の第1の文書、及び、第1の文書の翻訳文である第2の言語の第2の文書の入力を受け付ける機能と、第1の文書の第2の言語の直訳文を生成する機能と、第1の文書の第2の言語の意訳文を生成する機能と、第2の文書と直訳文の類似度を算出する機能と、第2の文書と意訳文の類似度を算出する機能と、直訳文の類似度と、意訳文の類似度を比較し、直訳文の類似度の方が高い場合に、第2の文書に第1の文書の意訳が含まれていないと判定する機能と、判定で意訳が含まれていないと判定された第2の文書に含まれる語句と第1の文書に含まれる語句の、対応関係を出力させるための指示文と、第1の文書と、第2の文書と、を含むプロンプトを第2の情報処理装置に送信することで、対応関係を取得する機能と、を有する、情報処理システムである。
 または、本発明の一態様は、第1の言語の第1の文書、及び、第1の文書の翻訳文である第2の言語の第2の文書の入力を受け付けるステップと、第2の文書に第1の文書の意訳が含まれるか、を判定させるための指示文と、第1の文書と、第2の文書と、を含む第1のプロンプトを作成するステップと、第1のプロンプトに従い、判定を行なうステップと、判定で意訳が含まれていないと判定される場合に、第2の文書に含まれる語句と第1の文書に含まれる語句の、対応関係を出力させるための指示文と、第1の文書と、第2の文書と、を含む第2のプロンプトを作成するステップと、第2のプロンプトに従い、対応関係を出力するステップと、を有する、情報処理方法である。
 上記の情報処理方法において、第2の文書に意訳を含むことを許容するか、の選択を受け付けるステップと、選択で意訳を含むことを許容すると選択された第2の文書と、第1の文書と、の同義判定を行なわせるための指示文と、第1の文書と、第2の文書と、を含む第3のプロンプトを作成するステップと、第3のプロンプトに従い、同義判定を行なうステップと、を有することが好ましい。
 または、上記の情報処理方法において、判定は、大規模言語モデルによって行われ、対応関係の出力は、大規模言語モデルによって行われることが好ましい。また、判定は、大規模言語モデルによって行われ、対応関係の出力は、大規模言語モデルによって行われ、同義判定は、大規模言語モデルによって行われることが好ましい。
 本発明の一態様によれば、翻訳文のチェック及び修正する作業を効率的に行うことが可能な情報処理システムを提供することができる。または、本発明の一態様によれば、翻訳文に意訳が含まれているかどうかを判定することが可能な情報処理システムを提供することができる。または、本発明の一態様によれば、翻訳文に意訳が含まれる場合に、原文と翻訳文が同義であるかを判定することが可能な情報処理システムを提供することができる。または、本発明の一態様によれば、翻訳文が直訳である場合に、原文に含まれる語句と、当該語句に対応する翻訳文の語句と、の対応関係を出力することが可能な情報処理システムを提供することができる。または、上記の何れか一又は複数の情報処理システムに用いられる情報処理方法を提供することができる。または、上記の何れか一又は複数の情報処理システムを備える情報処理装置を提供することができる。
 本発明の一態様によれば、利便性、有用性または信頼性に優れた新規な情報処理システムを提供することができる。または、本発明の一態様によれば、利便性、有用性または信頼性に優れた新規な情報処理方法を提供することができる。また、新規な情報処理システムを提供することができる。または、新規な情報処理方法を提供することができる。
 なお、これらの効果の記載は、他の効果の存在を妨げるものではない。なお、本発明の一態様は、必ずしも、これらの効果の全てを有する必要はない。なお、これら以外の効果は、明細書、図面、請求項などの記載から、自ずと明らかとなるものであり、明細書、図面、請求項などの記載から、これら以外の効果を抽出することが可能である。
図1は、情報処理システムの構成を説明する図である。
図2は、情報処理方法を説明する図である。
図3は、情報処理方法を説明する図である。
図4は、情報処理方法を説明する図である。
図5は、情報処理システムの構成を説明する図である。
図6は、情報処理装置の構成を説明する図である。
図7Aは、情報処理方法を説明する図である。図7Bは、情報処理システムの構成を説明する図である。
 本発明の一態様の情報処理システムは、ある自然言語の原文と、別の自然言語に翻訳された翻訳文と、を比較し、翻訳が正しく行われているかをチェックする際に用いることができる情報処理システムである。別言すると、本発明の一態様の情報処理システムは、翻訳文をチェックする翻訳者の手助けを行うことができる情報処理システムである。なお、上記の翻訳文は、機械翻訳による翻訳文に限らず、人手翻訳(人間翻訳ともいう)による翻訳文も含まれる。
 また、本発明の一態様の情報処理システムは、直訳と意訳に柔軟に対応することが可能なチェック機能を備える。つまり、直訳であることが求められる場合と、意訳を含むことが許容される場合と、で異なる情報処理を行う機能を有する。なお、本発明の一態様の情報処理システムは、直訳であることが求められる場合と、意訳を含むことが許容される場合と、の双方に対応する機能を有する情報処理システムであることが好ましいが、いずれか一方に対応する機能を有する情報処理システムであってもよい。
 本発明の一態様の情報処理システムを用いると、直訳であることが求められる場合(意訳を含むことを許容しない場合)において、翻訳文に意訳された箇所が含まれる場合に、当該意訳箇所を検出し、情報処理システムの使用者に修正を促すことができる。また、本発明の一態様の情報処理システムを用いることにより、意訳された箇所が含まれないと判定された翻訳文(または、意訳された箇所が含まれないと判定されるまで修正した翻訳文)と、原文と、を比較し、それぞれの文書に含まれる語句の対応関係を、情報処理システムの使用者に提示することで、情報抜け、情報過多などの誤訳を見つけ易くすることができ、当該誤訳を効率的に修正することができる。
 または、本発明の一態様の情報処理システムを用いると、意訳を含むことが許容される場合において、翻訳文が、原文と同義であるかの判定を行うことができる。また、同義判定の結果を、情報処理システムの使用者に提示することで、同義でない翻訳文(誤訳)を見落とすことを防ぐことができ、当該誤訳を効率的に修正することができる。
 一例として、本発明の一態様の情報処理システムは、直訳であることが求められる場合において、第1の言語の文書(原文)、及び当該原文が第2の言語に翻訳された第2の言語の文書(翻訳文)の入力を受け付ける機能を有する。また、本発明の一態様の情報処理システムは、入力された翻訳文に、原文の文書の意訳が含まれるか、を判定する機能を有し、この判定において意訳が含まれていない場合(翻訳文が直訳である場合)には、原文に含まれる語句と、翻訳文に含まれる語句の対応関係を出力する機能を有する。当該情報処理システムの使用者は、上記の対応関係を確認することで、翻訳が正しく行われているか、をチェックすることが可能である。また、当該情報処理システムの使用者は、入力された翻訳文に意訳が含まれると判定される場合に、翻訳文を修正することが可能である。また、当該情報処理システムの使用者は、上記の対応関係に誤りまたは差異が含まれている場合に、翻訳文を修正することが可能である。
 また、本発明の一態様の情報処理システムは、意訳を含むことが許容される場合において、第1の言語の文書(原文)、及び当該原文が第2の言語に翻訳された第2の言語の文書(翻訳文)の入力を受け付ける機能を有する。また、翻訳文が、原文と同義であるかを判定(同義判定)する機能を有する。また、上記の同義判定の結果を出力する機能を有する。当該情報処理システムの使用者は、上記の同義判定の結果を確認することで、翻訳が正しく行われているか、をチェックすることが可能である。また、当該情報処理システムの使用者は、入力された翻訳文に、原文と同義でない箇所が含まれている場合に、翻訳文を修正することが可能である。
 本発明の一態様の情報処理システムは、上記の入力から出力の一連の処理を、1文毎、2文以上の複数文毎、1段落毎、2段落以上の複数段落毎、などの単位で繰り返し行う機能を有する。使用者は、原文と、翻訳文と、語句の対応関係、又は同義判定の結果と、をセットとして確認し、必要に応じて修正することが可能であるが、この確認及び選択は、上記の繰り返し単位毎に行ってもよいし、複数回の繰り返し単位毎に行ってもよいし、全体の処理が完了してから行ってもよい。
 なお、本明細書にて用いる第1、第2、第3といった序数を用いた用語は、構成要素を識別するために便宜上付したものであり、その数を限定するものではない。
 実施の形態について、図面を用いて詳細に説明する。但し、本発明は以下の説明に限定されず、本発明の趣旨及びその範囲から逸脱することなくその形態及び詳細を様々に変更し得ることは当業者であれば容易に理解される。従って、本発明は以下に示す実施の形態の記載内容に限定して解釈されるものではない。なお、以下に説明する発明の構成において、同一部分又は同様な機能を有する部分には同一の符号を異なる図面間で共通して用い、その繰り返しの説明は省略する。
 本明細書に添付した図面では、構成要素を機能ごとに分類し、互いに独立したブロックとしてブロック図を示しているが、実際の構成要素は機能ごとに完全に切り分けることが難しく、一つの構成要素が複数の機能に係わることもあり得る。
(実施の形態1)
 本実施の形態では、本発明の一態様の情報処理システム、及び情報処理方法について、図1乃至図6を参照しながら説明する。
 図1は、本発明の一態様の情報処理システムの構成を説明する図である。
<情報処理システムの構成例1>
 本実施の形態で説明する情報処理システムは、図1に示すように、情報処理装置10と、情報処理装置40と、を有する。情報処理装置10と、情報処理装置40と、は、ネットワーク30を介して接続され、文書データ等の送信及び受信を行うことができる。
 また、本実施の形態で説明する情報処理システムは、使用者が、情報処理装置10を直接操作して文書等の入力を行うことができる構成であってもよいし、図1に示すように、情報処理装置10と、ネットワーク31を介して接続される情報端末20を用いて、文書等の入力を行うことができる構成であってもよい。
 本実施の形態で説明する情報処理システムで用いられる情報処理方法の一例を、図2乃至図4を用いて説明する。
[情報処理方法]
 図2乃至図4は、本発明の一態様の情報処理方法の一例を説明するフロー図である。
 本発明の一態様の情報処理方法が「開始」され、図2のステップS101において、第1の言語で表現された第1の文書(原文)と、第1の文書が第2の言語へと翻訳された第2の文書(翻訳文)が入力される。
 次に、図2のステップS111において、翻訳文(第2の文書)に意訳を含んでよいか、の選択がおこなわれる。なお、ここでの選択は、本情報処理システムの使用者が、都度選択してもよいし、ステップS101において、原文(第1の文書)及び翻訳文(第2の文書)と共に、予め入力をしておいてもよい。このようにして、翻訳文(第2の文書)に意訳を含むことを許容するか、の選択を受け付ける。
 図2のステップS111において、翻訳文(第2の文書)に意訳を含んでよい、と選択された場合(意訳を含むことが許容される場合)は、ステップS111から図4に示すフローチャートへ結合する結合子Aに進む。
 または、図2のステップS111において、翻訳文(第2の文書)に意訳を含むことを許容しない、と選択された場合(直訳であることが求められる場合)は、ステップS121に進む。
<直訳であることが求められる場合のフロー>
 次に、図2のステップS121において、翻訳文(第2の文書)に意訳が含まれているかの判定を行なう。
 図3を用いて、図2のステップS121における処理の一例として、ステップS201乃至ステップS204の処理を有するステップS121Aについて説明する。
 ステップS121Aの最初の処理として、図3のステップS201において、翻訳文(第2の文書)に、原文(第1の文書)の意訳が含まれるかを判定するためのプロンプトを作成する。
 意訳が含まれるかを判定するためのプロンプトは、原文(第1の文書)と、翻訳文(第2の文書)と、「以下の原文と翻訳文について、意訳されている箇所を指摘してください。」といった指示文と、を含む。なお、プロンプトとは、言語モデルに所望の動作をさせるための入力文といえる。言語モデルは、プロンプトが与えられると、当該プロンプトに基づいて、応答内容を生成する。
 ステップS201で作成したプロンプトを、ステップS202において大規模言語モデルに送信することで、ステップS203において、翻訳文(第2の文書)に、原文(第1の文書)の意訳が含まれるかを、大規模言語モデルに判定させることができる。その後、ステップS204において、大規模言語モデルから出力される判定結果を受信し、図2のステップS122に進むことができる。
 このようにして、情報処理システムは翻訳文(第2の文書)に意訳が含まれているかの判定を行なうことができる。
 次に、図2のステップS122において、ステップS121で取得した判定結果を、情報処理システムの使用者に提示する。翻訳文(第2の文書)に意訳が含まれる、と判定された場合は、意訳されていると判定された箇所を使用者に提示することができる。ステップS123において、使用者は、翻訳文(第2の文書)の意訳されていると判定された箇所を修正する。情報処理システムは修正された翻訳文(第2の文書)を受け付ける。なお、上記修正は、翻訳文(第2の文書またはステップS123で修正された第2の文書)に意訳が含まれないと判定されるまで、繰り返し行うことができる。
 ステップS121において、意訳を含まないと判定された場合には、ステップS131に進む。
 次に、図2のステップS131において、原文(第1の文書)と、翻訳文(第2の文書またはステップS123で修正された第2の文書)と、の対応関係を判定(対応判定)するためのプロンプトを作成する。
 対応判定のプロンプトは、原文(第1の文書)と、翻訳文(第2の文書またはステップS123で修正された第2の文書)と、「以下の原文と翻訳文は直訳されています。それぞれの語句について対応表を作成してください。」といった指示文と、を含む。このプロンプトを大規模言語モデルに送信することで、原文(第1の文書)と、翻訳文(第2の文書またはステップS123で修正された第2の文書)と、の対応関係を、大規模言語モデルから出力させることができる。
 次に、図2のステップS132において、大規模言語モデルは、原文(第1の文書)と、翻訳文(第2の文書またはステップS123で修正された第2の文書)と、の対応関係を判定する。また、大規模言語モデルは、当該対応関係の判定結果として、例えば対応関係の表を出力する。
 また、原文(第1の文書)と、翻訳文(第2の文書またはステップS123で修正された第2の文書)と、の対応関係に差異が含まれる場合に、その際の内容を説明するように、指示文を作成してもよい。この場合では、例えばプロンプトに含まれる指示文を「以下の原文と翻訳文は直訳されています。それぞれの語句について対応表を作成してください。そのあと、対応表から、対応関係に差異があれば報告してください。」といった内容にするとよい。
 上記プロンプトと、その結果の例を表1乃至表9に示す。表1乃至表3は、原文と翻訳文の対応関係に差異が含まれない場合の例であり、表4乃至表6は、原文と翻訳文の対応関係に差異が含まれる場合の第1の例であり、表7乃至表9は、原文と翻訳文の対応関係に差異が含まれる場合の第2の例である。なお、上記第1の例では、翻訳文に不足(a mouse)があり、上記第2の例では、翻訳文の数字(符号)に誤記(正:101、誤:102)がある。
 表1、表4及び表7は大規模言語モデルに入力するプロンプトの例であり、表2、表5及び表8は大規模言語モデルから出力される対応関係の表の例であり、表3、表6及び表9は大規模言語モデルから出力される対応関係の説明の例である。
Figure JPOXMLDOC01-appb-T000001
Figure JPOXMLDOC01-appb-T000002
Figure JPOXMLDOC01-appb-T000003
Figure JPOXMLDOC01-appb-T000004
Figure JPOXMLDOC01-appb-T000005
Figure JPOXMLDOC01-appb-T000006
Figure JPOXMLDOC01-appb-T000007
Figure JPOXMLDOC01-appb-T000008
Figure JPOXMLDOC01-appb-T000009
 次に、図2のステップS133において、翻訳文(第2の文書またはステップS123で修正された第2の文書)を修正するか、を選択する。
 ステップS133の選択として、表1乃至表3の例のように、ステップS132で大規模言語モデルから出力された対応関係に、差異が含まれない場合は、本発明の一態様の情報処理方法を「終了」することができる。または、表4乃至表9の例のように、ステップS133の選択として、ステップS132で大規模言語モデルから出力された対応関係に、差異が含まれる場合は、ステップS134に進み、ステップS134において、翻訳文を修正する。なお、上記修正は、原文と翻訳文との対応関係に差異が含まれないと判定されるまで、繰り返し行うことができる。
 なお、ステップS133の選択は、ステップS132で大規模言語モデルから出力された対応関係を参照して、本情報処理システムの使用者が、選択することもできる。
<意訳を含むことが許容される場合のフロー>
 ステップS111において、翻訳文(第2の文書)に意訳を含んでよい、と選択された場合(意訳を含むことが許容される場合)について、図4を用いて説明する。
 図4に示す結合子Aから接続されるステップS141において、原文(第1の文書)と、翻訳文(第2の文書)と、の同義判定を行なうためのプロンプトを作成する。
 同義判定のプロンプトは、原文(第1の文書)と、翻訳文(第2の文書)と、「以下の原文と翻訳文について、意訳されている箇所を表にまとめて指摘してください。そのあと、原文と翻訳文の意味が同じであるかを、同義である、同義でない、で答えてください。」といった指示文と、を含む。このプロンプトを大規模言語モデルに送信することで、原文(第1の文書)と、翻訳文(第2の文書)と、が同義であるかを、大規模言語モデルに判定させ、その判定結果を出力させることができる。
 同義判定としては、機械学習モデルを用いることができる。機械学習モデルとして、ニューラルネットワークモデルであるBERT(Bidirectional Encoder Representation from Transformers)をファインチューニングして用いることができる。
 機械学習モデルを同義判定に用いる場合は、含意関係認識タスクを用いて、含意である、含意でない、を、同義である、同義でないと置き換えて用いてもよい。例えば、二つの文を比較して、情報が同等であるものを同義である、情報が同等でないものを同義でないとラベル付けして学習すればよい。また、例えば、第1の文書と第2の文書を比較して、第1の文書に対して、第2の文書の情報が、過少である、過多である、同等である(同義である)、矛盾する、のラベルを付けて学習すればよい。
 または、同義判定として、文ベクトルの比較を行い、ベクトルのコサイン類似度で評価してもよい。類似度が一定値以上を同義、一定値未満を同義でないとしてもよい。
 または、同義判定として、LLMを用いることもできる。LLMを同義判定に用いる場合は、同義度の数値ではなく、第1の文書と第2の文書が、同義であるか、同義でないか、を2択で出力する方法としてもよい。また、LLMによる同義判定において、第1の文書に含まれる情報に対して第2の文書に含まれる情報が過多である場合、または欠落している場合に、その情報を付記して判定とともに出力する方法にしてもよい。なお、LLMによる同義判定において、第1の文書と第2の文書を比較して、第1の文書に対して、第2の文書の情報が、過少である、過多である、同等である(同義である)、矛盾する、のラベルを付けて学習すればよい。
 または、同義判定として、ベクトルのコサイン類似度と同義判定ラベルを組み合わせて用いてもよい。例えば、コサイン類似度をコサイン距離に変換し、過多ラベルを1、過小ラベルを−1、同義ラベルを0、矛盾ラベルをnanとして、ラベル評価値として、コサイン距離×ラベル評価値を算出してもよい。この場合、0に近いほど同義であって、情報の過不足も表すことができる。原文にない記載を余計に追加していることを特に嫌う場合など情報過不足を評価できるため、単に類似度を表すよりも効果的である。
 次に、図4のステップS143において、ステップS142で取得した同義判定の結果を、情報処理システムの使用者に提示し、使用者は翻訳文(第2の文書)を修正するか、を選択する。また、情報処理システムは、当該選択を受け付ける。
 ステップS143の選択として、ステップS142で大規模言語モデルから出力された同義判定に、同義ではない、の判定結果が含まれない場合は、本発明の一態様の情報処理方法を「終了」することができる。または、ステップS143の選択として、ステップS142で大規模言語モデルから出力された同義判定の結果に、同義ではない、の判定結果が含まれる場合は、同義ではないと判定された箇所を使用者に提示することができ、ステップS144において、使用者は、翻訳文(第2の文書)の同義ではないと判定された箇所を修正する。情報処理システムは修正された翻訳文(第2の文書)を受け付ける。なお、上記修正は、原文と翻訳文との同義判定の結果に、同義ではないと判定される結果が含まれなくなるまで、繰り返し行うことができる。
 なお、ステップS143の選択は、ステップS142で大規模言語モデルから出力された同義判定の結果を参照して、本情報処理システムの使用者が、選択することもできる。
 なお、図2及び図4に示すフローにおいて、「終了」とする前に、続けて次の原文の入力処理に進むことができる。全ての原文の処理が終わるまで、上記の処理を繰り返すことができる。
 図2乃至図4を用いて説明した本発明の一態様の情報処理方法において、各ステップの処理を、情報処理装置10、情報処理装置40、及び情報端末20のいずれの装置又は端末で行うかを、図5を用いて説明する。
 図5は、情報処理装置10、情報処理装置40、及び情報端末20で行われる処理を、図2乃至図4で用いた各ステップの符号を用いて示した関係図である。また、ネットワーク30を付したブロック矢印は、情報処理装置10と情報処理装置40の間のデータの送受信を示しており、また、ネットワーク31を付したブロック矢印は、情報処理装置10と情報端末20の間のデータの送受信を示している。
 図5に示すように、情報処理装置10は、ステップS201の処理を行う機能と、ステップS202の処理を行う機能と、ステップS204の処理を行う機能と、ステップS122の処理を行う機能と、ステップS131の処理を行う機能と、ステップS141の処理を行う機能と、を有する。
 また、図5に示すように、情報処理装置40は、ステップS203の処理を行う機能と、ステップS132の処理を行う機能と、ステップS142の処理を行う機能と、を有する。
 また、図5に示すように、情報端末20は、ステップS101の処理を行う機能と、ステップS111の処理を行う機能と、ステップS123の処理を行う機能と、ステップS133の処理を行う機能と、ステップS134の処理を行う機能と、ステップS143の処理を行う機能と、ステップS144の処理を行う機能と、を有する。
 情報処理装置10、情報処理装置40、情報端末20、ネットワーク30、及びネットワーク31の構成例を以下で説明する。
《情報処理装置10の構成例》
 本発明の一態様の情報処理システムが有する情報処理装置10の構成例について、図6を用いて説明する。
 図6に示すように、情報処理装置10は、入力部110、記憶部120、処理部130、出力部140、及び、伝送路150を有する。
[入力部110]
 入力部110は、情報処理装置10の外部からデータを受け付けることができる。例えば、入力部110は、情報端末20からデータを受け付けることができる。また、入力部110は、情報処理装置40からデータを受け付けることができる。具体的には、有線通信ポート、無線通信ポート、又は光通信ポート等のデバイスを用いることができる。
 入力部110は受け付けたデータを、伝送路150を介して、記憶部120及び処理部130の一方または双方に供給することができる。
[記憶部120]
 記憶部120は、処理部130が実行するプログラムを記憶する機能を有する。また、記憶部120は、処理部130が生成したデータ(例えば、演算結果、解析結果、推論結果)、及び、入力部110が受け付けたデータなどを記憶する機能を有していてもよい。
 記憶部120は、データベースを有していてもよい。また、情報処理装置10は、記憶部120とは別にデータベースを有していてもよい。情報処理装置10は、記憶部120の外部、情報処理装置10の外部、または情報処理システムの外部に存在するデータベースから、データを取り出す機能を有していてもよい。また、情報処理装置10は、自身が持つデータベースと、外部に存在するデータベースと、の双方からデータを取り出す機能を有していてもよい。
 ストレージ及びファイルサーバの一方または双方を記憶部120に用いることができる。また、ファイルサーバに保存されたファイルのパスを記録したデータベースを記憶部120に用いることができる。
 記憶部120は、揮発性メモリ及び不揮発性メモリのうち少なくとも一方を有する。揮発性メモリとしては、DRAM(Dynamic Random Access Memory)、及び、SRAM(Static Random Access Memory)等が挙げられる。不揮発性メモリとしては、ReRAM(Resistive Random Access Memory、抵抗変化型メモリともいう)、PRAM(Phase change Random Access Memory)、FeRAM(Ferroelectric Random Access Memory)、MRAM(Magnetoresistive Random Access Memory、磁気抵抗型メモリともいう)、及び、フラッシュメモリ等が挙げられる。また、記憶部120は、NOSRAM(登録商標)及びDOSRAM(登録商標)のうち少なくとも一方を有していてもよい。また、記憶部120は、記録メディアドライブを有していてもよい。記録メディアドライブとしては、ハードディスクドライブ(Hard Disk Drive:HDD)、及び、ソリッドステートドライブ(Solid State Drive:SSD)等が挙げられる。
 NOSRAMとは、「Nonvolatile Oxide Semiconductor Random Access Memory」の略称である。NOSRAMは、メモリセルが2トランジスタ型(2T)、又は3トランジスタ型(3T)ゲインセルであり、トランジスタが、金属酸化物をチャネル形成領域に用いたトランジスタ(OSトランジスタともいう)であるメモリのことをいう。OSトランジスタはオフ状態でソースとドレインとの間を流れる電流、つまりリーク電流が極めて小さい。NOSRAMは、リーク電流が極めて小さい特性を用いてデータに応じた電荷をメモリセル内に保持することで、不揮発性メモリとして用いることができる。特にNOSRAMは保持しているデータを破壊することなく読み出しすること(非破壊読み出し)が可能なため、データ読み出し動作のみを大量に繰り返す、演算処理に適している。NOSRAMは、積層して設けることでデータ容量を大きくできるため、大規模なキャッシュメモリ、メインメモリ、ストレージメモリとして用いることで半導体装置の高性能化を図ることができる。
 DOSRAMとは、「Dynamic Oxide Semiconductor RAM」の略称であり、1T(トランジスタ)1C(容量)型のメモリセルを有するRAMを指す。DOSRAMは、OSトランジスタを用いて形成されたDRAMであり、DOSRAMは、外部から送られてくる情報を一時的に格納するメモリである。DOSRAMは、OSトランジスタのオフ電流が小さいことを利用したメモリである。
 本明細書等において、金属酸化物(metal oxide)とは、広い意味での金属の酸化物である。金属酸化物は、酸化物絶縁体、酸化物導電体(透明酸化物導電体を含む)、酸化物半導体(Oxide Semiconductorまたは単にOSともいう)などに分類される。例えば、トランジスタの半導体層に金属酸化物を用いた場合、当該金属酸化物を酸化物半導体と呼称する場合がある。
 チャネル形成領域が有する金属酸化物はインジウム(In)を含むことが好ましい。チャネル形成領域が有する金属酸化物がインジウムを含む金属酸化物の場合、OSトランジスタのキャリア移動度(電子移動度)が高くなる。また、チャネル形成領域が有する金属酸化物は、元素Mを含む酸化物半導体であると好ましい。元素Mは、アルミニウム(Al)、ガリウム(Ga)及びスズ(Sn)の少なくとも1つであることが好ましい。元素Mに適用可能なその他の元素としては、ホウ素(B)、シリコン(Si)、チタン(Ti)、鉄(Fe)、ニッケル(Ni)、ゲルマニウム(Ge)、イットリウム(Y)、ジルコニウム(Zr)、モリブデン(Mo)、ランタン(La)、セリウム(Ce)、ネオジム(Nd)、ハフニウム(Hf)、タンタル(Ta)、及び、タングステン(W)などが挙げられる。ただし、元素Mとして、前述の元素を複数組み合わせても構わない場合がある。元素Mは、例えば、酸素との結合エネルギーが高い元素である。例えば、酸素との結合エネルギーがインジウムよりも高い元素である。また、チャネル形成領域が有する金属酸化物は、亜鉛(Zn)を含む金属酸化物であると好ましい。亜鉛を含む金属酸化物は結晶化しやすくなる場合がある。
 チャネル形成領域が有する金属酸化物は、インジウムを含む金属酸化物に限定されない。チャネル形成領域が有する金属酸化物は、例えば、亜鉛スズ酸化物、ガリウムスズ酸化物などの、インジウムを含まず、亜鉛を含む金属酸化物、ガリウムを含む金属酸化物、スズを含む金属酸化物などであっても構わない。
[処理部130]
 処理部130は、入力部110及び記憶部120の一方または双方から供給されたデータを用いて、演算、解析、及び推論などの処理を行う機能を有する。処理部130は、生成したデータ(例えば、演算結果、解析結果、推論結果)を、記憶部120及び出力部140の一方または双方に供給することができる。
 処理部130は、記憶部120からデータを取得する機能を有する。また、処理部130は、記憶部120に、データを記録する又は登録する機能を有してもよい。
 処理部130は、例えば、演算回路を有することができる。処理部130は、例えば、中央演算装置(CPU:Central Processing Unit)を有することができる。また、処理部130は、GPU(Graphics Processing Unit)を有することができる。
 処理部130は、DSP(Digital Signal Processor)等のマイクロプロセッサを有していてもよい。マイクロプロセッサは、FPGA(Field Programmable Gate Array)、FPAA(Field Programmable Analog Array)等のPLD(Programmable Logic Device)によって実現された構成であってもよい。また、処理部130は、量子プロセッサを有していてもよい。処理部130は、プロセッサにより種々のプログラムからの命令を解釈し実行することで、各種のデータ処理及びプログラム制御を行うことができる。プロセッサにより実行しうるプログラムは、プロセッサが有するメモリ領域及び記憶部120のうち少なくとも一方に格納される。
 処理部130はメインメモリを有していてもよい。メインメモリは、RAM(Random Access Memory)等の揮発性メモリ、及びROM(Read Only Memory)等の不揮発性メモリのうち少なくとも一方を有する。また、メインメモリは、上述したNOSRAM及びDOSRAMのうち少なくとも一方を有していてもよい。
 RAMとしては、例えばDRAM、SRAM等が用いられ、処理部130の作業空間として仮想的にメモリ空間が割り当てられ利用される。記憶部120に格納されたオペレーティングシステム、アプリケーションプログラム、プログラムモジュール、プログラムデータ、及びルックアップテーブル等は、実行のためにRAMにロードされる。RAMにロードされたこれらのデータ、プログラム、及びプログラムモジュールは、それぞれ、処理部130に直接アクセスされ、操作される。
 ROMには、書き換えを必要としない、BIOS(Basic Input/Output System)及びファームウェア等を格納することができる。ROMとしては、マスクROM、OTPROM(One Time Programmable Read Only Memory)、EPROM(Erasable Programmable Read Only Memory)等が挙げられる。EPROMとしては、紫外線照射により記憶データの消去を可能とするUV−EPROM(Ultra−Violet Erasable Programmable Read Only Memory)、EEPROM(Electrically Erasable Programmable Read Only Memory)、フラッシュメモリ等が挙げられる。
 処理部130は、OSトランジスタ、及び、チャネル形成領域にシリコンを有するトランジスタ(Siトランジスタ)の一方または双方を有することができる。
 処理部130は、OSトランジスタを有することが好ましい。OSトランジスタはオフ電流が極めて小さいため、OSトランジスタを記憶素子として機能する容量素子に流入した電荷(データ)を保持するためのスイッチとして用いることで、データの保持期間を長期にわたり確保することができる。この特性を、処理部が有するレジスタ及びキャッシュメモリのうち少なくとも一方に用いることで、必要なときだけ処理部を動作させ、他の場合には直前の処理の情報を当該記憶素子に待避させることにより処理部をオフにすることができる。すなわち、ノーマリーオフコンピューティングが可能となり、情報処理システムの低消費電力化を図ることができる。
 情報処理装置10は、少なくとも一部の処理にAIを用いることが好ましい。
 情報処理装置10は、特に、人工ニューラルネットワークを用いることが好ましい。ニューラルネットワークは、回路(ハードウェア)またはプログラム(ソフトウェア)により実現される。
 本明細書等において、ニューラルネットワークとは、生物の神経回路網を模し、学習によってニューロン同士の結合強度を決定し、問題解決能力を持たせるモデル全般を指す。ニューラルネットワークは、入力層、中間層(隠れ層)、及び出力層を有する。
 本明細書等において、ニューラルネットワークについて述べる際に、既にある情報からニューロンとニューロンの結合強度(重み係数ともいう)を決定することを「学習」と呼ぶ場合がある。
 本明細書等において、学習によって得られた結合強度を用いてニューラルネットワークを構成し、そこから新たな結論を導くことを「推論」と呼ぶ場合がある。
 情報処理装置10は、AIを用いた自然言語処理モデルを用いた処理を行うことができる。例えば、AIを用いた自然言語処理モデルを用いて機械翻訳の処理を行うことができる。機械翻訳の処理を行う自然言語処理モデルとして、Sequence−to−Sequence(seq2seq)、Transformer、BERT(Bidirectional Encoder Representations from Transformers)、T5(Text−to−Text Transfer Transformer)などのAIを用いた自然言語処理モデルを用いた処理を実行することができる。
[出力部140]
 出力部140は、処理部130における演算結果、解析結果、及び推論結果の少なくとも1つを、情報処理装置10の外部に出力することができる。具体的には、有線通信ポート、無線通信ポート、又は光通信ポート等のデバイスを用いることができる。
 例えば、出力部140は、情報処理装置40にデータを送信することができる。また、出力部140は、情報端末20にデータを送信することができる。
[伝送路150]
 伝送路150は、データを伝達する機能を有する。入力部110、記憶部120、処理部130、及び、出力部140の間のデータの送受信は、伝送路150を介して行うことができる。具体的には、マザーボード上のバスライン、有線通信ケーブル、または光通信ケーブルを用いることができる。
《情報処理装置40の構成例》
 情報処理装置40は、受信したデータを処理し、処理の結果を送信することができる。例えば、情報処理装置10から受信したデータを用いて、演算などの処理を行うことができる。また、情報処理装置40は処理の結果を、情報処理装置10に送信することができる。これにより、情報処理装置10における演算の負担を低減することができる。
 情報処理装置40は、AIを用いた自然言語処理モデルを用いた処理を行うことができる。例えば、BERT(Bidirectional Encoder Representations from Transformers)、T5(Text−to−Text Transfer Transformer)などのAIを用いた自然言語処理モデルを用いた処理を実行することができる。
 また、情報処理装置40は、大規模言語モデルを利用したモデル(文章生成モデル、対話モデルなど)を用いた処理を行うことができる。図2で説明したステップS132における、原文に含まれる語句と、翻訳文に含まれる語句との対応関係の判定は、大規模言語モデルを利用したモデルを用いて処理することが好ましい。また、図3で説明したステップS203における、翻訳文に原文の意訳が含まれているか、の判定は、大規模言語モデルを利用したモデルを用いて処理することが好ましい。また、図4で説明したステップS142おける、原文と翻訳文の同義判定は、大規模言語モデルを利用したモデルを用いて処理することが好ましい。例えば、GPT−3、GPT−3.5、GPT−4(登録商標)、LaMDA(Language Model for Dialogue Applications)、PaLM(Pathways Language Model)、Llama2、Llama3などの大規模言語モデルを用いて、処理を実行することができる。特に、GPT−4(登録商標)を用いることが好ましい。
 例えば、情報処理装置40は、言語Aで記述された文書を言語Bで記述された文書に翻訳することができる。また、情報処理装置40は、与えた指示文に従って、言語Aで記述された文書を言語Bで記述された文書に翻訳することができる。また、束縛条件を指示文に含めることができる。これにより、束縛条件を与えて翻訳の自由度を制御することができる。
 また、情報処理装置40は、様々な自然言語処理タスクを行うことができる汎用言語処理モデルを用いた処理を実行することができる。
 情報処理装置40は、サーバコンピュータ、スーパーコンピュータなどの大型のコンピュータである。また、情報処理装置40は並列計算機としての機能を有することが好ましい。情報処理装置40を並列計算機として用いることで、例えば、AIの学習及び推論に必要な大規模の計算を行うことができる。
 なお、情報処理装置40は、情報処理装置10と比較して、処理能力が高いコンピュータである。例えば、情報処理装置10および情報処理装置40の双方が並列計算機としての機能を有する場合、情報処理装置40は、情報処理装置10に比べて処理能力が高く、大規模な計算を行うことができる。また、例えば、情報処理装置10および情報処理装置40の双方が、大規模言語モデルを利用したモデルを用いた処理を行うことができる場合、情報処理装置40は、情報処理装置10に比べて、大規模なAIモデルを用いた処理を実行することができる。
 なお、サービスの提供者は、必ずしも、情報処理装置40を自前で所有する必要はない。例えば、サービスの提供者は、他の事業者等が情報処理装置40を用いて提供するサービスの一部を利用することができる。
《情報端末20の構成例》
 情報端末20は、本発明の一態様の情報処理システムのユーザが入力するデータを受け付けることができる。また、情報端末20は、本発明の一態様の情報処理システムが出力するデータを、ユーザに提供することができる。
 また、情報端末20は、ユーザから受け付けたデータを、情報処理装置10に送信することができる。また、情報端末20は、情報処理装置10から受信したデータを、ユーザに提供することができる。
 また、情報端末20は、ユーザから受け付けたデータを元に生成したデータを、情報処理装置10に送信することができる。また、情報端末20は、情報処理装置10から受信したデータを元に生成したデータを、ユーザに提供することができる。
 情報端末20には、例えば、専用のアプリケーションソフトウェアまたはウェブブラウザなどがインストールされる。ユーザは、いずれかを介して、情報処理装置10にアクセスすることができる。これにより、ユーザは、例えば、情報処理装置10に比べて処理能力が低いコンピュータを用いて、本発明の一態様の情報処理システムを用いたサービスを、享受することができる。
 情報端末20はクライアントコンピュータなどと呼ぶこともできる。情報端末20は、いずれも、本発明の一態様の情報処理システムのユーザが使用する情報端末装置である。
 例えば、デスクトップ型コンピュータ20a、ノート型コンピュータ20b、スマートフォン20c、タブレット型コンピュータ20dを、情報端末20に用いることができる。なお、タブレット型コンピュータ20dは、キーボードを有する筐体21と接続することで、ノート型コンピュータとして用いることもできる。
 これにより、情報処理システムの使用者は、例えば、情報端末20のような、情報処理装置10または情報処理装置40に比べて処理能力が低いコンピュータ等を用いて、本発明の一態様の情報処理システムを制御し、情報処理装置40に指示を与え、サービスを享受することができる。
《ネットワーク30》
 ネットワーク30は、情報処理装置10および情報処理装置40を接続する。これにより、両者の間で、入力されたデータおよび処理されたデータの送受信が可能になる。また、情報処理に係る負荷を分散することができる。
 なお、本実施の形態では、主に、ネットワーク30は、ネットワーク31よりも大規模なコンピュータネットワークである場合を説明する。例えば、グローバルネットワークを、ネットワーク30に用いることができる。具体的には、World Wide Web(WWW)の基盤であるインターネットを用いることができる。
《ネットワーク31》
 ネットワーク31は、複数の情報端末20および情報処理装置10を接続する。これにより、両者の間で、データの送受信が可能になる。また、情報処理に係る負荷を分散することができる。また、サービスの提供者は、例えば、ネットワーク31を介して、本発明の一態様の情報処理方法を用いたサービスを、ユーザに提供することができる。
 例えば、ローカルネットワークを、ネットワーク31に用いることができる。また、イントラネット又はエクストラネットを、ネットワーク31に用いることができる。また、PAN(Personal Area Network)、LAN(Local Area Network)、CAN(Campus Area Network)、MAN(Metropolitan Area Network)、WAN(Wide Area Network)、GAN(Global Area Network)等を、ネットワーク31に用いることができる。
 本発明の一態様の情報処理方法を用いたサービスの提供者と、当該サービスを享受するユーザとが、同じ企業等の組織内に属する場合、情報端末20と情報処理装置10の間で行われるデータの送受信は、例えば、当該組織内に構築されたネットワーク31を用いて、行われることが好ましい。これにより、インターネットを介して行う場合に比べて安全に、情報端末20と情報処理装置40の間でデータを送受信することができる。また、組織内の機密情報の外部への流出を防止することができる。
 なお、無線通信を行う場合、通信プロトコル又は通信技術として、第4世代移動通信システム(4G)、第5世代移動通信システム(5G)、第6世代移動通信システム(6G)などの通信規格、または、Wi−Fi(登録商標)、Bluetooth(登録商標)等のIEEEにより通信規格化された仕様を用いることができる。
(実施の形態2)
 本実施の形態では、本発明の一態様の情報処理システム、及び情報処理方法について、実施の形態1に記載した情報処理システムの構成例1と、一部の構成が異なる情報処理システムの構成例について説明する。
<情報処理システムの構成例2>
 図7A及び図7Bを用いて、上記の情報処理システムの構成例1とは別の情報処理システムの一例について説明する。ここでは、異なる部分について詳細に説明し、同じ構成を備える部分については、実施の形態1の説明を参照する。
 ここで説明する情報処理システムは、図2のステップS121の処理内容が、情報処理システムの構成例1の図3で説明したフローとは異なる。具体的には、図2のステップS121の処理内容を、図7Aに示すステップS301乃至ステップS304の処理を有するステップS121Bを用いて行う。図7Aを用いて、ステップS301乃至ステップS304の処理について説明する。
 ステップS121Bの最初の処理として、図7AのステップS301において、第1の文書(原文)を、第2の言語へと直訳で翻訳し、第3の文書(直訳文)を生成する。
 次に、図7AのステップS302において、第1の文書(原文)を、第2の言語へと意訳を含むように翻訳し、第4の文書(意訳文)を生成する。なお、ステップS301の処理とステップS302の処理は、先にステップS302の処理を行って、その後にステップS301の処理を行ってもよいし、ステップS301及びステップS302を並行して処理を行ってもよい。
 直訳文の生成(ステップS301)には、直訳から構成されるデータセットで学習した専用の翻訳モデルを用い、意訳文の生成(ステップS302)には、意訳から構成されるデータセットで学習した専用の翻訳モデルを用いるとよい。そうすることで、プロンプトの作成なども不要で計算量を抑制しつつ、異なるモデルによる生成であるため、直訳と意訳で異なる文を生成できる。翻訳モデルとしては、実施の形態1で説明した自然言語処理モデルを用いることができる。
 次に、図7AのステップS303において、図2のステップS101で入力された第2の文書(翻訳文)と、第3の文書(直訳文)及び第4の文書(意訳文)のそれぞれと、の類似度を算出する。
 類似度の算出としては、文をベクトル化してコサイン類似度を用いればよい。文のベクトル化には、senteceBERTなどを用いればよい。コサイン類似度は−1から1の間の実数で表され、値が1に近いほど、比較される2つの文書が似ていること(類似度が高い)を示し、コサイン類似度の値が−1に近いほど似ていないこと(類似度が低い)を示す。
 または、レーベンシュタイン距離、またはジャロ・ウィンクラー距離などの編集距離を類似度の代わりに用いてもよい。類似度の代わりに、編集距離を用いる場合、値が0に近いほど比較される2つの文書が似ていること(類似度が高い)を示し、編集距離の値が大きいほど似ていない(類似度が低い)ことを示す。
 または、BLEU(Bilingual Evaluation Understudy)スコアを類似度の代わりに用いてもよい。BLEUスコアは0から1の間の実数で表され、値が1に近いほど比較される2つの文書が似ていること(類似度が高い)を示し、BLEUスコアの値が0に近いほど似ていない(類似度が低い)ことを示す。
 次に、図7AのステップS304において、第2の文書(翻訳文)と第3の文書(直訳文)の類似度と、第2の文書(翻訳文)と第4の文書(意訳文)の類似度と、を比較する。比較の結果、第4の文書(意訳文)の類似度の方が、第3の文書(直訳文)の類似度よりも高い場合に、第2の文書(翻訳文)に意訳が含まれると判定する。または、比較の結果、第3の文書(直訳文)の類似度の方が、第4の文書(意訳文)の類似度よりも高い場合に、第2の文書(翻訳文)に意訳が含まれていないと判定する。このステップS304の判定の後に、図2のステップS122に進むことができる。
 以上が、実施の形態1に記載した情報処理システムの構成例1と、情報処理方法が異なる部分の説明である。次に、図7Bを用いて、本実施の形態における、情報処理装置10、情報処理装置40、及び情報端末20で行われる処理の関係について説明する。
 図7Bは、情報処理装置10、情報処理装置40、及び情報端末20で行われる処理を、図2、図4及び図7Aで用いた各ステップの符号を用いて示した関係図であり、図5とは異なる構成例である。また、ネットワーク30を付したブロック矢印は、情報処理装置10と情報処理装置40の間のデータの送受信を示しており、また、ネットワーク31を付したブロック矢印は、情報処理装置10と情報端末20の間のデータの送受信を示している。
 図7Bに示すように、情報処理装置10は、ステップS301の処理を行う機能と、ステップS302の処理を行う機能と、ステップS303の処理を行う機能と、ステップS304の処理を行う機能と、ステップS122の処理を行う機能と、ステップS131の処理を行う機能と、ステップS141の処理を行う機能と、を有する。
 また、図7Bに示すように、情報処理装置40は、ステップS132の処理を行う機能と、ステップS142の処理を行う機能と、を有する。
 また、図7Bに示すように、情報端末20は、ステップS101の処理を行う機能と、ステップS111の処理を行う機能と、ステップS123の処理を行う機能と、ステップS133の処理を行う機能と、ステップS134の処理を行う機能と、ステップS143の処理を行う機能と、ステップS144の処理を行う機能と、を有する。
 実施の形態1の情報処理システムの構成例1では、「ステップS121:翻訳文に意訳が含まれるか判定」として、情報処理装置10と情報処理装置40を用いて処理することができる。これに対して、実施の形態2の情報処理システムの構成例2では、「ステップS121:翻訳文に意訳が含まれるか判定」として、情報処理装置10のみで処理を行うことができる。
 以上に示した情報処理システム、情報処理方法、及び情報処理装置により、翻訳文のチェック及び修正する作業を効率的に行うことが可能となる。
10:情報処理装置、20a:デスクトップ型コンピュータ、20b:ノート型コンピュータ、20c:スマートフォン、20d:タブレット型コンピュータ、20:情報端末、21:筐体、30:ネットワーク、31:ネットワーク、40:情報処理装置、110:入力部、120:記憶部、130:処理部、140:出力部、150:伝送路

Claims (10)

  1.  第1の情報処理装置及び第2の情報処理装置を有し、
     前記第1の情報処理装置は、
     第1の言語の第1の文書、及び、前記第1の文書の翻訳文である第2の言語の第2の文書の入力を受け付ける機能と、
     前記第2の文書に前記第1の文書の意訳が含まれるか、を判定させるための指示文と、前記第1の文書と、前記第2の文書と、を含むプロンプトを前記第2の情報処理装置に送信することで、前記判定の結果を取得する機能と、
     前記判定で前記意訳が含まれていないと判定された前記第2の文書に含まれる語句と前記第1の文書に含まれる語句との、対応関係を出力させるための指示文と、前記第1の文書と、前記第2の文書と、を含むプロンプトを前記第2の情報処理装置に送信することで、前記対応関係を取得する機能と、
     前記対応関係を出力する機能と、を有する、
     情報処理システム。
  2.  請求項1において、
     前記第1の情報処理装置は、
     前記第2の文書に意訳を含むことを許容するか、の選択を受け付ける機能と、
     前記選択で意訳を含むことを許容すると選択された前記第2の文書と、前記第1の文書と、の同義判定を行なわせるための指示文と、前記第1の文書と、前記第2の文書と、を含むプロンプトを前記第2の情報処理装置に送信することで、前記同義判定の結果を取得する機能を有する、
     情報処理システム。
  3.  請求項1において、
     前記第1の情報処理装置は、
     前記意訳が含まれていると判定された前記第2の文書の修正を受け付け、前記意訳が含まれていないと判定されるまで、前記修正の受付と、前記判定のためのプロンプトの送信と、前記判定の結果の取得と、を繰り返す機能を有する、
     情報処理システム。
  4.  請求項1において、
     前記第1の情報処理装置は、
     前記第2の文書に含まれる語句と前記第1の文書に含まれる語句の前記対応関係に誤りがある場合に、前記第2の文書の修正を受け付ける機能を有する、
     情報処理システム。
  5.  請求項2において、
     前記第1の情報処理装置は、
     前記第2の文書と前記第1の文書の前記同義判定に誤りがある場合に、前記第2の文書の修正を受け付ける機能を有する、
     情報処理システム。
  6.  第1の情報処理装置及び第2の情報処理装置を有し、
     前記第1の情報処理装置は、
     第1の言語の第1の文書、及び、前記第1の文書の翻訳文である第2の言語の第2の文書の入力を受け付ける機能と、
     前記第1の文書の前記第2の言語の直訳文を生成する機能と、
     前記第1の文書の前記第2の言語の意訳文を生成する機能と、
     前記第2の文書と前記直訳文の類似度を算出する機能と、
     前記第2の文書と前記意訳文の類似度を算出する機能と、
     前記直訳文の類似度と、前記意訳文の類似度を比較し、前記直訳文の類似度の方が高い場合に、前記第2の文書に前記第1の文書の意訳が含まれていないと判定する機能と、
     前記判定で前記意訳が含まれていないと判定された前記第2の文書に含まれる語句と前記第1の文書に含まれる語句の、対応関係を出力させるための指示文と、前記第1の文書と、前記第2の文書と、を含むプロンプトを前記第2の情報処理装置に送信することで、前記対応関係を取得する機能と、を有する、
     情報処理システム。
  7.  第1の言語の第1の文書、及び、前記第1の文書の翻訳文である第2の言語の第2の文書の入力を受け付けるステップと、
     前記第2の文書に前記第1の文書の意訳が含まれるか、を判定させるための指示文と、前記第1の文書と、前記第2の文書と、を含む第1のプロンプトを作成するステップと、
     前記第1のプロンプトに従い、前記判定を行なうステップと、
     前記判定で前記意訳が含まれていないと判定される場合に、前記第2の文書に含まれる語句と前記第1の文書に含まれる語句の、対応関係を出力させるための指示文と、前記第1の文書と、前記第2の文書と、を含む第2のプロンプトを作成するステップと、
     前記第2のプロンプトに従い、前記対応関係を出力するステップと、を有する、
     情報処理方法。
  8.  請求項7において、
     前記第2の文書に意訳を含むことを許容するか、の選択を受け付けるステップと、
     前記選択で意訳を含むことを許容すると選択された前記第2の文書と、前記第1の文書と、の同義判定を行なわせるための指示文と、前記第1の文書と、前記第2の文書と、を含む第3のプロンプトを作成するステップと、
     前記第3のプロンプトに従い、前記同義判定を行なうステップと、を有する、
     情報処理方法。
  9.  請求項7において、
     前記判定は、大規模言語モデルによって行われ、
     前記対応関係の出力は、前記大規模言語モデルによって行われる、
     情報処理方法。
  10.  請求項8において、
     前記判定は、大規模言語モデルによって行われ、
     前記対応関係の出力は、前記大規模言語モデルによって行われ、
     前記同義判定は、前記大規模言語モデルによって行われる、
     情報処理方法。
PCT/IB2024/061478 2023-11-24 2024-11-18 情報処理システム、情報処理方法 Pending WO2025109448A1 (ja)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2023198950 2023-11-24
JP2023-198950 2023-11-24

Publications (1)

Publication Number Publication Date
WO2025109448A1 true WO2025109448A1 (ja) 2025-05-30

Family

ID=95826113

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/IB2024/061478 Pending WO2025109448A1 (ja) 2023-11-24 2024-11-18 情報処理システム、情報処理方法

Country Status (1)

Country Link
WO (1) WO2025109448A1 (ja)

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2009064137A (ja) * 2007-09-05 2009-03-26 Nippon Hoso Kyokai <Nhk> 対訳表現アラインメント装置およびそのプログラム

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2009064137A (ja) * 2007-09-05 2009-03-26 Nippon Hoso Kyokai <Nhk> 対訳表現アラインメント装置およびそのプログラム

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
IMAMURA, KENJI : "Automatic Construction of Machine Translation Knowledge Using Translation Literalness", JOURNAL OF NATURAL LANGUAGE PROCESSING, vol. 11, 1 April 2004 (2004-04-01), pages 85 - 99, XP093317513 *
IMAMURA, KENJI: "Bilingual Corpus Filtering Based on Translation Literality", FORUM ON INFORMATION TECHNOLOGY 2002, GENERAL LECTURE PAPERS, vol. 2, 1 January 2002 (2002-01-01), pages 185 - 186, XP093317509 *

Similar Documents

Publication Publication Date Title
CN109710951B (zh) 基于翻译历史的辅助翻译方法、装置、设备及存储介质
US20240231763A9 (en) Interactive editing of a machine-generated document
US11507760B2 (en) Machine translation method, machine translation system, program, and non-transitory computer-readable storage medium
CN113836192A (zh) 平行语料的挖掘方法、装置、计算机设备及存储介质
CN120338086A (zh) 一种基于对比监督和跨阶段蒸馏的通用信息抽取方法
US12578943B2 (en) Translation quality assurance based on large language models
US20250013829A1 (en) Bottom-up neural semantic parser
WO2025088438A1 (ja) 情報処理システム、情報処理方法
US20250094738A1 (en) Information processing system and information processing method
WO2025253263A1 (ja) 情報処理システム
US20260087273A1 (en) Data processing system and data processing method
US20260017446A1 (en) System for supporting specification preparation
CN114817469A (zh) 文本增强方法、文本增强模型的训练方法及装置
US20250156466A1 (en) Document reading support method
US20250165235A1 (en) Data processing system and data processing method
US20250200089A1 (en) Information processing system and information processing method
US20260087573A1 (en) Information processing system and information processing method
JP7859776B2 (ja) 機械翻訳システム
WO2025202851A1 (ja) 情報処理システム、及び情報処理方法
US20250258656A1 (en) Data processing system and data processing method
US20250110763A1 (en) Data processing system and data processing method
WO2025032472A1 (ja) 情報処理システム、情報処理方法
JP2025132683A (ja) 文書生成システム及びその動作方法
US20250231956A1 (en) Information processing system and information processing method
WO2026099696A1 (ja) 情報処理装置、及び情報処理方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24893689

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2025558914

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 2025558914

Country of ref document: JP