WO2022249256A1 - Api検出装置、api検出方法及びプログラム - Google Patents

Api検出装置、api検出方法及びプログラム Download PDF

Info

Publication number
WO2022249256A1
WO2022249256A1 PCT/JP2021/019661 JP2021019661W WO2022249256A1 WO 2022249256 A1 WO2022249256 A1 WO 2022249256A1 JP 2021019661 W JP2021019661 W JP 2021019661W WO 2022249256 A1 WO2022249256 A1 WO 2022249256A1
Authority
WO
WIPO (PCT)
Prior art keywords
api
source code
token
model
notation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2021/019661
Other languages
English (en)
French (fr)
Inventor
利行 倉林
卓弥 岩塚
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to JP2023523737A priority Critical patent/JPWO2022249256A1/ja
Priority to PCT/JP2021/019661 priority patent/WO2022249256A1/ja
Publication of WO2022249256A1 publication Critical patent/WO2022249256A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F8/00—Arrangements for software engineering
    • G06F8/70—Software maintenance or management
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00—Machine learning

Definitions

  • the present invention relates to an API detection device, an API detection method, and a program.
  • Non-Patent Document 1 there is a technique for testing by narrowing down the test target to only the changed parts.
  • static analysis is used to detect and test changed portions and APIs for executing the changed portions.
  • API notation differs depending on the web framework (hereinafter referred to as "FW") used by the software, so it is necessary to set individual analysis rules for detecting APIs for each FW.
  • FW web framework
  • the demand for the Web is very high, and there are various types of FWs. Therefore, setting individual analysis rules so as to correspond to all FWs is very costly.
  • Web technology is a rapidly changing field, and there is a high possibility that a new FW or a new notation will appear even in the same FW. Therefore, with the technique of Non-Patent Document 1, there is a high possibility that a huge amount of cost will be required to detect APIs in response to various FWs.
  • the present invention has been made in view of the above points, and aims to improve the efficiency of API detection.
  • the API detection device is based on learning data in which predetermined labels are assigned to tokens that are components of the API in the token string related to the first source code.
  • a learning unit that learns a model for detecting the component from a token string; and an API detection unit that causes the model to detect the
  • API detection can be made more efficient.
  • FIG. 4 is a flowchart for explaining an example of a processing procedure executed by the API detection device 10 during learning;
  • FIG. 2 is a diagram showing an example of the functional configuration of the API detecting device 10 when detecting an API according to the embodiment of the present invention;
  • 4 is a flowchart for explaining an example of a processing procedure executed by the API detection device 10 when detecting an API;
  • API Web API
  • FW Web framework
  • annotation notation the method of describing with annotations
  • function notation the method of describing with functions
  • FIG. 1 is a diagram showing a hardware configuration example of an API detection device 10 according to an embodiment of the present invention.
  • the API detection device 10 of FIG. 1 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a CPU 104, an interface device 105, a display device 106, an input device 107, and the like, which are connected to each other via a bus B, respectively.
  • a program that implements processing in the API detection device 10 is provided by a recording medium 101 such as a CD-ROM.
  • a recording medium 101 such as a CD-ROM.
  • the program is installed from the recording medium 101 to the auxiliary storage device 102 via the drive device 100 .
  • the program does not necessarily need to be installed from the recording medium 101, and may be downloaded from another computer via the network.
  • the auxiliary storage device 102 stores installed programs, as well as necessary files and data.
  • the memory device 103 reads and stores the program from the auxiliary storage device 102 when a program activation instruction is received.
  • the CPU 104 implements functions related to the API detection device 10 according to programs stored in the memory device 103 .
  • the interface device 105 is used as an interface for connecting to a network.
  • a display device 106 displays a program-based GUI (Graphical User Interface) or the like.
  • the input device 107 is composed of a keyboard, a mouse, etc., and is used to input various operational instructions.
  • FIG. 2 is a diagram showing a functional configuration example during learning of the API detection device 10 according to the embodiment of the present invention.
  • the API detection device 10 has a preprocessing section 11 and a learning section 12 . Each of these units is implemented by processing that one or more programs installed in the API detection device 10 cause the CPU 104 to execute.
  • a learning source code is prepared and input to the API detection device 10 .
  • the data structure of the learning source code is described below in a format based on the BNF (Backus-Naur form) notation.
  • the learning source code is a set of source code files in the project data of an actual program that utilizes (calls) arbitrary multiple FW APIs determined as learning targets.
  • FIG. 3 is a flowchart for explaining an example of a processing procedure executed by the API detection device 10 during learning.
  • the processing procedure of FIG. 3 is executed for each of the learning source code related to the annotation notation and the learning source code related to the function notation.
  • two models, an API detection model related to annotation notation and an API detection model related to function notation, are learned (generated). Note that FIG. 3 does not distinguish between annotation notation and function notation to show common processing.
  • the processing procedure in FIG. 3 includes loop processing L1 for each source code file included in the learning source code.
  • the source code stored in the source code file to be processed in the loop processing L1 is referred to as "target code”.
  • Loop processing L1 includes loop processing L2.
  • the loop processing L2 is loop processing for each block included in the target code.
  • a block to be processed in loop processing L2 is hereinafter referred to as a "target block”.
  • the preprocessing unit 11 abstracts each token in the token string (S102). This is to improve the accuracy of learning by executing abstraction. Any method may be used for abstraction. In the example below, lowercase and string masking are applied. Mask processing is, for example, a process of replacing character strings that are not reserved in the grammar of a program (for example, character strings enclosed by "", comments, etc.) with a predetermined character string ( ⁇ string> in the example below). Say. The following shows the token sequence before abstraction as input and the token sequence after abstraction as output.
  • API information is a set of a token and a label given to the token.
  • the labeling may be performed automatically using a labeling program (script) or the like, or may be performed manually.
  • the preprocessing unit 11 may receive an input of a label for each token from the user.
  • the learning unit 12 performs the API information sequence generated in the loop processing L1.
  • the API detection model is made to learn the correspondence between tokens and labels in API information (S104).
  • Any model neural network
  • LSTM, Transformer (Encoder part), etc. are suitable as the API detection model.
  • Word Embedding or the like may be used to embed the token.
  • FIG. 4 is a diagram showing an example of the functional configuration of the API detecting device 10 when detecting an API according to the embodiment of the present invention.
  • the API detection device 10 has a change file extraction unit 13 , a preprocessing unit 14 and an API detection unit 15 . Each of these units is implemented by processing that one or more programs installed in the API detection device 10 cause the CPU 104 to execute.
  • FIG. 5 is a flowchart for explaining an example of a processing procedure executed by the API detection device 10 when detecting an API.
  • step S201 the change file extraction unit 13 extracts source code files changed by the developer (hereinafter referred to as "change files") from the source code files of the application.
  • a change file may be a file that contains changes to an old version for a new version, or a file that contains changes made to a point in time during the course of development, for example.
  • the newly created source code file is itself a modified file.
  • loop processing L3 is executed for each change file.
  • target code the source code included in the change file to be processed in loop processing L3
  • Loop processing L3 includes loop processing L4.
  • the loop processing L4 is loop processing for each block included in the target code.
  • a block to be processed in loop processing L4 is hereinafter referred to as a "target block”.
  • step S202 the preprocessing unit 14 generates a token string of the target block by tokenizing the target block.
  • the content of this process is the same as that of step S101 in FIG.
  • the preprocessing unit 14 abstracts each token in the token string (S203).
  • the content of this process is the same as that of step S102 in FIG.
  • the preprocessing unit 14 stores the abstracted N-th token in the auxiliary storage device 102 or the memory device 103 in association with the token before abstraction and N. That is, the correspondence relationship before and after abstraction is maintained.
  • the API detection unit 15 inputs the abstracted token string obtained in step S202 to the learned API detection model, and outputs each token constituting the token string from the API detection model. obtain the label (S203). At this time, the API detection unit 15 inputs the token string (the same token string) to an API detection model that has been trained for annotation notation and an API detection model that has been trained for function notation. Get the token label. For each token, the API detection unit 15 determines the OR (logical sum) of two labels output from each API detection model for the token as the label of the token. The OR of two labels means that when at least one of the two labels is Path or HttpMethod and the other is None, Path or HttpMethod is adopted.
  • the API detection unit 15 generates API information including a combination of a token labeled Path and a token labeled HttpMethod in the token string restored to the pre-abstraction state, and generates API information is added to the test API (S206).
  • the data structure of the test API is described below in a format based on the BNF notation.
  • the API information is a set of a token whose Path label is estimated and a token whose HttpMethod label is estimated. For example, in the above example, the API information is ""/test","GET"".
  • the API detection unit 15 When loop processing L4 is completed for all blocks of the target code and loop processing L3 is completed for all change files, the API detection unit 15 outputs the contents of the test API (API information list) (S207).
  • the user can refer to the list of output API information to specify the API to be tested. For example, in the above example, the user can recognize that a test such as "send a GET request to "/test"" should be executed.
  • a test such as "send a GET request to "/test"" should be executed.
  • the technology disclosed in Non-Patent Document 1 may be used for such a test.
  • APIs can be detected without setting analysis rules for various types of FW. That is, API detection can be made more efficient. As a result, it is possible to test various types of FW at low cost, and shift left can be realized.
  • annotation notation and function notation are used as examples of API notation, but for APIs with other notations, learning and detection may be performed with respect to three or more notations.
  • API detection device 11 preprocessing unit 12 learning unit 13 change file extraction unit 14 preprocessing unit 15 API detection unit 100 drive device 101 recording medium 102 auxiliary storage device 103 memory device 104 CPU 105 interface device 106 display device 107 input device B bus

Landscapes

  • Engineering & Computer Science (AREA)
  • Software Systems (AREA)
  • Theoretical Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Artificial Intelligence (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Medical Informatics (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • Stored Programmes (AREA)

Abstract

API検出装置は、第1のソースコードに係るトークン列のうちAPIの構成要素であるトークンに対して所定のラベルが付与された学習データに基づいて、入力されたトークン列の中から前記構成要素を検出するモデルを学習する学習部と、第2のソースコードに係るトークン列を前記モデルへ入力することで、前記第2のソースコードに含まれるAPIの構成要素を前記モデルに検出させるAPI検出部と、を有することで、APIの検出を効率化する。

Description

API検出装置、API検出方法及びプログラム
 本発明は、API検出装置、API検出方法及びプログラムに関する。
 多様化する消費者ニーズに合ったサービスを迅速に提供するために、近年、ソフトウェア開発を短いサイクルで実施し、頻繁にサービスをリリースする開発スタイルが増加している.また、このような短いサイクルでのサービス提供を維持しながら品質を確保することも求められている.ソフトウェア開発において、テストは開発の終盤で行われることが多く、開発終盤で脆弱性が発見されると大きな手戻りが発生するという問題がある。そこで実装段階でテストを前倒しで行う、シフトレフトと呼ばれる考え方が存在する。
 従来、テスト対象を変更箇所のみに絞ってテストする技術が存在する(非特許文献1)。この技術では、静的解析により、変更箇所とその変更箇所を実行するためのAPIを検出してテストを行う。
岩塚 卓弥,倉林 利行,切貫 弘之、"テストプロセスの変革~セキュリティテストと探索的テスト~"、ビジネスコミュニケーション 2020 Vol.57 No.2、[online]、インターネット<URL:https://www.sic.ecl.ntt.co.jp/mt_assets/bc_202002/bc03.pdf>
 APIは、ソフトウェアが使用するWebフレームワーク(以下「FW」という。)によって記法が異なるため、FWごとにAPIを検出するための個別の解析ルールを設定する必要がある。しかしながら、現代社会において、Webは非常に需要が高く、FWも様々な種類が存在するため、あらゆるFWに対応するように個別に解析ルールを設定するためには大きなコストがかかる。またWeb技術は変化の激しい分野であり、新しいFWや、同一FWでも新しい記法が登場する可能性が高い。したがって、非特許文献1の技術では、様々なFWに対応させてAPIを検出するために膨大なコストを要する可能性が高い。
 本発明は、上記の点に鑑みてなされたものであって、APIの検出を効率化することを目的とする。
 そこで上記課題を解決するため、API検出装置は、第1のソースコードに係るトークン列のうちAPIの構成要素であるトークンに対して所定のラベルが付与された学習データに基づいて、入力されたトークン列の中から前記構成要素を検出するモデルを学習する学習部と、第2のソースコードに係るトークン列を前記モデルへ入力することで、前記第2のソースコードに含まれるAPIの構成要素を前記モデルに検出させるAPI検出部と、を有する。
 APIの検出を効率化することができる。
本発明の実施の形態におけるAPI検出装置10のハードウェア構成例を示す図である。 本発明の実施の形態におけるAPI検出装置10の学習時の機能構成例を示す図である。 API検出装置10が学習時に実行する処理手順の一例を説明するためのフローチャートである。 本発明の実施の形態におけるAPI検出装置10のAPIの検出時の機能構成例を示す図である。 API検出装置10がAPI検出時に実行する処理手順の一例を説明するためのフローチャートである。
 本実施の形態は,非特許文献1においてWebフレームワーク(以下「FW」という。)におけるWebAPI(以下、単に「API」という。)の記法を学習することで、少ないコストで様々なFWにおけるAPIの検出を可能とする。FWのAPI記法は大きく2つの種別にわけることができ、1つ目はアノテーションで記述する方法(以下、「アノテーション記法」という。)、2つ目は関数で記述する方法(以下、「関数記法」という。)である。アノテーション記法に分類されるFWのAPIの記法には類似する点が存在する。例えば、アノテーション記法のFWの1つであるMicronautは、以下に示すようにAPIを記述する。
@app.route("/test", methods=["GET"])
def test():
 このコントローラは、/test宛てにGETリクエストを受けたときに呼び出される。すなわち、パス及びHTTPメソッドがAPIの構成要素である。したがって、パスとして、"/test"が検出でき、HTTPメソッドとして"GET"が検出できればAPIを検出できることになる。
 同じくアノテーション記法のFWの1つであるSymfonyでは、以下のようにAPIを記述する。
/**
* @Route("/test", methods={"GET"})
*/
public function test(){
 それぞれ細部は異なるものの、「route」というキーワードの後にパス、「methods」というキーワードの後にHTTPメソッドが記述されている。その他、パスに関するキーワードとして「Controller」や「Path」、HTTPメソッドに関するキーワードとして「GET」や「POST」などが存在する。
 このように、相互に異なる複数のAPI記法には類似性が存在するため、アノテーション記法に分類される複数のFWでAPI記法を学習すれば、学習外のアノテーション記法に分類されるFWでもAPIが検出できると考えられる。関数記法も同様である。本実施の形態では、この類似性に着目し、アノテーション記法と関数記法のそれぞれのAPI記述を学習することで、幅広いFWでのAPI検出を可能とする。
 以下、図面に基づいて本発明の実施の形態を説明する。図1は、本発明の実施の形態におけるAPI検出装置10のハードウェア構成例を示す図である。図1のAPI検出装置10は、それぞれバスBで相互に接続されているドライブ装置100、補助記憶装置102、メモリ装置103、CPU104、インタフェース装置105、表示装置106、及び入力装置107等を有する。
 API検出装置10での処理を実現するプログラムは、CD-ROM等の記録媒体101によって提供される。プログラムを記憶した記録媒体101がドライブ装置100にセットされると、プログラムが記録媒体101からドライブ装置100を介して補助記憶装置102にインストールされる。但し、プログラムのインストールは必ずしも記録媒体101より行う必要はなく、ネットワークを介して他のコンピュータよりダウンロードするようにしてもよい。補助記憶装置102は、インストールされたプログラムを格納すると共に、必要なファイルやデータ等を格納する。
 メモリ装置103は、プログラムの起動指示があった場合に、補助記憶装置102からプログラムを読み出して格納する。CPU104は、メモリ装置103に格納されたプログラムに従ってAPI検出装置10に係る機能を実現する。インタフェース装置105は、ネットワークに接続するためのインタフェースとして用いられる。表示装置106はプログラムによるGUI(Graphical User Interface)等を表示する。入力装置107はキーボード及びマウス等で構成され、様々な操作指示を入力させるために用いられる。
 図2は、本発明の実施の形態におけるAPI検出装置10の学習時の機能構成例を示す図である。図2において、API検出装置10は、前処理部11及び学習部12を有する。これら各部は、API検出装置10にインストールされた1以上のプログラムが、CPU104に実行させる処理により実現される。
 学習時には、学習用ソースコードが用意され、API検出装置10に入力される。学習用ソースコードのデータ構造をBNF(Backus-Naur form)記法に基づく形式によって記すと以下の通りである。
学習用ソースコード::=ソースコードファイル+
 すなわち、学習用ソースコードは、学習対象として決定された任意の複数のFWのAPIを利用する(呼び出す)実際の或るプログラムのプロジェクトデータにおけるソースコードファイルの集合である。ここでは、相互にAPIの記法が異なる複数のプロジェクトデータが用意されることとする。したがって、学習用ソースコードは、アノテーション記法と関数記法とのそれぞれについて用意される。
 図3は、API検出装置10が学習時に実行する処理手順の一例を説明するためのフローチャートである。図3の処理手順は、アノテーション記法に係る学習用ソースコードと、関数記法に係る学習用ソースコードとのそれぞれについて実行される。その結果、アノテーション記法に係るAPI検出モデルと、関数記法に係るAPI検出モデルとの2つのモデルが学習(生成)される。なお、図3では、アノテーション記法及び関数記法について共通の処理を示すため、記法の区別は行わない。
 図3の処理手順は、学習用ソースコードに含まれるソースコードファイルごとのループ処理L1を含む。以下、ループ処理L1において処理対象とされているソースコードファイルが格納するソースコードを「対象コード」という。ループ処理L1は、ループ処理L2を含む。ループ処理L2は、対象コードが含むブロックごとのループ処理である。以下、ループ処理L2において処理対象とされているブロックを「対象ブロック」という。
 なお、ブロックとは、インデント長(字下げ幅)が同じである連続する行の集合をいう。例えば、
Def hoge():
  A=1
  B=2
というソースコードでは、「Def hoge():」が1つのブロックとなり、「A=1 B=2」が2つ目のブロックとなる。
 ステップS101において、前処理部11は、対象ブロックをトークナイズすることで、対象ブロックのトークン列を生成する。例えば、対象ブロックが以下の「入力」の通りである場合、以下の「出力」がトークナイズの結果であるトークン列として得られる。
入力:@Route("/test", methods={"GET"})
出力:@,Route,(, "/test",,, methods,=,{, "GET",},)]
 すなわち、トークン列は、対象ブロックに含まれる各トークンが、対象ブロックにおける出現順に配列されたデータであり、各トークンはカンマで区切られている。
 続いて、前処理部11は、当該トークン列の各トークンについて抽象化を実行する(S102)。抽象化を実行することで学習の精度を上げるためである。抽象化には任意の方法が用いられればよい。以下の例では、小文字化と文字列のマスク処理とが施されている。マスク処理とは、例えば、プログラムの文法において予約されていない文字列(例えば、""によって囲まれた文字列やコメント等)を所定の文字列(以下の例では<string>)に置換する処理をいう。以下は、入力として抽象化前のトークン列を示し、出力として抽象化後のトークン列を示す。
入力:[@,Route,(, "/test",,, methods,=,{, "GET",},)]
出力:[@,route,(, <string>,,, methods,=,{, <string>,},)]
 続いて、前処理部11は、抽象化されたトークン列の各トークンに対するラベル付けを行うことで、トークンごとに、API情報を生成する(S103)。以下にラベル付けの一例を示す。
入力:[@,route,(, <string>,,, methods,=,{, <string>,},)]
出力:[None,None,None,Path,None,None,None,None,HttpMethod,None,None]
 すなわち、ラベル付けは、各トークンに対して、None、Path又はHttpMethodのラベルを付与することである。APIの構成要素であるパス及びHTTPメソッドのいずれでもないトークンに対しては、Noneが付与される。APIの構成要素であるトークンに対しては、所定のトークンが付与される。具体的には、パスであるトークンに対しては、Pathが付与される。HTTPメソッドであるトークンに対しては、HttpMethodが付与される。なお、API情報のデータ構造をBNF記法に基づく形式によって記すと以下の通りである。
API情報::=(トークン,ラベル)+
ラベル::=(None|Path|HttpMethod)
 すなわち、API情報とは、トークンと当該トークンに付与されたラベルとの組である。なお、ラベル付けは、例えば、ラベル付けのためのプログラム(スクリプト)等を用いて自動的に行われてもよいし、人手で行われてもよい。人手で行われる場合、例えば、前処理部11は、各トークンに対するラベルの入力をユーザから受け付けてもよい。
 ループ処理L2が対象コードの全てのブロックについて終了し、ループ処理L1が学習用ソースコードの全てのソースコードファイルについて実行されると、学習部12は、ループ処理L1において生成されたAPI情報の系列(API情報がトークンの並び順で整列されたデータ)を学習データとして、API情報におけるトークンとラベルとの対応関係をAPI検出モデルに学習させる(S104)。API検出モデルは、任意のモデル(ニューラルネットワーク)を用いてよいが、ソースコードは時系列データであるため、LSTMやTransformer(Encoder部分)などがAPI検出モデルとして好適である。トークンの埋め込みにはWord Embeddingなどが用いられればよい。このようなモデルを利用することで、各トークンについて、当該トークンとラベルとの対応関係だけでなく、前後の文脈にも着目して各トークンに対応するラベルを推定可能とすることができる。
 図3の処理手順が、アノテーション記法と関数記法のそれぞれについて終了すると、それぞれの記法に対応した学習済みのAPI検出モデルが生成される。
 図4は、本発明の実施の形態におけるAPI検出装置10のAPIの検出時の機能構成例を示す図である。図4において、API検出装置10は、変更ファイル抽出部13、前処理部14及びAPI検出部15を有する。これら各部は、API検出装置10にインストールされた1以上のプログラムが、CPU104に実行させる処理により実現される。
 図5は、API検出装置10がAPI検出時に実行する処理手順の一例を説明するためのフローチャートである。
 ステップS201において、変更ファイル抽出部13は、アプリケーションのソースコードファイルの中から、開発者によって変更されたソースコードファイル(以下、「変更ファイル」という。)を抽出する。変更ファイルは、新バージョンについて旧バージョンに対する変更を含むファイルであってもよいし、例えば、開発の進行過程において、或る時点に対して行われた変更を含むファイルであってもよい。又は、新規に作成されたソースコードのファイルは、それ自体が変更ファイルである。
 続いて、変更ファイルごとに、ループ処理L3が実行される。以下、ループ処理L3において処理対象とされている変更ファイルが含むソースコードを「対象コード」という。ループ処理L3は、ループ処理L4を含む。ループ処理L4は、対象コードが含むブロックごとのループ処理である。以下、ループ処理L4において処理対象とされているブロックを「対象ブロック」という。
 ステップS202において、前処理部14は、対象ブロックをトークナイズすることで、対象ブロックのトークン列を生成する。この処理内容は、図3のステップS101と同様である。
 続いて、前処理部14は、当該トークン列の各トークンについて抽象化を実行する(S203)。この処理内容は、図3のステップS102と同様である。但し、ステップS203において、前処理部14は、抽象化されたN番目のトークンについて、抽象化前のトークンとNとを対応付けて補助記憶装置102又はメモリ装置103に保存しておく。すなわち、抽象化前後の対応関係が保持される。
 ステップS202及びS203の実行により、対象ブロックが以下の入力の通りであれば、ステップS203では、以下の出力が得られる。
入力:@app.route("/test", methods=["GET"])
出力:[@,app,route,(,<string>,,,methods,=,[, <string>,],)]
 この例では、"/test"が<string>に置換され、"GET"が<string>に置換されている。したがって、トークン列における"/test"の順番と"/test"との対応関係と、トークン列における"GET"の順番と"GET"との対応関係とが補助記憶装置102又はメモリ装置103に保存される。
 続いて、API検出部15は、ステップS202において得られた抽象化されたトークン列を学習済みのAPI検出モデルに入力することで、当該トークン列を構成する各トークンについて当該API検出モデルから出力されるラベルを取得する(S203)。この際、API検出部15は、当該トークン列(同一のトークン列)を、アノテーション記法について学習済みのAPI検出モデルと、関数記法について学習済みのAPI検出モデルとのそれぞれに入力し、それぞれから各トークンのラベルを取得する。API検出部15は、トークンごとに、当該トークンについて各API検出モデルから出力された2つのラベルのOR(論理和)を当該トークンのラベルとして判定する。2つのラベルのORとは、2つのラベルのうち少なくともいずれか一方がPath又はHttpMethodであり、他方がNoneである場合に、Path又はHttpMethodが採用されることをいう。2つのAPI検出モデルが使用されるのは、FWによってはアノテーション記法と関数記法の両方が存在するためである。以下に、ステップS203の入力と出力との例を示す。
入力:[@,app,route,(,<string>,,,methods,=,[, <string>,],)]
出力:[None, None, None, None, Path, None, None, None, None, HttpMethod, None, None]
 これによって、トークン列の各トークンに対するラベルが得られる。換言すれば、いずれのトークンがAPIの構成要素であるかが分かる。
 続いて、API検出部15は、当該トークン列において抽象化されたトークンを抽象化前の文字列に復元する(S205)。例えば、上記の入力が、以下のように復元される。
[@,app,route,(, "/test",,,methods,=,[, "GET",],)]
 続いて、API検出部15は、抽象化前の状態に復元されたトークン列のうち、ラベルがPathであるトークンとラベルがHttpMethodであるトークンとの組を含むAPI情報を生成し、当該API情報をテスト用APIへ追加する(S206)。テスト用APIのデータ構造をBNF記法に基づく形式によって記すと以下の通りである。
テスト用API::=API情報+
API情報::=<Path>,<HttpMethod>
 すなわち、テスト用APIは、API情報のリストである。API情報は、Pathのラベルが推定されたトークンと、HttpMethodのラベルが推定されたトークンとの組である。例えば、上記の例において、API情報は、「"/test","GET"」である。
 ループ処理L4が対象コードの全てのブロックについて終了し、ループ処理L3が全ての変更ファイルについて終了すると、API検出部15は、テスト用APIの内容(API情報のリスト)を出力する(S207)。
 ユーザは、出力されたAPI情報のリストを参照して、テストすべきAPIを特定することができる。例えば、上記の例では、ユーザは、「"/test"宛てにGETリクエストを送る」といったテストを実行すべきであることを認識することができる。このようなテストには、非特許文献1に開示された技術が利用されてもよい。
 上述したように、本実施の形態によれば、API記法を学習することで様々な種類のFWに対して解析ルールを設定することなくAPIを検出することができる。すなわち、APIの検出を効率化することができる。その結果、様々な種類のFWに対して少ないコストでテストを実施することができ、シフトレフトを実現できる。
 なお、上記では、APIの記法として、アノテーション記法と関数記法を例として説明したが、他の記法が存在するAPIについては、3つ以上の記法に関して学習及び検出が行われてもよい。
 

 以上、本発明の実施の形態について詳述したが、本発明は斯かる特定の実施形態に限定されるものではなく、請求の範囲に記載された本発明の要旨の範囲内において、種々の変形・変更が可能である。
10     API検出装置
11     前処理部
12     学習部
13     変更ファイル抽出部
14     前処理部
15     API検出部
100    ドライブ装置
101    記録媒体
102    補助記憶装置
103    メモリ装置
104    CPU
105    インタフェース装置
106    表示装置
107    入力装置
B      バス

Claims (7)

  1.  第1のソースコードに係るトークン列のうちAPIの構成要素であるトークンに対して所定のラベルが付与された学習データに基づいて、入力されたトークン列の中から前記構成要素を検出するモデルを学習する学習部と、
     第2のソースコードに係るトークン列を前記モデルへ入力することで、前記第2のソースコードに含まれるAPIの構成要素を前記モデルに検出させるAPI検出部と、
    を有することを特徴とするAPI検出装置。
  2.  前記学習部は、APIに関する複数の記法の種別ごとに、当該種別に係る前記学習データに基づいて前記モデルを学習し、
     前記API検出部は、前記第2のソースコードに係るトークン列を複数の前記モデルに入力することで、前記第2のソースコードに含まれるAPIの構成要素を前記複数のモデルに検出させる、
    ことを特徴とする請求項1記載のAPI検出装置。
  3.  前記複数の記法の種別は、アノテーション記法又は関数記法を含む、
    ことを特徴とする請求項2記載のAPI検出装置。
  4.  第1のソースコードに係るトークン列のうちAPIの構成要素であるトークンに対して所定のラベルが付与された学習データに基づいて、入力されたトークン列の中から前記構成要素を検出するモデルを学習する学習手順と、
     第2のソースコードに係るトークン列を前記モデルへ入力することで、前記第2のソースコードに含まれるAPIの構成要素を前記モデルに検出させるAPI検出手順と、
    をコンピュータが実行することを特徴とするAPI検出方法。
  5.  前記学習手順は、APIに関する複数の記法の種別ごとに、当該種別に係る前記学習データに基づいて前記モデルを学習し、
     前記API検出手順は、前記第2のソースコードに係るトークン列を複数の前記モデルに入力することで、前記第2のソースコードに含まれるAPIの構成要素を前記複数のモデルに検出させる、
    ことを特徴とする請求項4記載のAPI検出方法。
  6.  前記複数の記法の種別は、アノテーション記法又は関数記法を含む、
    ことを特徴とする請求項5記載のAPI検出方法。
  7.  請求項4乃至6いずれか一行記載のAPI検出方法をコンピュータに実行させることを特徴とするプログラム。
PCT/JP2021/019661 2021-05-24 2021-05-24 Api検出装置、api検出方法及びプログラム Ceased WO2022249256A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2023523737A JPWO2022249256A1 (ja) 2021-05-24 2021-05-24
PCT/JP2021/019661 WO2022249256A1 (ja) 2021-05-24 2021-05-24 Api検出装置、api検出方法及びプログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2021/019661 WO2022249256A1 (ja) 2021-05-24 2021-05-24 Api検出装置、api検出方法及びプログラム

Publications (1)

Publication Number Publication Date
WO2022249256A1 true WO2022249256A1 (ja) 2022-12-01

Family

ID=84229739

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2021/019661 Ceased WO2022249256A1 (ja) 2021-05-24 2021-05-24 Api検出装置、api検出方法及びプログラム

Country Status (2)

Country Link
JP (1) JPWO2022249256A1 (ja)
WO (1) WO2022249256A1 (ja)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2024134278A (ja) * 2023-03-20 2024-10-03 Lineヤフー株式会社 情報処理装置、情報処理方法、及び情報処理プログラム
JP2024134162A (ja) * 2023-03-20 2024-10-03 Lineヤフー株式会社 情報処理装置、情報処理方法、及び情報処理プログラム

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019028723A (ja) * 2017-07-31 2019-02-21 日本電信電話株式会社 設計確認装置及び設計確認方法
JP2020091865A (ja) * 2018-12-08 2020-06-11 富士通株式会社 メタデータに基づくapi属性抽出

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210064453A1 (en) * 2019-09-03 2021-03-04 Fujitsu Limited Automated application programming interface (api) specification construction

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019028723A (ja) * 2017-07-31 2019-02-21 日本電信電話株式会社 設計確認装置及び設計確認方法
JP2020091865A (ja) * 2018-12-08 2020-06-11 富士通株式会社 メタデータに基づくapi属性抽出

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2024134278A (ja) * 2023-03-20 2024-10-03 Lineヤフー株式会社 情報処理装置、情報処理方法、及び情報処理プログラム
JP2024134162A (ja) * 2023-03-20 2024-10-03 Lineヤフー株式会社 情報処理装置、情報処理方法、及び情報処理プログラム
JP7828916B2 (ja) 2023-03-20 2026-03-12 Lineヤフー株式会社 情報処理装置、情報処理方法、及び情報処理プログラム

Also Published As

Publication number Publication date
JPWO2022249256A1 (ja) 2022-12-01

Similar Documents

Publication Publication Date Title
US11036937B2 (en) Contraction aware parsing system for domain-specific languages
US12411905B2 (en) Browser extension with automation testing support
US9940104B2 (en) Automatic source code generation
CN108139891B (zh) 用于生成建议以纠正未定义标记错误的方法和系统
US8661416B2 (en) Method and apparatus for defining and instrumenting reusable Java server page code snippets for website testing and production
Liang et al. Neutron: an attention-based neural decompiler
WO2012032890A1 (ja) ソースコード変換方法およびソースコード変換プログラム
JP6878707B2 (ja) 試験装置、試験方法および試験プログラム
TWI826702B (zh) 定義和執行用於指定神經網路架構的程式碼之技術
US9311077B2 (en) Identification of code changes using language syntax and changeset data
Alizadehsani et al. Modern integrated development environment (ides)
US20220374212A1 (en) Indexing and accessing source code snippets contained in documents
US20240354068A1 (en) Video analytics pipeline development system with assistive feedback and annotation
Zhang et al. MobiAgent: A Systematic Framework for Customizable Mobile Agents
Teixeira et al. EasyTest: An approach for automatic test cases generation from UML Activity Diagrams
JP2009205523A (ja) プロパティ自動生成装置
CN121143863A (zh) 代码变更影响分析方法、装置、设备及存储介质
CN119179467B (zh) 一种人工智能辅助编程构建方法
US8819626B2 (en) Sharable development environment bookmarks for functional/data flow
WO2021205589A1 (ja) テストスクリプト生成装置、テストスクリプト生成方法及びプログラム
CN118502735A (zh) 用于代码编辑器的编辑辅助方法、系统及电子设备
Liu et al. Automatic Software Vulnerability Detection in Binary Code
CN116644427A (zh) 一种动静结合的跨架构固件漏洞检测方法及装置
JP6116983B2 (ja) エントリーポイント抽出装置
Shao et al. A survey of available information recovery of binary programs based on machine learning

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21942913

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2023523737

Country of ref document: JP

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 21942913

Country of ref document: EP

Kind code of ref document: A1