WO2021093643A1 - 版权认证方法、装置、设备、系统及计算机可读存储介质 - Google Patents

版权认证方法、装置、设备、系统及计算机可读存储介质 Download PDF

Info

Publication number
WO2021093643A1
WO2021093643A1 PCT/CN2020/126232 CN2020126232W WO2021093643A1 WO 2021093643 A1 WO2021093643 A1 WO 2021093643A1 CN 2020126232 W CN2020126232 W CN 2020126232W WO 2021093643 A1 WO2021093643 A1 WO 2021093643A1
Authority
WO
WIPO (PCT)
Prior art keywords
work
authenticated
review
digital work
preset
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/126232
Other languages
English (en)
French (fr)
Inventor
蔡远航
郑少杰
付勇
范增虎
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
WeBank Co Ltd
Original Assignee
WeBank Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by WeBank Co Ltd filed Critical WeBank Co Ltd
Publication of WO2021093643A1 publication Critical patent/WO2021093643A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F21/00Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
    • G06F21/10Protecting distributed programs or content, e.g. vending or licensing of copyrighted material ; Digital rights management [DRM]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/35Clustering; Classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/60Information retrieval; Database structures therefor; File system structures therefor of audio data
    • G06F16/65Clustering; Classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/24Classification techniques

Definitions

  • This application relates to the field of financial technology (Fintech) technology, and in particular to a copyright authentication method, device, equipment, system, and computer-readable storage medium.
  • Financial technology Fetech
  • the copyright certification of works is mainly completed through offline manual acceptance.
  • the specific process is: 1) the author of the work submits personal information and works to the intellectual property agent; 2) the intellectual property agent judges whether the work meets the registration conditions and determines the copyright Type of registration; 3) If the work meets the requirements for copyright application, the intellectual property agent submits a work registration application form to the Copyright Center; 4) After receiving the application, the Copyright Center will review the application materials and decide whether to issue the copyright registration of the work certificate.
  • the entire process takes about 20-30 working days, and the cycle is long. With the popularization of Internet technology, 100,000 original digital works are generated on the Internet every day.
  • the main purpose of this application is to provide a copyright authentication method, device, equipment, system, and computer-readable storage medium, aiming to shorten the copyright authentication cycle and improve the efficiency of copyright authentication.
  • this application provides a copyright authentication method, the copyright authentication method includes:
  • the digital work to be authenticated is compared with the authenticated digital work in the preset authentication work library to perform copyright authentication.
  • the step of determining the corresponding processing strategy and target review classification model according to the type of the work, and processing the digital work to be authenticated based on the processing strategy to obtain the target input object includes:
  • the corresponding processing strategy is determined to be the first processing strategy, and the target review classification model is determined to be the first review classification model;
  • the step of determining the corresponding processing strategy and target review classification model according to the type of the work, and processing the digital work to be authenticated based on the processing strategy to obtain the target input object includes:
  • the corresponding processing strategy is determined to be the second processing strategy, and the target review classification model is determined to be the second review classification model;
  • Preprocessing the digital work to be authenticated based on the second processing strategy to obtain an input picture wherein the preprocessing includes scaling processing and grayscale processing, and the target input object is the input picture.
  • the step of determining the corresponding processing strategy and target review classification model according to the type of the work, and processing the digital work to be authenticated based on the processing strategy to obtain the target input object includes:
  • the corresponding processing strategy is determined to be the third processing strategy, and the target review classification model is determined to be the third review classification model;
  • the target review classification model includes multiple, the number of the review results is the same as the number of the target review classification models, and the step of judging whether the review is passed based on the review results includes:
  • the step of comparing the digital work to be authenticated with the authenticated digital work in the preset authentication work library to perform copyright authentication includes:
  • the step of calculating the first similarity value between the digital work to be authenticated and the authenticated text work in the preset authentication work library by using a preset document search engine includes:
  • the scores of each word segmentation are added and processed to obtain the first similarity value between the digital work to be authenticated and the authenticated text work in the preset authentication work library.
  • the step of comparing the digital work to be authenticated with the authenticated digital work in the preset authentication work library to perform copyright authentication includes:
  • the step of comparing the digital work to be authenticated with the authenticated digital work in a preset authentication work library to perform copyright authentication includes:
  • the copyright authentication method further includes:
  • the data upload request is sent to the copyright authentication alliance chain, so that the copyright authentication alliance chain completes the upload operation of the digital work to be authenticated based on the data upload request.
  • this application also provides a copyright authentication device, which includes:
  • the first obtaining module is configured to obtain the digital work to be authenticated and the type of the work according to the digital work copyright authentication request when receiving the digital work copyright authentication request;
  • a processing module configured to determine a corresponding processing strategy and a target review classification model according to the type of the work, and process the digital work to be authenticated based on the processing strategy to obtain a target input object;
  • the review module is used to input the target input object into the target review classification model to obtain the review result, and determine whether the review is passed or not based on the review result;
  • the copyright authentication module is used to compare the digital work to be authenticated with the authenticated digital work in the preset authentication work library to perform copyright authentication when the review is passed.
  • the present application also provides a copyright authentication device
  • the copyright authentication device includes: a memory, a processor, and a copyright authentication program stored on the memory and running on the processor, so When the copyright authentication program is executed by the processor, the steps of the copyright authentication method described above are realized.
  • this application also provides a copyright certification system, which includes a copyright certification device and a copyright certification alliance chain; wherein,
  • the copyright authentication device is the copyright authentication device described above;
  • the copyright certification alliance chain is used to receive a data upload request sent by the copyright certification device
  • the present application also provides a computer-readable storage medium storing a copyright authentication program, which when executed by a processor, realizes the above-mentioned copyright authentication Method steps.
  • This application provides a copyright authentication method, device, equipment, system, and computer-readable storage medium.
  • receive a digital work copyright authentication request obtain the digital work to be authenticated and the type of the work according to the digital work copyright authentication request; determine according to the type of work Corresponding processing strategy and target review classification model, process the digital work to be certified based on the processing strategy to obtain the target input object; input the target input object into the target review classification model to obtain the review result, and judge whether the review is passed or not based on the review result; And when the review is passed, the digital work to be authenticated is compared with the authenticated digital work in the preset authentication work library for copyright authentication.
  • FIG. 1 is a schematic diagram of the device structure of the hardware operating environment involved in the solution of the embodiment of the application;
  • FIG. 2 is a schematic flowchart of the first embodiment of the copyright authentication method for applying for
  • Figure 3 is a schematic diagram of the system structure of the copyright certification system for the application.
  • FIG. 4 is a schematic diagram of functional modules of the first embodiment of the copyright authentication device of this application.
  • FIG. 1 is a schematic diagram of the device structure of the hardware operating environment involved in the solution of the embodiment of the application.
  • the copyright authentication device in the embodiment of the application may be a smart phone, or a terminal device such as a PC (Personal Computer, personal computer), a tablet computer, and a portable computer.
  • a terminal device such as a PC (Personal Computer, personal computer), a tablet computer, and a portable computer.
  • the copyright authentication device may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005.
  • the communication bus 1002 is used to implement connection and communication between these components.
  • the user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface.
  • the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface).
  • the memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as a magnetic disk memory.
  • the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
  • FIG. 1 does not constitute a limitation on the copyright authentication device, and may include more or less components than those shown in the figure, or a combination of certain components, or different components. Layout.
  • a memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a copyright authentication program.
  • the network interface 1004 is mainly used to connect to a back-end server and communicate with the back-end server;
  • the user interface 1003 is mainly used to connect to a client and communicate with the client;
  • the processor 1001 can be used to Call the copyright authentication program stored in the memory 1005, and execute the following steps of the copyright authentication method.
  • This application provides a copyright authentication method.
  • FIG. 2 is a schematic flowchart of the first embodiment of the copyright authentication method of this application.
  • the copyright authentication method includes:
  • Step S10 when receiving the digital work copyright certification request, obtain the digital work to be certified and the work type according to the digital work copyright certification request;
  • the copyright authentication method is applied to a copyright authentication system.
  • the copyright authentication system includes a copyright authentication device and a copyright authentication alliance chain.
  • the copyright authentication method in this embodiment is implemented by the copyright authentication device. Equipped with a copyright authentication system.
  • the copyright certification alliance chain can be composed of copyright certification agency nodes, notary agency nodes, judicial agency nodes, and external nodes. It is used to receive the data upload request sent by the copyright certification device, and then obtain the data to be uploaded based on the data upload request. Then, based on the consensus algorithm, the chain operation of the data information to be chained is completed, that is, the copyright authentication based on the blockchain is realized.
  • the copyright certification system when users need to perform copyright authentication on their works, they can upload their digital works (such as text works, picture works, audio works, etc.) through the corresponding software on the user terminal (such as PC personal computer, smart phone, etc.) , And fill in relevant information (including but not limited to the type of work, author information, etc.) to trigger a digital work copyright certification request.
  • the copyright certification system receives the digital work copyright certification request, it will obtain it according to the digital work copyright certification request Digital works to be certified and types of works.
  • users can only upload their digital works when they trigger a digital work copyright authentication request.
  • the copyright authentication system can determine the corresponding work type according to its format after obtaining the digital work to be authenticated. .
  • Step S20 Determine a corresponding processing strategy and a target review classification model according to the work type, and process the digital work to be authenticated based on the processing strategy to obtain a target input object;
  • the corresponding processing strategy and target review classification model are determined according to the work type, and the digital work to be certified is processed based on the processing strategy to obtain the target input object.
  • the corresponding processing strategy is determined as the first processing strategy
  • the target review classification model is determined as the first review classification model
  • the digital work to be certified is word-cut based on the first processing strategy , Get the first word text; input the first word text into the preset word vector model to get the first word vector of each segment in the first word text; then get the corresponding digital work to be authenticated according to the first word vector
  • the corresponding processing strategy is determined to be the second processing strategy, and the target review classification model is determined to be the second review classification model; the digital work to be certified is preprocessed based on the second processing strategy to obtain the input picture, where , Preprocessing includes scaling and grayscale processing, and the target input object is the input picture.
  • the corresponding processing strategy is determined to be the third processing strategy, and the target review classification model is determined to be the third review classification model; based on the third processing strategy, the digital work to be certified is converted into a text work type, and the conversion is obtained After the digital work to be certified; then process it according to the processing method of the text work, that is, perform word segmentation processing on the converted digital work to be certified to obtain the second word segmentation text; and then enter the second word segmentation text into the preset
  • the word vector model obtains the second word vector of each word in the second word segmentation text; obtains the second document vector corresponding to the converted digital work to be authenticated according to the second word vector, where the target input object is the second document vector .
  • Step S30 input the target input object into the target review classification model to obtain the review result, and judge whether the review is passed or not based on the review result;
  • the target input object is input into the target review classification model, and the review result is obtained, and based on the review result, it is judged whether the review is passed.
  • the main purpose of the review is to review whether the work submitted by the user involves terrorist propaganda, cult publications, political sensitivity, pornography, gambling, and drugs.
  • the corresponding target review classification model can include 6 categories. That is, the target review classification model may include a terrorist propaganda review classification model, a cult publications review classification model, a politically sensitive review classification model, a pornographic review classification model, a gambling-related review classification model, and a drug-related review classification model.
  • the number of review results is the same as the number of target review classification models, that is, the corresponding review results also include multiple.
  • the review is judged to be passed at this time; if at least one of the review results is unqualified, it means that the digital work to be certified involves bad information (terrorist publications/cult publications /Political Sensitive/Yellow/Gambling/Drugs). At this time, it is judged that the review is not passed.
  • an examination classification can also be constructed for each type of digital work, and correspondingly, there is only one examination result. In this case, it is only necessary to check whether the examination result is qualified.
  • Step S40 When the review is passed, the digital work to be authenticated is compared with the authenticated digital work in the preset authentication work library to perform copyright authentication.
  • the digital work to be authenticated is compared with the authenticated digital work in the preset authentication work library for copyright authentication.
  • different copyright authentication methods need to be adopted for different types of digital works to be authenticated.
  • For the specific copyright authentication process refer to the following fourth embodiment, which will not be repeated here.
  • the embodiment of the application provides a copyright authentication method.
  • a digital work copyright authentication request is received, the digital work to be authenticated and the work type are obtained according to the digital work copyright authentication request; the corresponding processing strategy and target review classification model are determined according to the work type, Process the digital work to be certified based on the processing strategy to obtain the target input object; input the target input object into the target review classification model to obtain the review result, and determine whether the review is passed based on the review result; and when the review is passed, the pending certification
  • the digital works are compared with the certified digital works in the preset certified works library for copyright certification.
  • step S20 may include:
  • Step a11 if the work type is a text work, determine the corresponding processing strategy as the first processing strategy, and determine the target review classification model as the first review classification model;
  • Step a12 performing word segmentation processing on the digital work to be authenticated based on the first processing strategy to obtain the first word text
  • Step a13 input the first word text into a preset word vector model to obtain the first word vector of each word segment in the first word text;
  • Step a14 Obtain a first document vector corresponding to the digital work to be authenticated according to the first word vector, wherein the target input object is the first document vector.
  • the corresponding processing strategy is determined as the first processing strategy
  • the target review classification model is determined as the first review classification model.
  • the first review classification model is pre-trained, and its type can be SVM (Support Vector Machine) model, Bayesian model, logistic regression model, convolutional neural network model and other two classification models, as follows
  • SVM Small Vector Machine
  • the target review classification model includes the terrorist propaganda review classification model, the cult propaganda review classification model, the politically sensitive review classification model, the pornographic review classification model, the gambling-related review classification model, and the drug-related review classification model.
  • the 6 categories of the review classification model are taken as examples to illustrate.
  • the first review classification model also includes 6 categories.
  • the training process is as follows: Take respectively 50,000 labeled text works that involve and 50,000 that do not involve terrorist propaganda. After cutting each word work, the word vector is obtained through a preset word vector model (in one embodiment, the word2vec model), and the word vectors are added according to the corresponding dimensions to obtain the document vector corresponding to each word work. Then train an SVM classification model based on the above 100,000 document vectors. Next, use the same method to train five SVM classification models to determine whether the work involves cult politicians, political sensitivity, pornography, gambling, and drugs.
  • the digital work to be certified is segmented based on the first processing strategy to obtain the first word text.
  • the word segmentation can be processed with preset tools, such as Chinese Academy of Sciences NLPIR, Harbin Institute of Technology LTP, and stammering. Participles and so on. The specific word segmentation process is consistent with the prior art, and will not be repeated here.
  • the preset word vector model is word2vec (word to vector, a correlation model used to generate word vectors) in one embodiment, and word2vec maps each Chinese vocabulary to a high-dimensional vector (usually a 200-dimensional vector), and For any two Chinese words, the closer they are semantically, the closer the vector distance after mapping. Therefore, the semantic similarity of Chinese vocabulary can be described according to the distance between word vectors.
  • the first document vector corresponding to the digital work to be authenticated is obtained according to the first word vector, where the target input object is the first document vector, that is, the first document vector is subsequently input into the corresponding first review classification model to obtain Review the results.
  • the first word vector is added according to the corresponding dimensions to obtain the corresponding first document vector.
  • step S20 may further include:
  • Step a21 if the work type is a picture work, determine the corresponding processing strategy as the second processing strategy, and determine the target review classification model as the second review classification model;
  • Step a22 preprocessing the digital work to be authenticated based on the second processing strategy to obtain an input picture, wherein the preprocessing includes scaling processing and grayscale processing, and the target input object is the input picture.
  • the corresponding processing strategy is determined to be the second processing strategy
  • the target review classification model is determined to be the second review classification model.
  • the second review classification model is pre-trained, and its type is a classification model based on a convolutional neural network in one embodiment.
  • the target review classification model includes the terrorist propaganda review classification model, the cult propaganda review classification model, the politically sensitive review classification model, the pornographic review classification model, the gambling-related review classification model, and the drug-related review classification model.
  • the 6 categories of the review classification model are taken as examples to illustrate.
  • the second review classification model also includes 6 categories.
  • the training process is as follows: Take respectively 50,000 labeled pictures related to and 50,000 pictures that do not involve terrorist propaganda. Preprocess each picture work.
  • the preprocessing process includes scaling and grayscale processing.
  • the scaling process refers to scaling the size of the picture to a preset size, such as 128 pixels*128 pixels.
  • the grayscale processing is about to be scaled. Convert the pictures into grayscale pictures, and then train a classification model based on convolutional neural network based on the above 100,000 preprocessed pictures. Next, use the same method to train 5 classification models based on convolutional neural networks to determine whether the work involves cult propaganda, political sensitivity, pornography, gambling, and drugs.
  • the digital work to be authenticated is preprocessed based on the second processing strategy to obtain the input picture.
  • the preprocessing includes scaling and grayscale processing.
  • the scaling process is to scale the size of the picture to a preset Set the size, for example, 128 pixels*128 pixels.
  • the grayscale processing is to convert the scaled image into a grayscale image, and the target input object is the input image, that is, the input image is subsequently input into the corresponding second review classification model to get the review result.
  • step S20 may further include:
  • Step a31 if the work type is an audio work, determine that the corresponding processing strategy is the third processing strategy, and determine that the target review classification model is the third review classification model;
  • Step a32 converting the digital work to be authenticated into a text work type based on the third processing strategy, to obtain the converted digital work to be authenticated;
  • Step a33 performing word segmentation processing on the converted digital work to be authenticated to obtain a second word segmentation text
  • Step a34 input the second word segmentation text into a preset word vector model to obtain the second word vector of each word segmentation in the second word segmentation text;
  • Step a35 Obtain a second document vector corresponding to the converted digital work to be authenticated according to the second word vector, wherein the target input object is the second document vector.
  • the processing procedure for audio works is as follows:
  • the corresponding processing strategy is determined to be the third processing strategy
  • the target review classification model is determined to be the third review classification model.
  • the third review classification model is pre-trained, and its type can be two classification models such as SVM (Support Vector Machine) model, Bayesian model, logistic regression model, and convolutional neural network model.
  • SVM Small Vector Machine
  • the third inspection classification model may be the same as the first inspection classification model, or other types of two classification models obtained by training based on the training method of the first inspection classification model.
  • the digital work to be authenticated is converted into a text work type based on the third processing strategy, and the converted digital work to be authenticated is obtained.
  • the audio work can be converted into a text work type through a voice recognition tool.
  • perform word segmentation processing on the converted digital work to be authenticated to obtain the second segmentation text where the word segmentation processing can use preset tools, such as Chinese Academy of Sciences NLPIR, Harbin Institute of Technology LTP, and stutter segmentation.
  • preset tools such as Chinese Academy of Sciences NLPIR, Harbin Institute of Technology LTP, and stutter segmentation.
  • the second word segmentation text is input into the preset word vector model to obtain the second word vector of each word segmentation in the second word segmentation text.
  • the preset word vector model is word2vec (word to vector, a correlation model used to generate word vectors) in one embodiment.
  • the second document vector corresponding to the digital work to be authenticated is obtained according to the second word vector, where the target input object is the second document vector, that is, the second document vector is subsequently input into the third review classification model corresponding to the value to obtain Review the results.
  • the second word vector is added according to the corresponding dimensions to obtain the corresponding second document vector.
  • the target review classification model includes multiple, and the number of the review results is the same as the number of the target review classification models, and the step of "determining whether the review is passed based on the review result" includes:
  • Step b1 detecting whether a plurality of the inspection results are all qualified
  • Step b2 if multiple of the review results are all qualified, then the review is determined to be passed;
  • step b3 if at least one of the multiple review results is unqualified, it is determined that the review is not passed.
  • the target review classification model may include multiple, and the number of review results is the same as the number of target review classification models, and the corresponding review results also include multiple.
  • the process of judging whether the review is passed is: testing whether multiple review results are all qualified.
  • At least one of the multiple review results is unqualified, it means that the digital work to be certified involves one or more of bad information (terrorist publications/cult publications/political sensitive/pornography/gambling/drugs). Determined that the review was not passed.
  • step S40 includes:
  • Step c11 Calculate the first similarity value between the digital work to be authenticated and the authenticated text work in the library of preset authentication works by using a preset document search engine;
  • Step c12 screening and obtaining a first preset number of similar text works from the authenticated text works according to the first similarity value
  • Step c13 Calculate the first longest common subsequence between the similar text work and the digital work to be authenticated, and calculate the similar text work and the to-be-certified digital work according to the length of the first longest common subsequence The length ratio between certified digital works is obtained, and the first calculation result is obtained;
  • Step c14 detecting whether there is a length ratio greater than a first preset threshold in the first calculation result
  • Step c15 if there is a length ratio greater than the first preset threshold, it is determined that the copyright authentication is not passed;
  • step c16 if there is no length ratio greater than the first preset threshold, it is determined that the copyright authentication is passed.
  • step c11 includes:
  • Step c111 performing word segmentation processing on the digital work to be authenticated through a preset document search engine to obtain a word segmentation set
  • Step c112 Perform an inverted index on the authenticated text works in the preset authenticated work library through the preset document search engine, and calculate the score corresponding to each word in the word segmentation set according to the inverted index result;
  • step c113 the scores of each word segmentation are added and processed to obtain the first similarity value between the digital work to be authenticated and the authenticated text work in the preset authentication work library.
  • the first similarity value between the digital work to be authenticated and the authenticated text work in the library of preset authentication works is calculated by the preset document search engine.
  • the preset document search engine is an ES (Elastic Search) search engine in one embodiment.
  • ES is a distributed, high-scalable, high-real-time search and data analysis engine, which can easily make a large amount of data Have the ability to search, analyze and explore.
  • the ES search engine is used to perform word segmentation processing on the digital work to be certified to obtain the word segmentation set, where the ES search engine has its own word segmentation device that can perform word segmentation on the digital work to be certified.
  • the inverted index is performed on the authenticated text works in the preset authentication library through the ES search engine, and the score corresponding to each word in the word segmentation set is calculated according to the inverted index result.
  • the word segmentation of the ES search engine is also performed first, and the word frequency information and location information of each word segmentation will be obtained, and then the word segmentation and the work document will be established.
  • the inverted index is a dictionary data structure (key-value), the key of the dictionary is a word segmentation, and the value (value) is a list of works containing the word segmentation, and the word segmentation is The location information and word frequency information in each work, through the inverted index, can quickly obtain a list of documents containing the word segmentation and word frequency information according to the word segmentation.
  • the document containing the word segmentation and the word frequency and inverse document frequency of the word segmentation in the document can be obtained according to the inverted index result, and each word frequency and inverse document frequency are calculated according to the The score of the word segmentation.
  • a word segmentation set containing the word segmentation "A", “B”, “C”, “D”, “E” and “F” is obtained, which is first found based on the inverted index of the certified text work
  • the portfolio contains the participle "A”, and the score corresponding to the participle "A” is calculated.
  • the score of each work in the portfolio can be calculated first.
  • the score of each work can be the word frequency and the word frequency of the participle A in the work.
  • the product of the inverse document frequency of word A (of course, other calculation methods can also be set according to the actual situation, such as calculation based on word frequency, inverse document frequency, and location information), and then add the scores of each work to obtain the word "A” Corresponding score; after getting the score corresponding to the participle "A”, perform the same operation on the participles "B", “C”, “D”, “E”, and “F” to get the participles "B", "C”, “D”, “E”, and “F” correspond to scores respectively.
  • the certified text works in the preset certified works library can be stored in advance using the ES search engine according to a specific index structure.
  • the certified text works in the library are inverted index. After the score corresponding to each word segmentation is obtained, the scores of each word segmentation are added and processed to obtain the first similarity value between the digital work to be authenticated and the authenticated text work in the preset authentication work library.
  • the first preset number of similar text works is obtained from the authenticated text works according to the first similarity value.
  • the set number can be set according to actual needs. For example, it can be set to a preset value, such as 1000, or it can be set to a number where the first similarity value is greater than a preset value, which is not specifically limited here.
  • screening the first similarity value is sorted in descending order, and the first preset number of text works ranked in the front are taken as similar text works. It needs to be explained that the premise of considering the plagiarism relationship between the two literary works is that the words used in the two works are the same, and the words before and after each word are also the same.
  • the ES search engine is used to preliminarily determine whether the words used are consistent for preliminary screening.
  • the purpose is to filter out the text works that contain all or most of the word cuts in the text work to be certified. Reduce the scope of authentication comparison, save server resources, and further improve the efficiency of copyright authentication of text works.
  • the ES search engine first performs a preliminary screening, and then Combined with the subsequent calculation of the longest common subsequence, compared to the method of calculating the similarity value directly based on the longest common subsequence, it can further improve the efficiency of copyright authentication of works.
  • the longest common subsequence is to find the longest common subsequence of two sequences.
  • the length ratio between the similar text work and the digital work to be authenticated is calculated according to the length of the first longest common subsequence to obtain the first calculation result.
  • the number of the first longest common subsequence corresponds to the same number of similar textual works.
  • the first preset threshold can be set according to actual needs, for example, it can be set to 0.8, which is not specifically limited here. If there is a length ratio greater than the first preset threshold, then it is considered that there is a citation/plagiarism relationship, and the copyright authentication is determined to be failed at this time; if there is no length ratio greater than the first preset threshold, the copyright authentication is determined to be passed.
  • Sequence and then calculate the length ratio between the similar text work and the digital work to be authenticated according to the length of the first longest common subsequence, and then whether the length ratio is greater than the first preset threshold, once a similar text work corresponding to a certain similar text work is detected If the length ratio is greater than the first preset threshold, it is considered that there is a citation/plagiarism relationship, and the copyright certification can be determined to be unsuccessful. In this case, there is no need to calculate the first longest common subsequence between other similar textual works and digital works to be certified. The subsequent steps can save server resources and further improve the efficiency of copyright authentication.
  • step S40 may further include:
  • Step c21 Calculate the second similarity value between the digital work to be authenticated and the authenticated photo work in the preset authentication work library by using a preset image search engine;
  • Step c22 selecting a second preset number of similar picture works from the certified picture works according to the second similarity value
  • Step c23 extract the first scale-invariant feature transform SIFT feature vector of the digital work to be authenticated, and extract the second SIFT feature vector of the similar picture work;
  • Step c24 Calculate the cosine distance between the first SIFT feature vector and the second SIFT feature vector to obtain a second calculation result
  • Step c25 detecting whether there is a cosine distance greater than a second preset threshold in the second calculation result
  • Step c26 if there is a cosine distance greater than the second preset threshold, it is determined that the copyright authentication is not passed;
  • Step c27 If there is no cosine distance greater than the second preset threshold, it is determined that the copyright authentication is passed.
  • the second similarity value between the digital work to be authenticated and the authenticated image work in the preset authentication work library is calculated by the preset image search engine.
  • the preset image retrieval engine is a CBIR (Content-based image retrieval) engine in one embodiment, and the core of the CBIR engine is to retrieve images using visual features of the image.
  • CBIR Content-based image retrieval
  • the core of the CBIR engine is to retrieve images using visual features of the image.
  • it is a kind of approximate matching technology, which combines the technical achievements of computer vision, image processing, image understanding and database.
  • the feature extraction and index establishment can be completed by the computer automatically, avoiding the manual description. Subjectivity.
  • the process of user retrieval is generally to provide a sample image (Queryby Example) or draw a sketch (Queryby Sketch), the system extracts the characteristics of the query image, and then compares with the characteristics in the database, and compares the image with the query characteristics. Return to the user. It should be noted that since the authenticated photo works in the preset authentication work library are all stored in a preset size, it is necessary to perform corresponding scaling processing on the digital works to be authenticated in the image category to obtain the same preset size to be authenticated For the digital work, the second similarity value between the zoomed digital work to be authenticated and the authenticated picture work in the preset authentication work library is calculated through the preset image retrieval engine.
  • a second preset number of similar image works is obtained from the authenticated image works according to the second similarity value.
  • the set number can be the same as or different from the first preset number, and can be set according to actual needs, which is not specifically limited here.
  • the second similarity value is sorted in descending order, and the second preset number of text works with the top ranking are taken as similar text works. It should be noted that the premise of considering the plagiarism relationship between two photo works is that the color, shape, texture and other low-level features of the two photo works are consistent, and the arrangement of the features is also consistent.
  • the purpose of the preliminary screening through the CBIR engine here is to filter out the certified image works that contain more features (such as color features, shape features, texture features, etc.) in the image works to be certified, so as to reduce the certification ratio In the right scope, it saves server resources and further improves the efficiency of copyright certification for image works.
  • the similarity value between the digital work to be authenticated and the authenticated image work in the preset authentication work library can also be calculated directly based on the SIFT feature vector, and then based on the similarity.
  • the degree value determines whether the copyright authentication is passed, but in comparison, the feature extraction and similarity calculation through the CBIR engine is simpler and more efficient than extracting the SIFT feature vector. Therefore, in this embodiment
  • the CBIR engine is used for preliminary screening first, and then combined with the subsequent calculation of the similarity value based on the SIFT feature vector. Compared with the method of directly calculating the similarity value based on the SIFT feature vector, it can further improve the copyright certification efficiency of the work.
  • SIFT Scale-invariant feature transform
  • the SIFT feature vector extraction process is: detect a series of key points in the image, these key points have nothing to do with scale scaling, rotation and brightness change, and then assign the gradient direction value to the above key points, you can get an image The SIFT feature vector.
  • the specific extraction process is consistent with the prior art and will not be described in detail here.
  • the cosine distance between the first SIFT feature vector and the second SIFT feature vector is calculated to obtain the second calculation result.
  • the cosine distance is used to characterize the similarity between two feature vectors.
  • other parameters such as Euclidean distance, can also be used to characterize the similarity.
  • the process of "distance” and "detecting whether there is a cosine distance greater than the second preset threshold” can be carried out at the same time. Once it is detected that the cosine distance corresponding to a similar picture work is greater than the second preset threshold, it is considered that there is a citation/plagiarism relationship. It can be judged that the copyright authentication is not passed.
  • step S40 may further include:
  • Step c31 converting the digital work to be authenticated into a text work type to obtain an audio text work to be authenticated
  • Step c32 Calculate the third similarity value between the audio text work to be authenticated and the authenticated audio text work in the preset authentication work library by using a preset document search engine;
  • Step c33 searching for a third preset number of similar audio text works from the authenticated audio text works according to the third similarity value
  • Step c34 Calculate the second longest common subsequence between the similar audio file work and the audio text work to be authenticated, and calculate the similar audio text work and the similar audio text work according to the length of the second longest common subsequence The length ratio between the audio text works to be authenticated to obtain the third calculation result;
  • Step c35 detecting whether there is a length ratio greater than a third preset threshold in the third calculation result
  • Step c36 if there is a length ratio greater than the third preset threshold, it is determined that the copyright authentication is not passed;
  • Step c37 If there is no length ratio greater than the third preset threshold, it is determined that the copyright authentication is passed.
  • the corresponding copyright authentication process is as follows:
  • the preset document search engine is an ES search engine in one embodiment.
  • the certified audio text works in the preset certified works library are obtained by converting the audio works into text works based on the voice recognition tool, and can be stored in a specific index structure using the ES search engine in advance to facilitate subsequent searches.
  • the authenticated audio text works After obtaining the third similarity value between the digital work to be authenticated and the authenticated audio text work, filter the authenticated audio text works according to the third similarity value to obtain a third preset number of similar text works, where the first 3.
  • the preset quantity can be the same as or different from the first preset quantity and the second preset quantity, and can be set according to actual needs, and there is no specific limitation here.
  • the third similarity value is sorted in descending order, and the third preset number of text works with the top ranking are selected as similar text works. It should be noted that the purpose of screening here is to screen out the text works that contain all or most of the word segmentation in the audio text works to be certified from the certified audio text works, so as to narrow the scope of authentication comparison and save server resources. Further improve the efficiency of copyright certification.
  • the "second" in the second longest common subsequence has no substantive meaning and is only used in conjunction with the above-mentioned first common subsequence.
  • a longest announcement sub-sequence is distinguished.
  • the third preset threshold may be the same as or different from the first preset threshold and the second preset threshold, and may be set according to actual needs, for example, it may be set to 0.8, which is not specifically limited here. If there is a length ratio greater than the third preset threshold, it is considered that there is a citation/plagiarism relationship, and the copyright certification is determined to be failed at this time; if there is no length ratio greater than the third preset threshold, the copyright certification is determined to be passed.
  • the process of "length ratio between certified digital works” and “detecting whether there is a length ratio greater than the third preset threshold” can be carried out simultaneously, that is, the second most common audio text works between each similar audio text work and the digital work to be certified are calculated in turn.
  • the copyright authentication method may further include:
  • Step A Obtain the work information of the digital work to be authenticated when the copyright authentication is passed;
  • the work information of the digital work to be authenticated is obtained.
  • some of the information in the work information (such as the author of the work, the type of the work, etc.) can be obtained according to the copyright authentication request of the digital work.
  • the corresponding prompt window can be generated and obtained after the user fills in the corresponding work information based on the prompt window; another part of the information can be generated by the copyright authentication system, such as the authentication time and the md5 (Message-Digest Algorithm) value of the work, where The authentication time can directly obtain the time when the copyright authentication is passed.
  • the md5 value of the work can be obtained through the corresponding program after obtaining the digital work to be authenticated.
  • the work information can include, but is not limited to: the author of the work, the type of work, the verification time, and the md5 value of the digital work to be certified.
  • Step B Generate corresponding copyright authentication information based on the work information, and generate a data upload request based on the copyright authentication information;
  • the corresponding copyright authentication information is generated based on the work information.
  • the work information can be converted into a preset data format, such as json (a lightweight data exchange format) data structure, to obtain the copyright authentication information.
  • a data upload request is generated according to the copyright authentication information.
  • sha256 Secure Hash Algorithm 256
  • the hash value generates a request for data on the chain.
  • Step C Send the data upload request to the copyright certification alliance chain, so that the copyright certification alliance chain completes the upload operation of the digital work to be authenticated based on the data upload request.
  • the data upload request is sent to the copyright authentication alliance chain, so that the copyright authentication alliance chain completes the upload operation of the digital work to be authenticated based on the data upload request.
  • the copyright certification alliance chain is mainly composed of copyright certification agency nodes, notary agency nodes, judicial agency nodes, and external nodes. Data upload requests can be sent to external nodes in the copyright certification alliance chain, and then all nodes are based on The general consensus algorithm completes the chaining operation of the digital works to be authenticated, that is, the hash value in the data chaining request is written into the consortium for permanent retention.
  • the digital work to be authenticated that has passed the copyright authentication can also be stored in the preset authenticated work library for testing. Whether the subsequent work has plagiarized the authenticated work; at the same time, a message indicating that the copyright authentication is successful is returned to the user terminal, and the user has been notified.
  • the copyright authentication system when it determines that the copyright authentication is passed, it can generate a corresponding data upload request and send it to the copyright authentication alliance chain, so that the copyright authentication alliance chain completes the upload operation based on the data upload request.
  • Blockchain has unmodifiable characteristics, which can realize the protection of the copyright of digital works. Once there is an unauthorized dissemination of other people’s works in the future, the copyright owner can appeal the infringement based on the authentication information on the blockchain. , To reduce the difficulty of protecting rights.
  • this application also provides a copyright authentication system.
  • Figure 3 is a schematic diagram of the system architecture of the copyright authentication system of the application.
  • the copyright authentication system includes a copyright authentication device and a copyright authentication alliance chain; of course, it may also include a user terminal.
  • the copyright authentication device is the copyright authentication device shown in FIG. 1; it is used to execute the steps in the above copyright authentication method embodiment. For specific functions and implementation processes, please refer to the above embodiment, which will not be repeated here.
  • the copyright certification alliance chain is used to receive a data upload request sent by the copyright certification device
  • the copyright certification alliance chain can be used to receive data on-chain requests sent by the copyright certification device.
  • the copyright certification alliance chain is mainly composed of copyright certification authority nodes, notary authority nodes, judicial authority nodes, and external nodes.
  • the data on-chain request sent by the copyright authentication device can be received through the external link node. Then obtain the data information to be uploaded based on the data upload request, where the data information to be uploaded can be a hash value generated based on the copyright authentication information, and then complete the operation of the hash value based on the consensus algorithm.
  • the hash value is written into the consortium chain and kept forever.
  • the external node in order to ensure the authority and security of copyright authentication, usually only respond to requests from copyright authentication devices.
  • the external node when it receives the data upload request, it can obtain the corresponding device IP (Internet Protocol, Internet Protocol address), and then check whether the device IP is the IP of the preset copyright authentication device, if so, obtain the data information to be uploaded based on the data upload request, and complete the upload operation of the data information to be uploaded based on the consensus algorithm .
  • IP Internet Protocol, Internet Protocol address
  • the online authentication of the copyright of digital works can be realized through the copyright authentication equipment, and the copyright authentication can be completed within seconds of the time complexity for different types of digital works, which is compared with the manual process in the prior art.
  • Copyright certification this application can reduce labor costs, shorten the copyright certification cycle, and improve the efficiency of copyright certification, so that the copyright of the author of the work can be protected in time.
  • self-media people or creators do not need to apply for and join the blockchain platform to realize copyright certification and protection, thereby reducing the creator’s use cost, and at the same time, others cannot obtain the creator’s behavioral operations. , To ensure data privacy.
  • This application also provides a copyright authentication device.
  • FIG. 4 is a schematic diagram of the functional modules of the first embodiment of the copyright authentication device of this application.
  • the copyright authentication device includes:
  • the first obtaining module 10 is configured to obtain the digital work to be authenticated and the type of the work according to the digital work copyright authentication request when receiving the digital work copyright authentication request;
  • the processing module 20 is configured to determine a corresponding processing strategy and a target review classification model according to the work type, and process the digital work to be authenticated based on the processing strategy to obtain a target input object;
  • the review module 30 is configured to input the target input object into the target review classification model to obtain the review result, and determine whether the review is passed or not based on the review result;
  • the copyright authentication module 40 is configured to compare the digital work to be authenticated with the authenticated digital work in the preset authentication work library to perform copyright authentication when the review is passed.
  • processing module 20 is specifically configured to:
  • the corresponding processing strategy is determined to be the first processing strategy, and the target review classification model is determined to be the first review classification model;
  • processing module 20 is also specifically configured to:
  • the corresponding processing strategy is determined to be the second processing strategy, and the target review classification model is determined to be the second review classification model;
  • Preprocessing the digital work to be authenticated based on the second processing strategy to obtain an input picture wherein the preprocessing includes scaling processing and grayscale processing, and the target input object is the input picture.
  • processing module 20 is also specifically configured to:
  • the corresponding processing strategy is determined to be the third processing strategy, and the target review classification model is determined to be the third review classification model;
  • the target review classification model includes multiple, the number of the review results is the same as the number of the target review classification model, and the review module 30 is specifically configured to:
  • the copyright authentication module 40 is specifically configured to:
  • copyright authentication module 40 is also specifically configured to:
  • the scores of each word segmentation are added and processed to obtain the first similarity value between the digital work to be authenticated and the authenticated text work in the preset authentication work library.
  • the copyright authentication module 40 is also specifically used for:
  • the copyright authentication module 40 is also specifically configured to:
  • the copyright authentication device further includes:
  • the second obtaining module is used to obtain the work information of the digital work to be authenticated when the copyright authentication is passed;
  • a generating module configured to generate corresponding copyright authentication information based on the work information, and generate a data upload request based on the copyright authentication information
  • the sending module is configured to send the data upload request to the copyright authentication alliance chain, so that the copyright authentication alliance chain completes the upload operation of the digital work to be authenticated based on the data upload request.
  • each module in the above copyright authentication device corresponds to each step in the above copyright authentication method embodiment, and the functions and realization process thereof will not be repeated here.
  • the present application also provides a computer-readable storage medium that stores a copyright authentication program, and when the copyright authentication program is executed by a processor, the copyright authentication method as described in any of the above embodiments is implemented. step.
  • the technical solution of this application essentially or the part that contributes to the existing technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium (such as ROM/RAM) as described above. , Magnetic disks, optical disks), including several instructions to make a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) execute the method described in each embodiment of the present application.
  • a terminal device which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Data Mining & Analysis (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Software Systems (AREA)
  • Multimedia (AREA)
  • Databases & Information Systems (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Artificial Intelligence (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Computation (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Technology Law (AREA)
  • Computer Hardware Design (AREA)
  • Computer Security & Cryptography (AREA)
  • Collating Specific Patterns (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)
  • Storage Device Security (AREA)

Abstract

一种版权认证方法、装置、设备、系统及计算机可读存储介质,涉及金融科技技术领域。该版权认证方法包括:在接收到数字作品版权认证请求时,根据所述数字作品版权认证请求获取待认证数字作品和作品类型(S10);根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象(S20);将所述目标输入对象输入至所述目标审查分类模型中,得到审查结果,并基于所述审查结果判断审查是否通过(S30);在审查通过时,将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证(S40)。

Description

版权认证方法、装置、设备、系统及计算机可读存储介质
优先权信息
本申请要求于2019年11月11日申请的、申请号为201911093190.5、名称为“版权认证方法、装置、设备、系统及计算机可读存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及金融科技(Fintech)技术领域,尤其涉及一种版权认证方法、装置、设备、系统及计算机可读存储介质。
背景技术
随着计算机技术的发展,越来越多的技术(大数据、分布式、区块链Blockchain、人工智能等)应用在金融领域,传统金融业正在逐步向金融科技(Fintech)转变,但由于金融行业的安全性、实时性要求,也对技术提出了更高的要求。
目前作品版权认证主要通过线下人工受理的方式完成,具体流程为:1)作品著作人向知识产权代理人提交个人信息及作品;2)由知识产权代理人判断作品是否具备登记条件,判定版权登记类型;3)若作品符合版权申请条件,由知识产权代理人向版权中心提交作品登记申请表;4)版权中心接到申请后,对申请资料进行审核,并决定是否下发作品的著作权登记证书。整个流程大约需要20~30个工作日,周期较长。随着互联网技术的普及,网络上每天都会产生十万级的原创数字作品,线下人工受理作品版权认证的方式由于存在耗时耗力、成本高、效率低和周期长等缺陷,已经无法满足当下的需求,使得大量侵权盗版作品在网络上传播。因此,亟需一种数字作品的版权认证方法,以缩短版权认证周期、提高版权认证效率,及时对作品著作人的版权进行保护。
发明内容
本申请的主要目的在于提供一种版权认证方法、装置、设备、系统及计算机可读存储介质,旨在缩短版权认证周期,提高版权认证效率。
为实现上述目的,本申请提供一种版权认证方法,所述版权认证方法包括:
在接收到数字作品版权认证请求时,根据所述数字作品版权认证请求获取待认证数字作品和作品类型;
根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象;
将所述目标输入对象输入至所述目标审查分类模型中,得到审查结果,并基于所述审查结果判断审查是否通过;
在审查通过时,将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证。
在一实施例中,所述根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象的步骤包括:
若所述作品类型为文字作品,则确定对应的处理策略为第一处理策略,并确定目标审查分类模型为第一审查分类模型;
基于所述第一处理策略对所述待认证数字作品进行切词处理,得到第一切词文本;
将所述第一切词文本输入预设词向量模型,得到所述第一切词文本中各切词的第一词向量;
根据所述第一词向量得到所述待认证数字作品对应的第一文档向量,其中,所述目标输入对象为所述第一文档向量。
在一实施例中,所述根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象的步骤包括:
若所述作品类型为图片作品,则确定对应的处理策略为第二处理策略,并确定目标审查分类模型为第二审查分类模型;
基于所述第二处理策略对所述待认证数字作品进行预处理,得到输入图片,其中,所述预处理包括缩放处理和灰度处理,所述目标输入对象为所述输入图片。
在一实施例中,所述根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象的步骤包括:
若所述作品类型为音频作品,则确定对应的处理策略为第三处理策略,并确定目标审查分类模型为第三审查分类模型;
基于所述第三处理策略将所述待认证数字作品转换成文字作品类型,得到转换后的待认证数字作品;
对所述转换后的待认证数字作品进行切词处理,得到第二切词文本;
将所述第二切词文本输入预设词向量模型,得到所述第二切词文本中各切词的第二词向量;
根据所述第二词向量得到所述转换后的待认证数字作品对应的第二文档向量,其中,所述目标输入对象为所述第二文档向量。
在一实施例中,所述目标审查分类模型包括多个,所述审查结果的数量与所述目标审查分类模型的数量相同,所述基于所述审查结果判断审查是否通过的步骤包括:
检测多个所述审查结果是否均为审查合格;
若多个所述审查结果均为审查合格,则判定审查通过;
若多个所述审查结果中至少存在一个为审查不合格,则判定审查不通过。
在一实施例中,若所述作品类型为文字作品,所述将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证的步骤包括:
通过预设文档搜索引擎计算所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值;
根据所述第一相似度值从所述已认证文字作品中筛选得到第一预设数量的相似文字作品;
计算所述相似文字作品与所述待认证数字作品之间的第一最长公共子序列,并根据所述第一最长公共子序列的长度计算所述相似文字作品与所述待认证数字作品之间的长度比值,得到第一计算结果;
检测所述第一计算结果中是否存在大于第一预设阈值的长度比值;
若存在大于第一预设阈值的长度比值,则判定版权认证不通过;
若不存在大于第一预设阈值的长度比值,则判定版权认证通过。
在一实施例中,所述通过预设文档搜索引擎计算所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值的步骤包括:
通过预设文档搜索引擎对所述待认证数字作品进行分词处理,得到分词集;
通过所述预设文档搜索引擎对预设认证作品库中的已认证文字作品进行倒排索引,并根据所述倒排索引结果计算所述分词集中各分词对应的分值;
对各分词的分值进行加和处理,得到所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值。
在一实施例中,若所述作品类型为图片作品,所述将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证的步骤包括:
通过预设图像检索引擎计算所述待认证数字作品与预设认证作品库中的已认证图片作品之间的第二相似度值;
根据所述第二相似度值从所述已认证图片作品中筛选得到第二预设数量的相似图片作品;
提取所述待认证数字作品的第一尺度不变特征变换SIFT特征向量,并提取所述相似图片作品的第二SIFT特征向量;
计算所述第一SIFT特征向量与所述第二SIFT特征向量之间的余弦距离,得到第二计算结果;
检测所述第二计算结果中是否存在大于第二预设阈值的余弦距离;
若存在大于第二预设阈值的余弦距离,则判定版权认证不通过;
若不存在大于第二预设阈值的余弦距离,则判定版权认证通过。
在一实施例中,若所述作品类型为音频作品,所述将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证的步骤包括:
将所述待认证数字作品转换成文字作品类型,得到待认证音频文字作品;
通过预设文档搜索引擎计算所述待认证音频文字作品与预设认证作品库中的已认证音频文字作品之间的第三相似度值;
根据所述第三相似度值从所述已认证音频文字作品中检索得到第三预设数量的相似音频文字作品;
计算所述相似音频文件作品与所述待认证音频文字作品之间的第二最长公共子序列,并根据所述第二最长公共子序列的长度计算所述相似音频文字作品与所述待认证音频文字作品之间的长度比值,得到第三计算结果;
检测所述第三计算结果中是否存在大于第三预设阈值的长度比值;
若存在大于第三预设阈值的长度比值,则判定版权认证不通过;
若不存在大于第三预设阈值的长度比值,则判定版权认证通过。
在一实施例中,所述版权认证方法还包括:
在版权认证通过时,获取所述待认证数字作品的作品信息;
基于所述作品信息生成对应的版权认证信息,并根据所述版权认证信息生成数据上链请求;
将所述数据上链请求发送至版权认证联盟链,以使得所述版权认证联盟链基于所述数据上链请求完成对所述待认证数字作品的上链操作。
此外,为实现上述目的,本申请还提供一种版权认证装置,所述版权认证装置包括:
第一获取模块,用于在接收到数字作品版权认证请求时,根据所述数字作品版权认证请求获取待认证数字作品和作品类型;
处理模块,用于根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象;
审查模块,用于将所述目标输入对象输入至所述目标审查分类模型中,得到审查结果,并基于所述审查结果判断审查是否通过;
版权认证模块,用于在审查通过时,将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证。
此外,为实现上述目的,本申请还提供一种版权认证设备,所述版权认证设备包括:存储器、处理器及存储在所述存储器上并可在所述处理器上运行的版权认证程序,所述版权认证程序被所述处理器执行时实现如上所述的版权认证方法的步骤。
此外,为实现上述目的,本申请还提供一种版权认证系统,所述版权认证系统包括版 权认证设备和版权认证联盟链;其中,
所述版权认证设备为如上所述的版权认证设备;
所述版权认证联盟链,用于接收所述版权认证设备发送的数据上链请求;
基于所述数据上链请求获取待上链数据信息,并基于共识算法完成对所述待上链数据信息的上链操作。
此外,为实现上述目的,本申请还提供一种计算机可读存储介质,所述计算机可读存储介质上存储有版权认证程序,所述版权认证程序被处理器执行时实现如上所述的版权认证方法的步骤。
本申请提供一种版权认证方法、装置、设备、系统及计算机可读存储介质,在接收到数字作品版权认证请求时,根据数字作品版权认证请求获取待认证数字作品和作品类型;根据作品类型确定对应的处理策略和目标审查分类模型,基于处理策略对待认证数字作品进行处理,得到目标输入对象;将目标输入对象输入至目标审查分类模型中,得到审查结果,并基于审查结果判断审查是否通过;并在在审查通过时,将待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证。通过上述方式,可实现对数字作品版权的在线认证,可以针对不同类型的数字作品在秒级的时间复杂度内完成版权认证,相比于现有技术中通过人工进行版权认证,本申请可降低人力成本、缩短版权认证周期、提高版权认证效率,从而可及时对作品著作人的版权进行保护。
附图说明
图1为本申请实施例方案涉及的硬件运行环境的设备结构示意图;
图2为本申请版权认证方法第一实施例的流程示意图;
图3为本申请版权认证系统的系统结构示意图;
图4为本申请版权认证装置第一实施例的功能模块示意图。
本申请目的的实现、功能特点及优点将结合实施例,参照附图做进一步说明。
具体实施方式
应当理解,此处所描述的具体实施例仅仅用以解释本申请,并不用于限定本申请。
参照图1,图1为本申请实施例方案涉及的硬件运行环境的设备结构示意图。
本申请实施例版权认证设备可以是智能手机,也可以是PC(Personal Computer,个人计算机)、平板电脑、便携计算机等终端设备。
如图1所示,该版权认证设备可以包括:处理器1001,例如CPU,通信总线1002,用户接口1003,网络接口1004,存储器1005。其中,通信总线1002用于实现这些组件之间的连接通信。用户接口1003可以包括显示屏(Display)、输入单元比如键盘(Keyboard),可选用户接口1003还可以包括标准的有线接口、无线接口。网络接口1004可选的可以包括标准的有线接口、无线接口(如Wi-Fi接口)。存储器1005可以是高速RAM存储器,也可以是稳定的存储器(non-volatile memory),例如磁盘存储器。存储器1005可选的还可以是独立于前述处理器1001的存储装置。
本领域技术人员可以理解,图1中示出的版权认证设备结构并不构成对版权认证设备的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置。
如图1所示,作为一种计算机存储介质的存储器1005中可以包括操作系统、网络通信模块、用户接口模块以及版权认证程序。
在图1所示的终端中,网络接口1004主要用于连接后台服务器,与后台服务器进行 数据通信;用户接口1003主要用于连接客户端,与客户端进行数据通信;而处理器1001可以用于调用存储器1005中存储的版权认证程序,并执行以下版权认证方法的各个步骤。
基于上述硬件结构,提出本申请版权认证方法的各实施例。
本申请提供一种版权认证方法。
参照图2,图2为本申请版权认证方法第一实施例的流程示意图。
在本实施例中,该版权认证方法包括:
步骤S10,在接收到数字作品版权认证请求时,根据所述数字作品版权认证请求获取待认证数字作品和作品类型;
在本实施例中,该版权认证方法应用于版权认证系统中,该版权认证系统包括版权认证设备和版权认证联盟链,其中,本实施例的版权认证方法是由版权认证设备实现的,该设备搭载有版权认证系统。版权认证联盟链可以由版权认证机构节点、公证机构节点、司法机构节点以及外联节点组成,用于接收版权认证设备发送的数据上链请求,进而基于数据上链请求获取待上链数据信息,然后基于共识算法完成对该待上链数据信息的上链操作,即实现基于区块链的版权认证。
在本实施例中,当用户需要对其作品进行版权认证时,可通过用户端(如PC个人计算机、智能手机等)的对应软件上传其数字作品(如文字作品、图片作品和音频作品等),并填写相关信息(包括但不限于作品类型、著作人信息等),进而触发数字作品版权认证请求,此时,版权认证系统在接收到数字作品版权认证请求时,根据数字作品版权认证请求获取待认证数字作品和作品类型。当然,可以理解,在具体实施时,用户在触发数字作品版权认证请求时,可以只上传其数字作品,版权认证系统可在获取到待认证数字作品后,可根据其格式判断其对应的作品类型。
步骤S20,根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象;
在获取到待认证数字作品和作品类型后,根据作品类型确定对应的处理策略和目标审查分类模型,基于处理策略对待认证数字作品进行处理,得到目标输入对象。
具体的,若作品类型为文字作品,则确定对应的处理策略为第一处理策略,并确定目标审查分类模型为第一审查分类模型;然后,基于第一处理策略对待认证数字作品进行切词处理,得到第一切词文本;将第一切词文本输入预设词向量模型,得到第一切词文本中各切词的第一词向量;进而根据第一词向量得到待认证数字作品对应的第一文档向量,其中,目标输入对象为第一文档向量。
若作品类型为图片作品,则确定对应的处理策略为第二处理策略,并确定目标审查分类模型为第二审查分类模型;基于第二处理策略对待认证数字作品进行预处理,得到输入图片,其中,预处理包括缩放处理和灰度处理,目标输入对象为输入图片。
若作品类型为音频作品,则确定对应的处理策略为第三处理策略,并确定目标审查分类模型为第三审查分类模型;基于第三处理策略将待认证数字作品转换成文字作品类型,得到转换后的待认证数字作品;然后按与文字作品的处理方法进行处理,即,对转换后的待认证数字作品进行切词处理,得到第二切词文本;进而将第二切词文本输入预设词向量模型,得到第二切词文本中各切词的第二词向量;根据第二词向量得到转换后的待认证数字作品对应的第二文档向量,其中,目标输入对象为第二文档向量。
具体的执行过程可参照下述第二实施例,此处不作赘述。
步骤S30,将所述目标输入对象输入至所述目标审查分类模型中,得到审查结果,并基于所述审查结果判断审查是否通过;
在得到目标输入对象之后,将目标输入对象输入至目标审查分类模型中,得到审查结果,并基于审查结果判断审查是否通过。其中,在审查时,主要是为了审查用户提交的作 品是否涉及恐怖宣传、邪教宣传、政治敏感以及黄、赌、毒,对应的目标审查分类模型可以包括6类。即,目标审查分类模型可以包括恐怖宣传类审查分类模型、邪教宣传类审查分类模型、政治敏感类审查分类模型、涉黄类审查分类模型、涉赌类审查分类模型、涉毒类审查分类模型。当目标审查分类模型可以包括多个时,审查结果的数量与目标审查分类模型的数量相同,即对应的审查结果也包括多个。在基于多个审查结果判断审查是否通过时,需检测多个审查结果是否均为审查合格;若多个所述审查结果均为审查合格,则说明待认证数字作品不涉及不良信息(恐怖宣传/邪教宣传/政治敏感/黄/赌/毒),此时判定审查通过;若多个所述审查结果中至少存在一个为审查不合格,则说明待认证数字作品涉及不良信息(恐怖宣传/邪教宣传/政治敏感/黄/赌/毒)中一种或多种,此时判定审查不通过。
当然,在具体实施例中,还可以针对各类型数字作品分别构建一个审查分类,对应的,审查结果只有一个,此时,则只需检测该审查结果是否为审查合格。
步骤S40,在审查通过时,将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证。
在审查通过时,将待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证。具体的,需针对不同类型的待认证数字作品采用不同的版权认证方法,具体的版权认证过程可参照下述第四实施例,此处不作赘述。通过将待认证数字作品与预设认证作品库中的已认证数字作品进行比对,可检测待认证作品与已认证作品之间是否存处在引用/抄袭关系,以确定是否对版本进行认证。
本申请实施例提供一种版权认证方法,在接收到数字作品版权认证请求时,根据数字作品版权认证请求获取待认证数字作品和作品类型;根据作品类型确定对应的处理策略和目标审查分类模型,基于处理策略对待认证数字作品进行处理,得到目标输入对象;将目标输入对象输入至目标审查分类模型中,得到审查结果,并基于审查结果判断审查是否通过;并在在审查通过时,将待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证。通过上述方式,可实现对数字作品版权的在线认证,可以针对不同类型的数字作品在秒级的时间复杂度内完成版权认证,相比于现有技术中通过人工进行版权认证,本申请实施例可降低人力成本、缩短版权认证周期、提高版权认证效率,从而可及时对作品著作人的版权进行保护。
进一步的,基于图2所示的第一实施例,提出本申请版权认证方法的第二实施例。
在本实施例中,作为其中一实施方式,步骤S20可以包括:
步骤a11,若所述作品类型为文字作品,则确定对应的处理策略为第一处理策略,并确定目标审查分类模型为第一审查分类模型;
步骤a12,基于所述第一处理策略对所述待认证数字作品进行切词处理,得到第一切词文本;
步骤a13,将所述第一切词文本输入预设词向量模型,得到所述第一切词文本中各切词的第一词向量;
步骤a14,根据所述第一词向量得到所述待认证数字作品对应的第一文档向量,其中,所述目标输入对象为所述第一文档向量。
在本实施例中,对于文字作品的处理过程如下:
若作品类型为文字作品,则确定对应的处理策略为第一处理策略,并确定目标审查分类模型为第一审查分类模型。其中,第一审查分类模型是预先训练好的,其类型可以为SVM(Support Vector Machine,支持向量机)模型、贝叶斯模型、逻辑回归模型、卷积神经网络模型等二分类模型,下述训练过程以SVM模型进行说明。对于目标审查分类模型,以目标审查分类模型包括恐怖宣传类审查分类模型、邪教宣传类审查分类模型、政治敏感类审查分类模型、涉黄类审查分类模型、涉赌类审查分类模型、涉毒类审查分类模型这6类 为例进行说明,对应的,第一审查分类模型也包括6类,其训练过程为:分别取已标注的5万条涉及和5万条不涉及恐怖宣传的文字作品,对每个文字作品进行切词后,通过预设词向量模型(在一实施例中为word2vec模型)得到词向量后,对词向量按对应维度进行相加,得到各文字作品对应的文档向量,然后根据上述10万个文档向量训练一个SVM分类模型。接下来利用同样的方法,分别训练用于判断作品是否涉及邪教宣传、是否涉及政治敏感、是否涉黄、是否涉赌、是否涉毒的5个SVM分类模型。
在确定得到第一处理策略后,基于第一处理策略对待认证数字作品进行切词处理,得到第一切词文本,其中,切词处理可采用预设的工具,如中科院NLPIR、哈工大LTP、结巴分词等。具体的分词过程与现有技术相一致,此处不作赘述。
然后,将第一切词文本输入预设词向量模型,得到第一切词文本中各切词的第一词向量。其中,预设词向量模型在一实施例中为word2vec(word to vector,用来产生词向量的相关模型),word2vec将每一个中文词汇映射为一个高维向量(通常取200维向量),且对于任意两个中文词汇,语义上越相近,映射后得到的向量距离也越近。因此可以根据词向量之间的距离来描述中文词汇的语义相似性。
最后,根据第一词向量得到待认证数字作品对应的第一文档向量,其中,目标输入对象为第一文档向量,即后续将第一文档向量输入至对应的第一审查分类模型中,以得到审查结果。对于第一文档向量的获取,是对第一词向量按对应维度进行相加,即可得到对应的第一文档向量。
通过上述方式,可实现对文字作品类的数字作品进行处理,得到对应的目标输入对象,以便于后续输入至目标审查分类模型中,得到审查结果。
作为又一实施方式,步骤S20还可以包括:
步骤a21,若所述作品类型为图片作品,则确定对应的处理策略为第二处理策略,并确定目标审查分类模型为第二审查分类模型;
步骤a22,基于所述第二处理策略对所述待认证数字作品进行预处理,得到输入图片,其中,所述预处理包括缩放处理和灰度处理,所述目标输入对象为所述输入图片。
在本实施例中,对于图片作品的处理过程如下:
若作品类型为图片作品,则确定对应的处理策略为第二处理策略,并确定目标审查分类模型为第二审查分类模型。其中,第二审查分类模型是预先训练好的,其类型在一实施例中为基于卷积神经网络的分类模型。对于目标审查分类模型,以目标审查分类模型包括恐怖宣传类审查分类模型、邪教宣传类审查分类模型、政治敏感类审查分类模型、涉黄类审查分类模型、涉赌类审查分类模型、涉毒类审查分类模型这6类为例进行说明,对应的,第二审查分类模型也包括6类,其训练过程为:分别取已标注的5万条涉及和5万条不涉及恐怖宣传的图片作品,对每个图片作品进行预处理,预处理过程包括缩放处理和灰度处理,其中,缩放处理即为将图片的大小缩放为一预设尺寸,例如128像素*128像素,灰度处理即将缩放后的图片转换为灰度图片,然后根据上述10万张预处理后的图片训练一个基于卷积神经网络的分类模型。接下来利用同样的方法,分别训练用于判断作品是否涉及邪教宣传、是否涉及政治敏感、是否涉黄、是否涉赌、是否涉毒的5个基于卷积神经网络的分类模型。
在确定得到第二处理策略后,基于第二处理策略对待认证数字作品进行预处理,得到输入图片,其中,预处理包括缩放处理和灰度处理,缩放处理即为将图片的大小缩放为一预设尺寸,例如128像素*128像素,灰度处理即将缩放后的图片转换为灰度图片,目标输入对象为输入图片,即后续将输入图片输入至对应的第二审查分类模型中,以得到审查结果。
通过上述方式,可实现对图片作品类的数字作品进行处理,得到对应的目标输入对象, 以便于后续输入至目标审查分类模型中,得到审查结果。
作为另一实施方式,步骤S20还可以包括:
步骤a31,若所述作品类型为音频作品,则确定对应的处理策略为第三处理策略,并确定目标审查分类模型为第三审查分类模型;
步骤a32,基于所述第三处理策略将所述待认证数字作品转换成文字作品类型,得到转换后的待认证数字作品;
步骤a33,对所述转换后的待认证数字作品进行切词处理,得到第二切词文本;
步骤a34,将所述第二切词文本输入预设词向量模型,得到所述第二切词文本中各切词的第二词向量;
步骤a35,根据所述第二词向量得到所述转换后的待认证数字作品对应的第二文档向量,其中,所述目标输入对象为所述第二文档向量。
在本实施例中,对于音频作品的处理过程如下:
若作品类型为音频作品,则确定对应的处理策略为第三处理策略,并确定目标审查分类模型为第三审查分类模型。其中,第三审查分类模型是预先训练好的,其类型可以为SVM(Support Vector Machine,支持向量机)模型、贝叶斯模型、逻辑回归模型、卷积神经网络模型等二分类模型,其中,第三审查分类模型可以与第一审查分类模型相同,也可以基于第一审查分类模型的训练方法训练得到的其他类型的二分类模型。
在确定得到第三处理策略后,先基于第三处理策略将待认证数字作品转换成文字作品类型,得到转换后的待认证数字作品。具体的,可以通过语音识别工具将音频作品转换为文字作品类型。然后,对转换后的待认证数字作品进行切词处理,得到第二切词文本,其中,切词处理可采用预设的工具,如中科院NLPIR、哈工大LTP、结巴分词等。具体的分词过程与现有技术相一致,此处不作赘述。
然后,将第二切词文本输入预设词向量模型,得到第二切词文本中各切词的第二词向量。其中,预设词向量模型在一实施例中为word2vec(word to vector,用来产生词向量的相关模型)。
最后,根据第二词向量得到待认证数字作品对应的第二文档向量,其中,目标输入对象为第二文档向量,即后续将第二文档向量输入值对应的第三审查分类模型中,以得到审查结果。对于第二文档向量的获取,是对第二词向量按对应维度进行相加,即可得到对应的第二文档向量。
通过上述方式,可实现对音频作品类的数字作品进行处理,得到对应的目标输入对象,以便于后续输入至目标审查分类模型中,得到审查结果。
基于上述第一实施例和第二实施例,提出本申请版权认证方法的第三实施例。
在本实施例中,所述目标审查分类模型包括多个,所述审查结果的数量与所述目标审查分类模型的数量相同,步骤“基于所述审查结果判断审查是否通过”包括:
步骤b1,检测多个所述审查结果是否均为审查合格;
步骤b2,若多个所述审查结果均为审查合格,则判定审查通过;
步骤b3,若多个所述审查结果中至少存在一个为审查不合格,则判定审查不通过。
在本实施例中,目标审查分类模型可以包括多个,审查结果的数量与目标审查分类模型的数量相同,对应的审查结果也包括多个。对于审查是否通过的判断过程为:检测多个审查结果是否均为审查合格。
若多个所述审查结果均为审查合格,则说明待认证数字作品不涉及不良信息(恐怖宣传/邪教宣传/政治敏感/黄/赌/毒),此时判定审查通过;
若多个所述审查结果中至少存在一个为审查不合格,则说明待认证数字作品涉及不良信息(恐怖宣传/邪教宣传/政治敏感/黄/赌/毒)中一种或多种,此时判定审查不通过。
进一步的,基于上述第一实施例,提出本申请版权认证方法的第四实施例。
在本实施例中,若所述作品类型为文字作品,步骤S40包括:
步骤c11,通过预设文档搜索引擎计算所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值;
步骤c12,根据所述第一相似度值从所述已认证文字作品中筛选得到第一预设数量的相似文字作品;
步骤c13,计算所述相似文字作品与所述待认证数字作品之间的第一最长公共子序列,并根据所述第一最长公共子序列的长度计算所述相似文字作品与所述待认证数字作品之间的长度比值,得到第一计算结果;
步骤c14,检测所述第一计算结果中是否存在大于第一预设阈值的长度比值;
步骤c15,若存在大于第一预设阈值的长度比值,则判定版权认证不通过;
步骤c16,若不存在大于第一预设阈值的长度比值,则判定版权认证通过。
其中,步骤c11包括:
步骤c111,通过预设文档搜索引擎对所述待认证数字作品进行分词处理,得到分词集;
步骤c112,通过所述预设文档搜索引擎对预设认证作品库中的已认证文字作品进行倒排索引,并根据所述倒排索引结果计算所述分词集中各分词对应的分值;
步骤c113,对各分词的分值进行加和处理,得到所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值。
在本实施例中,若作品类型为文字作品,其对应的版权认证过程如下:
通过预设文档搜索引擎计算待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值。其中,预设文档搜索引擎在一实施例中为ES(Elastic Search,弹性检索)搜索引擎,ES是一个分布式、高扩展、高实时的搜索与数据分析引擎,它能很方便的使大量数据具有搜索、分析和探索的能力。具体的,先通过ES搜索引擎对待认证数字作品进行分词处理,得到分词集,其中ES搜索引擎中自带分词器,可以对待认证数字作品进行分词。然后,通过ES搜索引擎对预设认证作品库中的已认证文字作品进行倒排索引,并根据倒排索引结果计算分词集中各分词对应的分值。在通过ES搜索引擎对已认证文字作品进行倒排索引时,也是先通过ES搜索引擎自带的分词器进行分词,并会得到各分词的词频信息和位置信息,进而建立各分词与作品文档之间的倒排索引,其中倒排索引是一个字典类数据结构(key-value),字典的键(key)是一个个的分词,值(value)是包含该分词的作品列表,以及该分词在每个作品中的位置信息与词频信息,通过倒排索引,可以根据分词快速获取包含这个分词的文档列表及词频信息。在根据倒排索引结果计算分词集中各分词对应的分值时,可以根据倒排索引结果获取包含分词的文档、及分词在文档中的词频和逆文档频率,进而根据词频和逆文档频率计算各分词的分值。例如,对待认证数字作品进行分词后得到包含分词“A”、“B”、“C”、“D”、“E”和“F”的分词集,先基于已认证文字作品的倒排索引找到包含分词“A”的作品集,并计算分词“A”对应的分值,具体的,可以先计算作品集中每个作品的分值,每个作品的分值可以为作品中分词A的词频与词A的逆文档频率的乘积(当然也可以根据实际情况设定其他计算方式,如基于词频、逆文档频率和位置信息计算),进而对每个作品的分值进行加和得到分词“A”对应的分值;在得到分词“A”对应的分值之后,对分词“B”、“C”、“D”、“E”、“F”执行相同的操作,以得到分词“B”、“C”、“D”、“E”、“F”分别对应的分值。此外,可以理解,预设认证作品库中的已认证文字作品,可预先采用ES搜索引擎按特定的索引结构进行存储,此时则无需执行“通过所述预设文档搜索引擎对预设认证作品库中的已认证文字作品进行倒排索引”这一步骤。在得到的各分词对应的分值之后,对各分词的分值进行加和处理,得到待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值。
在得到待认证数字作品与已认证文字作品之间的第一相似度值之后,根据第一相似度值从已认证文字作品中筛选得到第一预设数量的相似文字作品,其中,第一预设数量可以根据实际需要进行设定,例如,可设定为一预设数值,如1000,还可以设定为第一相似度值大于一预设值的数量,此处不作具体限定。在筛选时,将第一相似度值按从大到小的顺序进行排序,取排名在前的第一预设数量的文字作品,作为相似文字作品。需要说明的是,考虑到两篇文字作品存在抄袭关系的前提是两篇作品的用词是一致的,且每个词语前后相邻的词语也是一致的。如果两篇作品的用词都不相同,那么这两篇新闻也一定不会存在抄袭关系。因此,此处通过ES搜索引擎初步判断用词是否一致,以进行初步的筛选,其目的在于从已认证文字作品中筛选得到包含待认证文字作品中的全部或大部分切词的文字作品,以缩小认证比对的范围,节省服务器资源,进一步提高文字作品版权认证的效率。此外,需要说明的是,在具体实施时,当然也可以直接基于计算最长公告子序列的方式来计算待认证数字作品与预设认证作品库中的已认证文字作品之间的相似度值,进而基于该相似度值判定版权认证是否通过,但是相比而言,通过ES搜索引擎进行初筛的计算过程的复杂程度显然低于最长公共子序列的计算复杂程度,借助ES搜索引擎从千万级的已认证文字作品中检索出与待认证数字作品用词一致的相似度值为前1000作品集的耗时为毫秒级,因此,本实施例中通过ES搜索引擎先进行初筛,再与后续计算最长公共子序列相结合,相比于直接基于最长公共子序列计算相似度值的方式,可进一步提高作品版权认证效率。
然后,计算相似文字作品与待认证数字作品之间的第一最长公共子序列,其中,第一最长公共子序列中的“第一”并无实质含义,仅用于与后续的第二最长公告子序列进行区分。最长公共子序列,即找出两个序列的最长的公共子序列,例如,给定两个字符串X=<x1,x2,x3,…,xn>,Y=<y1,y2,y3,…,ym>,存在两个长度为k的下标序列<i1,i2,…,ik>,<j1,j2,…,jk>,使得字符串X和Y在对应下标i和j的位置的字符相等,满足上述要求的最长的下标序列对应的子字符串即为X和Y之间的最长公共子序列。如字符串ABCDE和字符串XAYCDZ之间的最长公共子序列为ACD。然后,根据第一最长公共子序列的长度计算相似文字作品与待认证数字作品之间的长度比值,得到第一计算结果。第一最长公共子序列的数量对应的与相似文字作品的数量相同。对于长度比值的计算,例如,某一相似文字作品的第一最长公共子序列包括100个字符时,则其长度为100,若待认证数字作品包括1000个字符,则长度比值为100/1000=0.1。
然后,检测第一计算结果中是否存在大于第一预设阈值的长度比值,其中,第一预设阈值可根据实际需要设定,例如可设为0.8,此处不作具体限定。若存在大于第一预设阈值的长度比值,则认为存在引用/抄袭关系,此时判定版权认证不通过;若不存在大于第一预设阈值的长度比值,则判定版权认证通过。
需要说明的是,在具体实施时,“计算相似文字作品与待认证数字作品之间的第一最长公共子序列,并根据第一最长公共子序列的长度计算相似文字作品与待认证数字作品之间的长度比值”与“检测是否存在大于第一预设阈值的长度比值”的过程可同时进行,即,依次计算各相似文字作品与待认证数字作品之间的第一最长公共子序列,然后根据第一最长公共子序列的长度计算相似文字作品与待认证数字作品之间的长度比值,进而该长度比值是否大于第一预设阈值,一旦检测到某一相似文字作品对应的长度比值大于第一预设阈值,则认为存在引用/抄袭关系,即可判定版权认证不通过,此时则无需计算其他相似文字作品与待认证数字作品之间的第一最长公共子序列及后续步骤,可节省服务器资源,进一步提高版权认证的效率。
通过上述方式,可检测待认证作品与已认证作品之间是否存处在引用/抄袭关系,以确定是否对版本进行认证。
在本实施例中,若所述作品类型为图片作品,步骤S40还可以包括:
步骤c21,通过预设图像检索引擎计算所述待认证数字作品与预设认证作品库中的已认证图片作品之间的第二相似度值;
步骤c22,根据所述第二相似度值从所述已认证图片作品中筛选得到第二预设数量的相似图片作品;
步骤c23,提取所述待认证数字作品的第一尺度不变特征变换SIFT特征向量,并提取所述相似图片作品的第二SIFT特征向量;
步骤c24,计算所述第一SIFT特征向量与所述第二SIFT特征向量之间的余弦距离,得到第二计算结果;
步骤c25,检测所述第二计算结果中是否存在大于第二预设阈值的余弦距离;
步骤c26,若存在大于第二预设阈值的余弦距离,则判定版权认证不通过;
步骤c27,若不存在大于第二预设阈值的余弦距离,则判定版权认证通过。
在本实施例中,若所述作品类型为图片作品,其对应的版权认证过程如下:
通过预设图像检索引擎计算待认证数字作品与预设认证作品库中的已认证图片作品之间的第二相似度值。其中,预设图像检索引擎在一实施例中为CBIR(Content-based image retrieval,基于内容的图像检索)引擎,CBIR引擎的核心是使用图像的可视特征对图像进行检索。本质上讲,它是一种近似匹配技术,融合了计算机视觉、图像处理、图像理解和数据库等多个领域的技术成果,其中的特征提取和索引的建立可由计算机自动完成,避免了人工描述的主观性。用户检索的过程一般是提供一个样例图像(Queryby Example)或描绘一幅草图(Queryby Sketch),系统抽取该查询图像的特征,然后与数据库中的特征进行比较,并将与查询特征相似的图像返回给用户。需要说明的是,由于预设认证作品库中的已认证图片作品都是按预设尺寸存储的,故需对图片类的待认证数字作品进行对应的缩放处理,得到同样预设尺寸的待认证数字作品,再通过预设图像检索引擎计算缩放后的待认证数字作品与预设认证作品库中的已认证图片作品之间的第二相似度值。
在得到待认证数字作品与已认证图片作品之间的第二相似度值之后,根据第二相似度值从已认证图片作品中筛选得到第二预设数量的相似图片作品,其中,第二预设数量可以与第一预设数量相同或不同,可以根据实际需要进行设定,此处不作具体限定。在筛选时,将第二相似度值按从大到小的顺序进行排序,取排名在前的第二预设数量的文字作品,作为相似文字作品。需要说明的是,考虑到两个图片作品存在抄袭关系的前提是两个图片作品的颜色、形状、纹理等低层次特征是一致的,且各特征的排列方式也是一致的。如果两个图片作品的低层次特征都不相似,那么这两个图片作品也一定不会存在抄袭关系。因此,此处通过CBIR引擎进行初步筛选的目的在于从已认证图片作品中筛选得到包含较多待认证图片作品中特征(如颜色特征、形状特征、纹理特征等)的图片作品,以缩小认证比对的范围,节省服务器资源,进一步提高图片作品版权认证的效率。当然,需要说明的是,在具体实施时,也可以直接基于SIFT特征向量的方式来计算待认证数字作品与预设认证作品库中的已认证图片作品之间的相似度值,进而基于该相似度值判定版权认证是否通过,但是相比而言,通过CBIR引擎进行特征提取和相似度计算,相比于提取SIFT特征向量,其计算过程更为简单,效率更高,因此,本实施例中先通过CBIR引擎先进行初筛,再与后续基于SIFT特征向量计算相似度值相结合,相比于直接基于SIFT特征向量计算相似度值的方式,可进一步提高作品版权认证效率。
然后,分别提取待认证数字作品和相似图片作品的SIFT(Scale-invariant feature transform,尺度不变特征变换)特征向量,即提取待认证数字作品的第一SIFT特征向量,并提取相似图片作品的第二SIFT特征向量。其中,SIFT特征向量的提取过程为:在图像中检测出一系列的关键点,这些关键点对尺度缩放,旋转以及亮度变化无关,然后对上述关键点分配梯度方向值,就可以得到一幅图像的SIFT特征向量。具体的提取过程与现有 技术相一致,此处不作具体描述。
进而,计算第一SIFT特征向量与第二SIFT特征向量之间的余弦距离,得到第二计算结果。本实施例中采用余弦距离表征两特征向量之间的相似度,在具体实施例中,还可以采用其他参数来表征,如欧式距离等。最后,检测第二计算结果中是否存在大于第二预设阈值的余弦距离;其中,第二预设阈值可以与第一预设阈值相同或不同,可根据实际需要设定,例如也可设为0.8,此处不作具体限定。若存在大于第二预设阈值的余弦距离,则认为存在引用/抄袭关系,判定版权认证不通过;若不存在大于第二预设阈值的余弦距离,则判定版权认证通过。
同样的,在具体实施时,“提取待认证数字作品的第一SIFT特征向量,并提取相似图片作品的第二SIFT特征向量,然后计算第一SIFT特征向量与第二SIFT特征向量之间的余弦距离”与“检测是否存在大于第二预设阈值的余弦距离”的过程可同时进行,一旦检测到某一相似图片作品对应的余弦距离大于第二预设阈值,则认为存在引用/抄袭关系,即可判定版权认证不通过,此时则无需提取其他相似图片作品与待认证数字作品的SIFT特征向量,也无需计算对应的余弦距离及后续检测步骤,可节省服务器资源,进一步提高版权认证的效率。
在本实施例中,若所述作品类型为音频作品,步骤S40还可以包括:
步骤c31,将所述待认证数字作品转换成文字作品类型,得到待认证音频文字作品;
步骤c32,通过预设文档搜索引擎计算所述待认证音频文字作品与预设认证作品库中的已认证音频文字作品之间的第三相似度值;
步骤c33,根据所述第三相似度值从所述已认证音频文字作品中检索得到第三预设数量的相似音频文字作品;
步骤c34,计算所述相似音频文件作品与所述待认证音频文字作品之间的第二最长公共子序列,并根据所述第二最长公共子序列的长度计算所述相似音频文字作品与所述待认证音频文字作品之间的长度比值,得到第三计算结果;
步骤c35,检测所述第三计算结果中是否存在大于第三预设阈值的长度比值;
步骤c36,若存在大于第三预设阈值的长度比值,则判定版权认证不通过;
步骤c37,若不存在大于第三预设阈值的长度比值,则判定版权认证通过。
在本实施例中,若作品类型为音频作品,其对应的版权认证过程如下:
将所述待认证数字作品转换成文字作品类型,得到待认证音频文字作品,然后,通过预设文档搜索引擎计算待认证数字作品与预设认证作品库中的已认证音频文字作品之间的第三相似度值。其中,预设文档搜索引擎在一实施例中为ES搜索引擎。预设认证作品库中的已认证音频文字作品,是基于语音识别工具将音频作品转换为文字作品类型得到的,可预先采用ES搜索引擎按特定的索引结构进行存储,以便于后续进行搜索。
在得到待认证数字作品与已认证音频文字作品之间的第三相似度值之后,根据第三相似度值从已认证音频文字作品中筛选得到第三预设数量的相似文字作品,其中,第三预设数量可以与第一预设数量和第二预设数量相同或不同,可根据实际需要进行设定,此处不作具体限定。在筛选时,将第三相似度值按从大到小的顺序进行排序,取排名在前的第三预设数量的文字作品,作为相似文字作品。需要说明的是,此处筛选的目的在于从已认证音频文字作品中筛选得到包含待认证音频文字作品中的全部或大部分切词的文字作品,以缩小认证比对的范围,节省服务器资源,进一步提高版权认证的效率。
然后,计算相似文字音频作品与待认证音频文字作品之间的第二最长公共子序列,其中,第二最长公共子序列中的“第二”并无实质含义,仅用于与上述第一最长公告子序列进行区分。进而根据第二最长公共子序列的长度计算相似文字音频作品与待认证数字作品之间的长度比值,得到第三计算结果,并检测第三计算结果中是否存在大于第三预设阈值 的长度比值,其中,第三预设阈值可以与第一预设阈值和第二预设阈值相同或不同,可根据实际需要设定,例如可设为0.8,此处不作具体限定。若存在大于第三预设阈值的长度比值,则认为存在引用/抄袭关系,此时判定版权认证不通过;若不存在大于第三预设阈值的长度比值,则判定版权认证通过。
需要说明的是,在具体实施时,“计算相似音频文字作品与待认证数字作品之间的第二最长公共子序列,并根据第二最长公共子序列的长度计算相似音频文字作品与待认证数字作品之间的长度比值”与“检测是否存在大于第三预设阈值的长度比值”的过程可同时进行,即,依次计算各相似音频文字作品与待认证数字作品之间的第二最长公共子序列,然后根据第二最长公共子序列的长度计算相似音频文字作品与待认证数字作品之间的长度比值,进而该长度比值是否大于第三预设阈值,一旦检测到某一相似音频文字作品对应的长度比值大于第三预设阈值,则认为存在引用/抄袭关系,即可判定版权认证不通过,此时则无需计算其他相似音频文字作品与待认证数字作品之间的第二最长公共子序列及后续步骤,可节省服务器资源,进一步提高版权认证的效率。
进一步地,基于上述第一、第二和第四实施例,提出本申请版权认证方法的第五实施例。
在本实施例中,在步骤S40之后,该版权认证方法还可以包括:
步骤A,在版权认证通过时,获取所述待认证数字作品的作品信息;
在本实施例中,在版权认证通过时,获取待认证数字作品的作品信息,其中,作品信息中的其中一部分信息(如作品作者、作品类型等)可以根据数字作品版权认证请求获取得到,也可以生成对应的提示窗口,在用户基于提示窗口填写对应的作品信息后获取得到;另一部分信息可由版权认证系统生成,例如认证时间、作品的md5(Message-Digest Algorithm,信息摘要算法)值,其中认证时间可直接获取版权认证通过时的时间,作品的md5值可在获取到待认证数字作品后,通过对应的程序获取,例如,获取文件的byte(字节)信息,第二步通过MessageDigest(信息摘要)类进行md5加密,第三步转换成16进制的md5码值。作品信息可以包括但不限于:作品作者、作品类型、认证时间及待认证数字作品的md5值。
步骤B,基于所述作品信息生成对应的版权认证信息,并根据所述版权认证信息生成数据上链请求;
然后,基于作品信息生成对应的版权认证信息,具体的,可以将作品信息转换成预设数据格式,如json(一种轻量级的数据交换格式)数据结构,以得到版权认证信息。在生成版权认证信息之后,进而根据版权认证信息生成数据上链请求,具体的,可以采用sha256(Secure Hash Algorithm 256,安全散列算法256)算法生成该版权认证信息的哈希值,进而基于该哈希值生成数据上链请求。
步骤C,将所述数据上链请求发送至版权认证联盟链,以使得所述版权认证联盟链基于所述数据上链请求完成对所述待认证数字作品的上链操作。
最后,将数据上链请求发送至版权认证联盟链,以使得版权认证联盟链基于数据上链请求完成对待认证数字作品的上链操作。其中,版权认证联盟链主要由版权认证机构节点、公证机构节点、司法机构节点以及外联节点组成,可以将数据上链请求发送至版权认证联盟链中的外联节点,然后通过所有节点一起基于通用的共识算法完成对待认证数字作品的上链操作,即将数据上链请求中的哈希值写入联盟连中永久保留。
当然,可以理解的是,在版权认证通过时,除上报版权认证联盟链对数据上链外,还可以将该版权认证通过的待认证数字作品存储至预设已认证作品库中,用于检测后续作品是否存在抄袭已认证作品的情况;同时,向用户端返回版权认证成功的消息提示,已告知用户。
在本实施例中,版权认证系统在判定版权认证通过时,可生成对应的数据上链请求,并发送至版权认证联盟链,以使得版权认证联盟链基于数据上链请求完成上链操作,基于区块链具有不可修改的特性,可实现对数字作品版权的保护,后续一旦出现未经授权对他人作品进行传播的情况,版权所有者即可根据区块链上的认证信息对侵权行为进行申诉,降低维权的难度。
在现有技术中,也存在基于区块链的版权认证方案,例如,版权认证机构、公证机构、司法机构以及若干自媒体人分别作为一个节点,共同构成一个版权认证区块链平台,以对数字作品的版权进行审查认证。但是其对作品的审查以及版权认证环节还是在线下通过人工完成,而仅仅把最终的认证结果写入区块链平台,这样的方式并没有缩短认证周期,提高版权认证效率。同时,当前基于区块链的方案,多采用公链方式组织区块链,以此来保证系统运行的稳定性,此外还要求自媒体人或者创作者申请并加入区块链平台,增加创作者的使用成本,且他人可知晓各节点的操作行为,无法有效保证数据隐私性。对此,本申请还提供一种版权认证系统。
参照图3,图3为本申请版权认证系统的系统架构示意图。
在本实施例中,如图3所示,版权认证系统包括版权认证设备和版权认证联盟链;当然,还可以包括用户端。
其中,版权认证设备为如图1所示的版权认证设备;用于执行上述版权认证方法实施例中的各步骤,具体的功能和实现过程可参照上述实施例,此处不作赘述。
所述版权认证联盟链,用于接收所述版权认证设备发送的数据上链请求;
基于所述数据上链请求获取待上链数据信息,并基于共识算法完成对所述待上链数据信息的上链操作。
在本实施例中,版权认证联盟链,可以用于接收版权认证设备发送的数据上链请求,其中,版权认证联盟链主要由版权认证机构节点、公证机构节点、司法机构节点以及外联节点组成,可通过外链节点接收版权认证设备发送的数据上链请求。进而基于数据上链请求获取待上链数据信息,其中,该待上链数据信息可以为基于版权认证信息生成的哈希值,然后基于共识算法完成对该哈希值的上链操作,即将该哈希值写入联盟链中永久保留。
此外,为保证版权认证的权威性和安全性,通常只对来自版权认证设备的请求进行响应,对应的,外联节点在接收到数据上链请求时,可获取对应的设备IP(Internet Protocol,互联网协议地址),进而检测该设备IP是否为预设的版权认证设备的IP,若是,则基于数据上链请求获取待上链数据信息,并基于共识算法完成对待上链数据信息的上链操作。
通过构建上述版权认证系统,可通过版权认证设备实现对数字作品版权的在线认证,可以针对不同类型的数字作品在秒级的时间复杂度内完成版权认证,相比于现有技术中通过人工进行版权认证,本申请可降低人力成本、缩短版权认证周期、提高版权认证效率,从而可及时对作品著作人的版权进行保护。此外,本实施例中,自媒体人或者创作者无需申请并加入区块链平台、即可实现版权认证和保护,从而可减少创作者的使用成本,同时,他人也无法获取创作者的行为操作,可保证数据隐私性。
本申请还提供一种版权认证装置。
参照图4,图4为本申请版权认证装置第一实施例的功能模块示意图。
如图4所示,所述版权认证装置包括:
第一获取模块10,用于在接收到数字作品版权认证请求时,根据所述数字作品版权认证请求获取待认证数字作品和作品类型;
处理模块20,用于根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象;
审查模块30,用于将所述目标输入对象输入至所述目标审查分类模型中,得到审查 结果,并基于所述审查结果判断审查是否通过;
版权认证模块40,用于在审查通过时,将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证。
进一步地,所述处理模块20具体用于:
若所述作品类型为文字作品,则确定对应的处理策略为第一处理策略,并确定目标审查分类模型为第一审查分类模型;
基于所述第一处理策略对所述待认证数字作品进行切词处理,得到第一切词文本;
将所述第一切词文本输入预设词向量模型,得到所述第一切词文本中各切词的第一词向量;
根据所述第一词向量得到所述待认证数字作品对应的第一文档向量,其中,所述目标输入对象为所述第一文档向量。
进一步地,所述处理模块20还具体用于:
若所述作品类型为图片作品,则确定对应的处理策略为第二处理策略,并确定目标审查分类模型为第二审查分类模型;
基于所述第二处理策略对所述待认证数字作品进行预处理,得到输入图片,其中,所述预处理包括缩放处理和灰度处理,所述目标输入对象为所述输入图片。
进一步地,所述处理模块20还具体用于:
若所述作品类型为音频作品,则确定对应的处理策略为第三处理策略,并确定目标审查分类模型为第三审查分类模型;
基于所述第三处理策略将所述待认证数字作品转换成文字作品类型,得到转换后的待认证数字作品;
对所述转换后的待认证数字作品进行切词处理,得到第二切词文本;
将所述第二切词文本输入预设词向量模型,得到所述第二切词文本中各切词的第二词向量;
根据所述第二词向量得到所述转换后的待认证数字作品对应的第二文档向量,其中,所述目标输入对象为所述第二文档向量。
进一步地,所述目标审查分类模型包括多个,所述审查结果的数量与所述目标审查分类模型的数量相同,所述审查模块30具体用于:
检测多个所述审查结果是否均为审查合格;
若多个所述审查结果均为审查合格,则判定审查通过;
若多个所述审查结果中至少存在一个为审查不合格,则判定审查不通过。
进一步地,若所述作品类型为文字作品,所述版权认证模块40具体用于:
通过预设文档搜索引擎计算所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值;
根据所述第一相似度值从所述已认证文字作品中筛选得到第一预设数量的相似文字作品;
计算所述相似文字作品与所述待认证数字作品之间的第一最长公共子序列,并根据所述第一最长公共子序列的长度计算所述相似文字作品与所述待认证数字作品之间的长度比值,得到第一计算结果;
检测所述第一计算结果中是否存在大于第一预设阈值的长度比值;
若存在大于第一预设阈值的长度比值,则判定版权认证不通过;
若不存在大于第一预设阈值的长度比值,则判定版权认证通过。
进一步地,所述版权认证模块40还具体用于:
通过预设文档搜索引擎对所述待认证数字作品进行分词处理,得到分词集;
通过所述预设文档搜索引擎对预设认证作品库中的已认证文字作品进行倒排索引,并 根据所述倒排索引结果计算所述分词集中各分词对应的分值;
对各分词的分值进行加和处理,得到所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值。
进一步地,若所述作品类型为图片作品,所述版权认证模块40还具体用于:
通过预设图像检索引擎计算所述待认证数字作品与预设认证作品库中的已认证图片作品之间的第二相似度值;
根据所述第二相似度值从所述已认证图片作品中筛选得到第二预设数量的相似图片作品;
提取所述待认证数字作品的第一尺度不变特征变换SIFT特征向量,并提取所述相似图片作品的第二SIFT特征向量;
计算所述第一SIFT特征向量与所述第二SIFT特征向量之间的余弦距离,得到第二计算结果;
检测所述第二计算结果中是否存在大于第二预设阈值的余弦距离;
若存在大于第二预设阈值的余弦距离,则判定版权认证不通过;
若不存在大于第二预设阈值的余弦距离,则判定版权认证通过。
进一步地,若所述作品类型为音频作品,所述版权认证模块40还具体用于:
将所述待认证数字作品转换成文字作品类型,得到待认证音频文字作品;
通过预设文档搜索引擎计算所述待认证音频文字作品与预设认证作品库中的已认证音频文字作品之间的第三相似度值;
根据所述第三相似度值从所述已认证音频文字作品中检索得到第三预设数量的相似音频文字作品;
计算所述相似音频文件作品与所述待认证音频文字作品之间的第二最长公共子序列,并根据所述第二最长公共子序列的长度计算所述相似音频文字作品与所述待认证音频文字作品之间的长度比值,得到第三计算结果;
检测所述第三计算结果中是否存在大于第三预设阈值的长度比值;
若存在大于第三预设阈值的长度比值,则判定版权认证不通过;
若不存在大于第三预设阈值的长度比值,则判定版权认证通过。
进一步地,所述版权认证装置还包括:
第二获取模块,用于在版权认证通过时,获取所述待认证数字作品的作品信息;
生成模块,用于基于所述作品信息生成对应的版权认证信息,并根据所述版权认证信息生成数据上链请求;
发送模块,用于将所述数据上链请求发送至版权认证联盟链,以使得所述版权认证联盟链基于所述数据上链请求完成对所述待认证数字作品的上链操作。
其中,上述版权认证装置中各个模块的功能实现与上述版权认证方法实施例中各步骤相对应,其功能和实现过程在此处不再一一赘述。
本申请还提供一种计算机可读存储介质,该计算机可读存储介质上存储有版权认证程序,所述版权认证程序被处理器执行时实现如以上任一项实施例所述的版权认证方法的步骤。
本申请计算机可读存储介质的具体实施例与上述版权认证方法各实施例基本相同,在此不作赘述。
需要说明的是,在本文中,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者系统不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者系统所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包 括该要素的过程、方法、物品或者系统中还存在另外的相同要素。
上述本申请实施例序号仅仅为了描述,不代表实施例的优劣。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在如上所述的一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端设备(可以是手机,计算机,服务器,空调器,或者网络设备等)执行本申请各个实施例所述的方法。
以上仅为本申请的优选实施例,并非因此限制本申请的专利范围,凡是利用本申请说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本申请的专利保护范围内。

Claims (14)

  1. 一种版权认证方法,其中,所述版权认证方法包括:
    在接收到数字作品版权认证请求时,根据所述数字作品版权认证请求获取待认证数字作品和作品类型;
    根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象;
    将所述目标输入对象输入至所述目标审查分类模型中,得到审查结果,并基于所述审查结果判断审查是否通过;
    在审查通过时,将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证。
  2. 如权利要求1所述的版权认证方法,其中,所述根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象的步骤包括:
    若所述作品类型为文字作品,则确定对应的处理策略为第一处理策略,并确定目标审查分类模型为第一审查分类模型;
    基于所述第一处理策略对所述待认证数字作品进行切词处理,得到第一切词文本;
    将所述第一切词文本输入预设词向量模型,得到所述第一切词文本中各切词的第一词向量;
    根据所述第一词向量得到所述待认证数字作品对应的第一文档向量,其中,所述目标输入对象为所述第一文档向量。
  3. 如权利要求1所述的版权认证方法,其中,所述根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象的步骤包括:
    若所述作品类型为图片作品,则确定对应的处理策略为第二处理策略,并确定目标审查分类模型为第二审查分类模型;
    基于所述第二处理策略对所述待认证数字作品进行预处理,得到输入图片,其中,所述预处理包括缩放处理和灰度处理,所述目标输入对象为所述输入图片。
  4. 如权利要求1所述的版权认证方法,其中,所述根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象的步骤包括:
    若所述作品类型为音频作品,则确定对应的处理策略为第三处理策略,并确定目标审查分类模型为第三审查分类模型;
    基于所述第三处理策略将所述待认证数字作品转换成文字作品类型,得到转换后的待认证数字作品;
    对所述转换后的待认证数字作品进行切词处理,得到第二切词文本;
    将所述第二切词文本输入预设词向量模型,得到所述第二切词文本中各切词的第二词向量;
    根据所述第二词向量得到所述转换后的待认证数字作品对应的第二文档向量,其中,所述目标输入对象为所述第二文档向量。
  5. 如权利要求1至4中任一项所述的版权认证方法,其中,所述目标审查分类模型包括多个,所述审查结果的数量与所述目标审查分类模型的数量相同,所述基于所述审查结果判断审查是否通过的步骤包括:
    检测多个所述审查结果是否均为审查合格;
    若多个所述审查结果均为审查合格,则判定审查通过;
    若多个所述审查结果中至少存在一个为审查不合格,则判定审查不通过。
  6. 如权利要求1所述的版权认证方法,其中,若所述作品类型为文字作品,所述将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证的步骤包括:
    通过预设文档搜索引擎计算所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值;
    根据所述第一相似度值从所述已认证文字作品中筛选得到第一预设数量的相似文字作品;
    计算所述相似文字作品与所述待认证数字作品之间的第一最长公共子序列,并根据所述第一最长公共子序列的长度计算所述相似文字作品与所述待认证数字作品之间的长度比值,得到第一计算结果;
    检测所述第一计算结果中是否存在大于第一预设阈值的长度比值;
    若存在大于第一预设阈值的长度比值,则判定版权认证不通过;
    若不存在大于第一预设阈值的长度比值,则判定版权认证通过。
  7. 如权利要求6所述的版权认证方法,其中,所述通过预设文档搜索引擎计算所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值的步骤包括:
    通过预设文档搜索引擎对所述待认证数字作品进行分词处理,得到分词集;
    通过所述预设文档搜索引擎对预设认证作品库中的已认证文字作品进行倒排索引,并根据所述倒排索引结果计算所述分词集中各分词对应的分值;
    对各分词的分值进行加和处理,得到所述待认证数字作品与预设认证作品库中的已认证文字作品之间的第一相似度值。
  8. 如权利要求1所述的版权认证方法,其中,若所述作品类型为图片作品,所述将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证的步骤包括:
    通过预设图像检索引擎计算所述待认证数字作品与预设认证作品库中的已认证图片作品之间的第二相似度值;
    根据所述第二相似度值从所述已认证图片作品中筛选得到第二预设数量的相似图片作品;
    提取所述待认证数字作品的第一尺度不变特征变换SIFT特征向量,并提取所述相似图片作品的第二SIFT特征向量;
    计算所述第一SIFT特征向量与所述第二SIFT特征向量之间的余弦距离,得到第二计算结果;
    检测所述第二计算结果中是否存在大于第二预设阈值的余弦距离;
    若存在大于第二预设阈值的余弦距离,则判定版权认证不通过;
    若不存在大于第二预设阈值的余弦距离,则判定版权认证通过。
  9. 如权利要求1所述的版权认证方法,其中,若所述作品类型为音频作品,所述将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证的步骤包括:
    将所述待认证数字作品转换成文字作品类型,得到待认证音频文字作品;
    通过预设文档搜索引擎计算所述待认证音频文字作品与预设认证作品库中的已认证音频文字作品之间的第三相似度值;
    根据所述第三相似度值从所述已认证音频文字作品中检索得到第三预设数量的相似音频文字作品;
    计算所述相似音频文件作品与所述待认证音频文字作品之间的第二最长公共子序 列,并根据所述第二最长公共子序列的长度计算所述相似音频文字作品与所述待认证音频文字作品之间的长度比值,得到第三计算结果;
    检测所述第三计算结果中是否存在大于第三预设阈值的长度比值;
    若存在大于第三预设阈值的长度比值,则判定版权认证不通过;
    若不存在大于第三预设阈值的长度比值,则判定版权认证通过。
  10. 如权利要求1至4、6至9中任一项所述的版权认证方法,其中,所述版权认证方法还包括:
    在版权认证通过时,获取所述待认证数字作品的作品信息;
    基于所述作品信息生成对应的版权认证信息,并根据所述版权认证信息生成数据上链请求;
    将所述数据上链请求发送至版权认证联盟链,以使得所述版权认证联盟链基于所述数据上链请求完成对所述待认证数字作品的上链操作。
  11. 一种版权认证装置,其中,所述版权认证装置包括:
    第一获取模块,用于在接收到数字作品版权认证请求时,根据所述数字作品版权认证请求获取待认证数字作品和作品类型;
    处理模块,用于根据所述作品类型确定对应的处理策略和目标审查分类模型,基于所述处理策略对所述待认证数字作品进行处理,得到目标输入对象;
    审查模块,用于将所述目标输入对象输入至所述目标审查分类模型中,得到审查结果,并基于所述审查结果判断审查是否通过;
    版权认证模块,用于在审查通过时,将所述待认证数字作品与预设认证作品库中的已认证数字作品进行比对,以进行版权认证。
  12. 一种版权认证设备,其中,所述版权认证设备包括:存储器、处理器及存储在所述存储器上并可在所述处理器上运行的版权认证程序,所述版权认证程序被所述处理器执行时实现如权利要求1至10中任一项所述的版权认证方法的步骤。
  13. 一种版权认证系统,其中,所述版权认证系统包括版权认证设备和版权认证联盟链;其中,
    所述版权认证设备为如权利要求12所述的版权认证设备;
    所述版权认证联盟链,用于接收所述版权认证设备发送的数据上链请求;
    基于所述数据上链请求获取待上链数据信息,并基于共识算法完成对所述待上链数据信息的上链操作。
  14. 一种计算机可读存储介质,其中,所述计算机可读存储介质上存储有版权认证程序,所述版权认证程序被处理器执行时实现如权利要求1至10中任一项所述的版权认证方法的步骤。
PCT/CN2020/126232 2019-11-11 2020-11-03 版权认证方法、装置、设备、系统及计算机可读存储介质 Ceased WO2021093643A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201911093190.5 2019-11-11
CN201911093190.5A CN110781460B (zh) 2019-11-11 2019-11-11 版权认证方法、装置、设备、系统及计算机可读存储介质

Publications (1)

Publication Number Publication Date
WO2021093643A1 true WO2021093643A1 (zh) 2021-05-20

Family

ID=69390433

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/126232 Ceased WO2021093643A1 (zh) 2019-11-11 2020-11-03 版权认证方法、装置、设备、系统及计算机可读存储介质

Country Status (2)

Country Link
CN (1) CN110781460B (zh)
WO (1) WO2021093643A1 (zh)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113987418A (zh) * 2021-09-26 2022-01-28 支付宝(杭州)信息技术有限公司 一种混编作品的发布方法及装置
CN114611071A (zh) * 2022-02-23 2022-06-10 北京大学 一种基于联盟链的众包式数字内容版权检测方法
CN115033278A (zh) * 2022-06-29 2022-09-09 苏州浪潮智能科技有限公司 一种产品硬件认证配置管控方法、装置及存储介质
CN115495712A (zh) * 2022-09-28 2022-12-20 支付宝(杭州)信息技术有限公司 数字作品处理方法及装置
CN116186263A (zh) * 2023-03-01 2023-05-30 上海喜马拉雅科技有限公司 文档检测方法、装置、计算机设备及计算机可读存储介质
CN118568686A (zh) * 2024-05-16 2024-08-30 南京马特沃斯数字科技有限公司 基于区块链的数字版权管理平台

Families Citing this family (15)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110781460B (zh) * 2019-11-11 2025-05-23 深圳前海微众银行股份有限公司 版权认证方法、装置、设备、系统及计算机可读存储介质
CN111488555A (zh) * 2020-04-02 2020-08-04 上海七印信息科技有限公司 版权认证方法、装置、计算机设备和存储介质
CN111552829B (zh) * 2020-05-07 2023-06-27 京东科技信息技术有限公司 用于分析图像素材的方法和装置
CN112131859A (zh) * 2020-08-25 2020-12-25 中央民族大学 藏文作文抄袭检测原型系统
CN112163243A (zh) * 2020-10-09 2021-01-01 成都乐链科技有限公司 基于区块链的数字资产审查与存证方法、确权方法及装置
CN112199951A (zh) * 2020-11-04 2021-01-08 支付宝(杭州)信息技术有限公司 一种事件信息生成的方法及装置
CN112559975A (zh) * 2020-11-25 2021-03-26 山东浪潮质量链科技有限公司 一种基于区块链的数字媒体版权实现方法、设备及介质
CN112487088B (zh) * 2020-11-26 2021-08-24 中国搜索信息科技股份有限公司 一种基于区块链的融媒体资源版权保护方法
CN113536288B (zh) * 2021-06-23 2023-10-27 上海派拉软件股份有限公司 数据认证方法、装置、认证设备及存储介质
CN113949515A (zh) * 2021-09-09 2022-01-18 卓尔智联(武汉)研究院有限公司 一种数字版权信息处理方法及装置、存储介质
CN113515664A (zh) * 2021-09-14 2021-10-19 北京远鉴信息技术有限公司 异常音频的确定方法、装置、电子设备及可读存储介质
CN114495139B (zh) * 2022-01-24 2025-01-17 东软教育科技集团有限公司 一种基于图像的作业查重系统及方法
CN115221473A (zh) * 2022-07-05 2022-10-21 黄冈市学海园文化发展有限公司 一种数字版权管理设备认证系统
CN116205219B (zh) * 2023-03-09 2025-09-16 上海喜马拉雅科技有限公司 文本相似度检测方法、装置、计算机设备及可读存储介质
CN120493220B (zh) * 2025-07-16 2025-11-14 浪潮云洲工业互联网有限公司 一种基于区块链的ai数字内容版权存证方法、设备及介质

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101645117A (zh) * 2008-08-06 2010-02-10 武汉大学 一种媒体发布网络中的发布内容控制方法
CN101826099A (zh) * 2010-02-04 2010-09-08 蓝盾信息安全技术股份有限公司 一种相似文档识别、文档扩散度确定的方法及系统
CN105550381A (zh) * 2016-03-17 2016-05-04 北京工业大学 一种基于改进sift特征的高效图像检索方法
CN106874253A (zh) * 2015-12-11 2017-06-20 腾讯科技(深圳)有限公司 识别敏感信息的方法及装置
CN107832384A (zh) * 2017-10-28 2018-03-23 北京安妮全版权科技发展有限公司 侵权检测方法、装置、存储介质和电子设备
CN109145529A (zh) * 2018-09-12 2019-01-04 重庆工业职业技术学院 一种用于版权认证的文本相似性分析方法与系统
CN110781460A (zh) * 2019-11-11 2020-02-11 深圳前海微众银行股份有限公司 版权认证方法、装置、设备、系统及计算机可读存储介质

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20040098876A (ko) * 2003-05-16 2004-11-26 (주)노마글로벌 네트워크 상에서 디지털 저작권 보호를 위한 원격 인증 시스템 구성 방법.
TWI247518B (en) * 2004-04-08 2006-01-11 Jau-Ming Shr Copyright protection method of digital publication and system thereof
CN103390121B (zh) * 2012-05-10 2016-12-14 北京大学 数字作品权属认证方法和系统
CN109684786A (zh) * 2018-11-05 2019-04-26 深圳变设龙信息科技有限公司 一种基于区块链的版权登记方法、装置及终端设备
CN109492351A (zh) * 2018-11-23 2019-03-19 北京奇眸科技有限公司 基于区块链的版权保护方法、装置及可读存储介质
CN110188515A (zh) * 2019-05-16 2019-08-30 中细软集团有限公司 一种区块链网络数字作品登记方法和客户端

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101645117A (zh) * 2008-08-06 2010-02-10 武汉大学 一种媒体发布网络中的发布内容控制方法
CN101826099A (zh) * 2010-02-04 2010-09-08 蓝盾信息安全技术股份有限公司 一种相似文档识别、文档扩散度确定的方法及系统
CN106874253A (zh) * 2015-12-11 2017-06-20 腾讯科技(深圳)有限公司 识别敏感信息的方法及装置
CN105550381A (zh) * 2016-03-17 2016-05-04 北京工业大学 一种基于改进sift特征的高效图像检索方法
CN107832384A (zh) * 2017-10-28 2018-03-23 北京安妮全版权科技发展有限公司 侵权检测方法、装置、存储介质和电子设备
CN109145529A (zh) * 2018-09-12 2019-01-04 重庆工业职业技术学院 一种用于版权认证的文本相似性分析方法与系统
CN110781460A (zh) * 2019-11-11 2020-02-11 深圳前海微众银行股份有限公司 版权认证方法、装置、设备、系统及计算机可读存储介质

Cited By (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113987418A (zh) * 2021-09-26 2022-01-28 支付宝(杭州)信息技术有限公司 一种混编作品的发布方法及装置
CN114611071A (zh) * 2022-02-23 2022-06-10 北京大学 一种基于联盟链的众包式数字内容版权检测方法
CN114611071B (zh) * 2022-02-23 2024-10-18 北京大学 一种基于联盟链的众包式数字内容版权检测方法
CN115033278A (zh) * 2022-06-29 2022-09-09 苏州浪潮智能科技有限公司 一种产品硬件认证配置管控方法、装置及存储介质
CN115033278B (zh) * 2022-06-29 2024-04-30 苏州浪潮智能科技有限公司 一种产品硬件认证配置管控方法、装置及存储介质
CN115495712A (zh) * 2022-09-28 2022-12-20 支付宝(杭州)信息技术有限公司 数字作品处理方法及装置
CN115495712B (zh) * 2022-09-28 2024-04-16 支付宝(杭州)信息技术有限公司 数字作品处理方法及装置
CN116186263A (zh) * 2023-03-01 2023-05-30 上海喜马拉雅科技有限公司 文档检测方法、装置、计算机设备及计算机可读存储介质
CN118568686A (zh) * 2024-05-16 2024-08-30 南京马特沃斯数字科技有限公司 基于区块链的数字版权管理平台

Also Published As

Publication number Publication date
CN110781460B (zh) 2025-05-23
CN110781460A (zh) 2020-02-11

Similar Documents

Publication Publication Date Title
WO2021093643A1 (zh) 版权认证方法、装置、设备、系统及计算机可读存储介质
US12183056B2 (en) Adversarially robust visual fingerprinting and image provenance models
CN106650799B (zh) 一种电子证据分类提取方法及系统
WO2022174491A1 (zh) 基于人工智能的病历质控方法、装置、计算机设备及存储介质
US20110119293A1 (en) Method And System For Reverse Pattern Recognition Matching
WO2022142032A1 (zh) 手写签名校验方法、装置、计算机设备及存储介质
WO2022134584A1 (zh) 房产图片验证方法、装置、计算机设备及存储介质
CN107609389B (zh) 一种基于图像内容相关性的验证方法及系统
US11880798B2 (en) Determining section conformity and providing recommendations
CN110532543B (zh) 证据材料的分析处理方法、装置、计算机设备和存储介质
CN114513355A (zh) 恶意域名检测方法、装置、设备及存储介质
CN111259115B (zh) 内容真实性检测模型的训练方法、装置和计算设备
WO2020000752A1 (zh) 仿冒移动应用程序的判别方法及系统
WO2024120245A1 (zh) 视频信息摘要生成方法、装置、存储介质及计算机设备
CN111597490A (zh) Web指纹识别方法、装置、设备及计算机存储介质
CN114238643A (zh) 敏感信息识别模型的构建、敏感信息识别方法及装置
CN113705185B (zh) 一种动态合同签署方法、装置、设备和介质
CN115689811A (zh) 基于区块链的电子遗嘱生成方法、资产继承方法及系统
CN116340991A (zh) Ip图库素材资源的大数据管理方法、装置以及电子设备
US10803115B2 (en) Image-based domain name system
CN119441553A (zh) 数字资产检索方法、装置、计算机设备、可读存储介质和程序产品
Song et al. A multi-modality feature fusion method for android malware detection
CN117669582A (zh) 一种基于深度学习的工程咨询处理方法、装置及电子设备
CN117591770B (zh) 政策的推送方法、装置以及计算机设备
CN115883205B (zh) 一种电力监控系统弱口令检测方法和装置

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20886598

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20886598

Country of ref document: EP

Kind code of ref document: A1