WO2017128101A1 - 名称列表展示、处理方法及客户端、服务器 - Google Patents

名称列表展示、处理方法及客户端、服务器 Download PDF

Info

Publication number
WO2017128101A1
WO2017128101A1 PCT/CN2016/072307 CN2016072307W WO2017128101A1 WO 2017128101 A1 WO2017128101 A1 WO 2017128101A1 CN 2016072307 W CN2016072307 W CN 2016072307W WO 2017128101 A1 WO2017128101 A1 WO 2017128101A1
Authority
WO
WIPO (PCT)
Prior art keywords
predetermined
name
unicode
character
category
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/072307
Other languages
English (en)
French (fr)
Inventor
张辰
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba Group Holding Ltd
Original Assignee
Alibaba Group Holding Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba Group Holding Ltd filed Critical Alibaba Group Holding Ltd
Priority to PCT/CN2016/072307 priority Critical patent/WO2017128101A1/zh
Priority to CN201680080048.5A priority patent/CN108702407A/zh
Publication of WO2017128101A1 publication Critical patent/WO2017128101A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04MTELEPHONIC COMMUNICATION
    • H04M1/00Substation equipment, e.g. for use by subscribers
    • H04M1/26Devices for calling a subscriber
    • H04M1/27Devices whereby a plurality of signals may be stored simultaneously
    • H04M1/274Devices whereby a plurality of signals may be stored simultaneously with provision for storing more than one subscriber number at a time, e.g. using toothed disc
    • H04M1/2745Devices whereby a plurality of signals may be stored simultaneously with provision for storing more than one subscriber number at a time, e.g. using toothed disc using static electronic memories, e.g. chips
    • H04M1/275Devices whereby a plurality of signals may be stored simultaneously with provision for storing more than one subscriber number at a time, e.g. using toothed disc using static electronic memories, e.g. chips implemented by means of portable electronic directories

Definitions

  • the present application relates to the field of data processing, and in particular, to a name list display, a processing method, and a client and a server.
  • the Internet has changed people's daily lives so that people can communicate without meeting each other.
  • This kind of communication is not limited to dialogue between language and text, but has been extended to commercial trade, such as online shopping, financial management and so on.
  • a list of contacts is usually set up in the software to provide the user with information about the contact.
  • the data of the contact list is sorted and sorted, so that the user can quickly query the target contact.
  • a pinyin4j dictionary is integrated in the software, which includes pinyin corresponding to each Chinese character, so that the input Chinese characters can find the corresponding pinyin through the search dictionary.
  • a list of contacts may exist in some software for the user to view the contacts, thereby facilitating interaction between the users. For example, doing instant messaging, or making financial transfers, and so on.
  • users in different countries may use different languages.
  • users in the same country may also store different contact names in different languages. For example, a Chinese user has Japanese friends, American friends, North Korean friends, and so on. In such a case, when generating a contact list, it is necessary to sort the list for the contact name without the language. If a dictionary of words and pronunciations is set for each language, making a dictionary and subsequent updates to update the dictionary requires a lot of manual labor, which is time consuming and laborious.
  • the embodiment of the present application provides a name list display, a processing method, a client, and a server, which can effectively adapt to a multi-language generated name list.
  • the present application provides a method for displaying a name list, which includes: obtaining name data; wherein the name data includes at least one name string; and obtaining a uniformity of a first predetermined character in the name string Determining a predetermined code segment corresponding to the Unicode of the first predetermined character, dividing the name string a predetermined category corresponding to the predetermined code segment; wherein the number of the predetermined code segments is at least one, each of the predetermined code segments corresponds to a predetermined category; and the category identifier of the predetermined category is correspondingly displayed in the display interface, And a name string corresponding to the predetermined category.
  • the embodiment of the present application further provides a client, which includes: a data acquisition module, configured to acquire name data; wherein the name data includes at least one name string; a Unicode acquisition module, configured to acquire the name character a Unicode of the first predetermined character in the string; a predetermined class division module, configured to determine a predetermined code segment corresponding to the Unicode of the first predetermined character, and dividing the name string into a predetermined category corresponding to the predetermined code segment Wherein the number of the predetermined code segments is at least one, each of the predetermined code segments corresponds to a predetermined category; a display module, configured to display the category identifiers of the predetermined category in the display interface, and the predetermined The name string corresponding to the category.
  • a data acquisition module configured to acquire name data
  • the name data includes at least one name string
  • a Unicode acquisition module configured to acquire the name character a Unicode of the first predetermined character in the string
  • a predetermined class division module configured to determine a predetermined code segment corresponding
  • the embodiment of the present application further provides a client, including: a processor, a display, the processor is configured to obtain name data; wherein the name data includes at least one name string; and is used to obtain the name string a unified code of the first predetermined character; determining a predetermined code segment corresponding to the Unicode of the first predetermined character, dividing the name string into a predetermined category corresponding to the predetermined code segment; wherein the predetermined code segment The number of the at least one, each of the predetermined code segments corresponds to a predetermined category; controlling the display to display the category identifier of the predetermined category in the display interface, and a name string corresponding to the predetermined category.
  • the embodiment of the present application further provides a name list processing method, including: acquiring name data; wherein the name data includes at least one name string; obtaining a Unicode of the first predetermined character in the name string; a predetermined code segment corresponding to the Unicode of the first predetermined character, the name string is divided into a predetermined category corresponding to the predetermined code segment; wherein the number of the predetermined code segments is at least one, each of the The predetermined code segment corresponds to a predetermined category.
  • the embodiment of the present application further provides a server, including: a data obtaining module, configured to acquire name data; wherein the name data includes at least one name string; a Unicode obtaining module, configured to obtain the name string a Unicode of the first predetermined character; a predetermined class division module, configured to determine a predetermined code segment corresponding to the Unicode of the first predetermined character, and divide the name string into a predetermined category corresponding to the predetermined code segment; The number of the predetermined code segments is at least one, and each of the predetermined code segments corresponds to a predetermined category.
  • the embodiment of the present application further provides a server, including: a processor, where the processor is configured to obtain name data; wherein the name data includes at least one name string; and is used to obtain the first name string a Unicode of the predetermined character; determining a predetermined code segment corresponding to the Unicode of the first predetermined character, and dividing the name string into a predetermined category corresponding to the predetermined code segment; wherein the number of the predetermined code segments is At least one of each of the predetermined code segments corresponds to a predetermined category.
  • the embodiment of the present application is based on the name string.
  • a Unicode of a predetermined character dividing the name string into a predetermined category corresponding to the predetermined code segment, and correspondingly setting a category identifier to the predetermined category, and correspondingly displaying the category identifier of the predetermined category in the display interface
  • the name string corresponding to the predetermined category the one-to-one correspondence between the Unicode and the name string is realized, and the languages of the plurality of countries are classified and displayed in the same name list.
  • the user can quickly locate the specific name string in the name list by displaying the category identifier, and find the corresponding target name.
  • the name list display method is freed from the traditional preset dictionary in the process of implementation, and uses the existing unified code to sort and sort, which reduces a large amount of human labor and saves human resource costs.
  • FIG. 1 is a flowchart of a method for displaying a name list according to an embodiment of the present application
  • FIG. 2 is a flowchart of a method for displaying a name list according to an embodiment of the present application
  • FIG. 3 is a flowchart of a method for displaying a name list according to an embodiment of the present application
  • FIG. 4 is a flowchart of a method for displaying a name list according to an embodiment of the present application
  • FIG. 5 is a block diagram of a client provided by an embodiment of the present application.
  • FIG. 6 is a block diagram of a client provided by an embodiment of the present application.
  • FIG. 7 is a flowchart of a method for processing a name list according to an embodiment of the present application.
  • FIG. 8 is a flowchart of a method for processing a name list according to an embodiment of the present application.
  • FIG. 9 is a flowchart of a method for processing a name list according to an embodiment of the present application.
  • FIG. 10 is a flowchart of a method for processing a name list according to an embodiment of the present application.
  • FIG. 11 is a block diagram of a server according to an embodiment of the present application.
  • FIG. 12 is a schematic diagram of a predetermined class division module according to an embodiment of the present application.
  • an embodiment of the present application provides a method for displaying a name list, which may include the following steps.
  • Step S10 Acquire name data; wherein the name data includes at least one name string.
  • the client may be a communication device having a network communication function, such as a desktop computer, a notebook computer, a tablet computer, a smart phone, and a smart wearable device.
  • the client can also be software running in the above communication device.
  • the client can be used by the user, and can perform interactive activities such as instant messaging, and can perform various e-commerce activities such as online transaction and online electronic payment in the instant communication process.
  • the name data may be pre-stored in the client.
  • the pre-stored name data may be retrieved by the processor.
  • the name data can also be obtained by downloading from the server side.
  • the name data can be downloaded from the server through the data communication module.
  • the name data may also be temporarily obtained by the user in an input manner.
  • the manner of obtaining the name data may be other methods, for example, accepting the name data sent by other clients, and the like, which is not specifically limited herein.
  • the name data may be a set of names that the user needs to use when performing interactive activities such as instant messaging and e-commerce.
  • the name may be specifically: name, nickname, company name, account name, location name, etc., and the application is not specifically limited herein.
  • the name data includes at least one name string.
  • the name string may be a specific name in the name, nickname, company name, account name, and place name.
  • the name string may be a specific name, such as “Wang Erxiao”.
  • the name string may also be a specific account name, such as a bank card number of a bank: "622848440637874213".
  • the form of the name string may be at least one word, word, phrase, or the like in a certain language, or may be a symbol, a number, a letter, a plurality of combinations, etc., of course, the name string may also be in other forms. This application is not specifically limited herein.
  • Step S12 Acquire a unified code of the first predetermined character in the name string.
  • the Unicode is used to set a uniform and unique encoding for characters in each language to meet the requirements of text conversion and processing across languages and platforms.
  • the Unicode has a character set that can be mapped with a predetermined number of characters such that characters in the character set of the Unicode have a one-to-one correspondence with characters in each language. That is, a Unicode unique corresponds to one character.
  • the predetermined number of digits may be hexadecimal, binary, etc., which are not stated here.
  • the Unicode code is a typical Unicode code that uses hexadecimal digits 0-0x10FFFF to map characters in various languages.
  • Chinese, Japanese, and Korean CJK Choinese Japanese Korean three characters occupy the 0x3000 to 0x9FFF part of Unicode.
  • the UCS-2 specification code is commonly used in Unicode. Specifically, it encodes one character with two bytes. For example, the encoding of the Chinese character "jing" is 0x7ECF. Since character encoding is generally expressed in hexadecimal, in order to be in decimal Differentiate, hexadecimal starts with 0x, 0x7ECF converts to decimal is 32463, UCS-2 uses two bytes to encode characters, two bytes are 16-bit binary, and 2's 16th power equals 65536, so UCS-2 It can encode up to 65536 characters.
  • a unified character set may be stored in the client, and the unified character set may include a character and a Unicode corresponding to the character.
  • the unified character set may also be stored on the server end, and the client may obtain related characters in the unified character set and a unified code corresponding to the characters after establishing communication with the server end through the communication module.
  • Each of the first predetermined characters may uniquely correspond to a unified code.
  • the first predetermined character may be a first character of the name string, or may be a combination of a first character and a second character in the name string, or may be a second character alone, or may also be Is a predetermined character in the name string pre-specified by the user.
  • the first predetermined character is not limited to the above description.
  • Other changes may be made by those skilled in the art in light of the technical spirit of the present application.
  • the functions and effects thereof are the same or similar to the present application, they should be covered by the present application.
  • the first predetermined character may uniquely have a unified code, according to the first predetermined character, by using a lookup manner or by a mapping relationship between the first predetermined character and the Unicode, A Unicode of the first predetermined character in the name string.
  • the first predetermined character can be converted into a hexadecimal number, and then the search is performed in the unified character set. If the same hexadecimal number in the unified character set is found, the Unicode of the first predetermined character can be uniquely determined.
  • the manner of obtaining is not limited to the above description.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • Step S14 determining a predetermined code segment corresponding to the Unicode of the first predetermined character, and dividing the name string into a predetermined category corresponding to the predetermined code segment; wherein the number of the predetermined code segments is at least one.
  • Each of the predetermined code segments corresponds to a predetermined category.
  • the language ordering mechanism in the programming language API can be used to divide the Unicode of the first predetermined character according to the pronunciation, spelling and the like of the language. At least one predetermined code segment is determined. Each of the predetermined code segments may be correspondingly set to a predetermined category, or a plurality of predetermined code segments may be correspondingly set to the same predetermined category, thereby classifying the characters of the plurality of languages into a limited category.
  • unified coding is not arranged according to the pronunciation of the language.
  • a sorter such as a collator in Java, may be used to sort the characters according to the characteristics of the pronunciation of the language.
  • the Unicode corresponding to the sorted characters is divided into at least one pre- The specified code segment. Specifically, when the sorter collator performs sorting, the order of the two characters can be compared according to the characteristics of the language, and the characters in the unified character set can be segmented according to the language features.
  • the first predetermined character of all "ah” (a) pronunciations can be ranked in front of the first predetermined characters of all "bar” pronunciations according to their pronunciation.
  • the first predetermined character of all "bar” pronunciations is placed in front of the first predetermined character of all "can” (ca) pronunciations.
  • the first predetermined character in the name data is divided into "a” to "z” 26 predetermined code segments according to the pronunciation of the characters.
  • the predetermined category is a set of a plurality of name strings.
  • the predetermined category may be implemented in a data structure by a queue, a data stack, an array, or the like.
  • the name strings of the first predetermined character pronunciation belonging to the same predetermined code segment may be stored under the same array to form a predetermined category.
  • "Chen Yi”, “Cheng Er”, “Cai San”, and “Cheng Si” are stored under the same array to form a predetermined category.
  • the manner in which the data structure of the predetermined category is formed is not limited to the above description. Other changes may be made by those skilled in the art in light of the technical spirit of the present application. However, as long as the functions and effects thereof are the same or similar to the present application, they should be covered by the present application.
  • the collator in Java may be used to sequentially arrange the Unicode codes corresponding to the first predetermined characters according to the pronunciation, spelling and the like of the language, and is divided into At least one predetermined code segment.
  • the first predetermined character may be compared with a character corresponding to the Unicode of the segment header and/or the segment end of the predetermined code segment. If the order of the first predetermined character is after the first character of the predetermined code segment, but before the end of the paragraph character, or in the order before the first character of the next segment, it may be determined that the first predetermined character belongs to the predetermined Code segment.
  • a correspondence relationship may be formed between the predetermined code segment header and/or the segment end Unicode and the predetermined class. Please refer to Table 1.
  • the predetermined category of the name string may be determined according to the correspondence between the predetermined code segment and the predetermined category.
  • the collator in the existing Java is utilized, and the sorting according to the pronunciation, spelling, and the like of the language is implemented.
  • the specific ordering manner is not limited to the foregoing description, and those skilled in the art may also make other manners, such as selecting other existing sorters or writing according to the sorting function, etc., inspired by the technical essence of the present application.
  • the functions and effects achieved by the same or similar to the present application are within the scope of the present application.
  • a unified character set may be integrated in the operating system, that is, a Unicode corresponding to each character may be obtained from an operating system.
  • the segment header Unicode of the predetermined code segment may be stored in the client, so that when the Unicode segment is segmented, the segment header Unicode may be used to distinguish different predetermined code segments; or the predetermined code segment is a segment header Unicode and a segment end segment Unicode of the predetermined code segment.
  • the client stores a Unicode of 26 predetermined code segments, a total of 26; or a segment header Unicode and a segment end of the predetermined code segment
  • the Unicode code a total of 52. Therefore, the storage space required is relatively small compared to the client.
  • the unified coding is not arranged according to the pronunciation and pronunciation of each language.
  • the Unicode code corresponding to the first predetermined character may be corresponding to an index value according to the pronunciation and other characteristics of the language, and may be sorted according to the index value, and then the index value may be sorted.
  • the latter Unicode code is divided into at least one predetermined code segment.
  • a corresponding index value corresponding to the Unicode of the first predetermined character may be given according to the pronunciation thereof.
  • the index value can be viewed as a number. For example, “ah” (a), its index value may be 1; and “bar” (ba), its index value may be 10, then all "ah” (a) to "bar” (ba) between the pronunciation of Chinese characters Its index value is between 1-10.
  • the predetermined index corresponding to the first index value of the encoded segment is [1, 10) is the A category, that is, the word whose initial letter a is pronounced, such as "ah", "an”, etc., is assigned to the category A.
  • the predetermined index value corresponding to the first index value of the code segment segment [10, 20) is the B category, that is, the word whose initial letter b is pronounced, such as "bar", "being", etc., is assigned to the B category.
  • the Z category is compiled.
  • the entire Chinese code block is divided into 26 code segments, corresponding to 26 predetermined categories.
  • the first reservation may be determined by performing a lookup comparison between the index value of the Unicode of the first predetermined character and the segment header index value and/or the segment tail index value of the preset code segment.
  • the predetermined code segment corresponding to the character. Since the predetermined code segment corresponds to a predetermined category, the name string may be divided into a predetermined category corresponding to the predetermined code segment.
  • the index value of the corresponding Unicode code is compared with a predetermined code segment header index value and/or a segment tail index value, thereby determining the first predetermined character.
  • the corresponding predetermined code segment and the predetermined category Please refer to Table 2 for details. For example, if the index value corresponding to the Unicode of the first predetermined character is greater than or equal to 1 and less than 10, the predetermined code segment corresponding to the first predetermined character and the predetermined category may be determined, and the The name string corresponding to a predetermined character is divided into the A category of the predetermined code segment.
  • a unified character set may be integrated in the operating system, that is, the unified code corresponding to each first predetermined character may be obtained from an operating system.
  • the segment header index value of the predetermined code segment that can be stored only in the client, so that when the Unicode segment is segmented, different predetermined code segments are distinguished by the segment header index value and/or the segment tail index value.
  • the client stores an index value of 26 predetermined code segments, for a total of 26; or a segment header index value and a segment tail of a predetermined code segment.
  • the index value is 52 in total. Therefore, the storage space required is relatively small compared to the client.
  • the manner of determining the predetermined code segment corresponding to the Unicode of the first predetermined character is not limited to the above description.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • Step S16 correspondingly displaying the category identifier of the predetermined category in the display interface, and a name string corresponding to the predetermined category.
  • the predetermined category may correspond to a category identifier, and when displayed in the display interface, Displaying the category identifier of the predetermined category together with the name string corresponding to the predetermined category to form a list of names that are convenient for the user to find.
  • the category identifier may be an identifier that is convenient for the user to identify according to the user's language habits. For example, for a user who uses Chinese, the category identifier may be the first letter of the first predetermined character pronunciation. For a user who uses English, the category identifier may be the first letter of the first predetermined character. Of course, for other language types, the category identifier may also be the phonetic symbol of the first predetermined character or the initials of the spelling, etc., and the examples are not repeated here.
  • the name string corresponding to the predetermined category may all display or display part of the name string or not display the name string.
  • the predetermined category includes a plurality of name strings.
  • a predetermined number of name strings can be displayed under the category mark corresponding to the predetermined category.
  • a maximum of three name strings can be set.
  • the name list display method is based on the Unicode of the first predetermined character in the name string, and divides the name string into a predetermined category corresponding to the predetermined code segment, and correspondingly sets the predetermined category.
  • a category identifier correspondingly displaying the category identifier of the predetermined category and the name string corresponding to the predetermined category in the display interface, realizing the one-to-one correspondence between the Unicode and the name string, and the languages of the plurality of countries are Classified display in the same name list.
  • the user can quickly locate the specific name string in the name list by displaying the category identifier, and find the corresponding target name.
  • the name list display method is freed from the traditional preset dictionary in the process of implementation, and uses the existing unified code to sort and sort, which reduces a large amount of human labor and saves human resource costs.
  • a Unicode code is usually integrated to serve as a basis for encoding and decoding text data when flowing data across regions or data platforms.
  • a Unicode code is usually integrated to serve as a basis for encoding and decoding text data when flowing data across regions or data platforms.
  • the Unicode by sorting and sorting the name list by using the Unicode, it is possible to implement the sorting of the name list by way of not implementing the additional dictionary introduced in the existing client. Specifically, for example, a pinyin4j dictionary may not be introduced, which saves storage space.
  • the specific content of the name data may include at least one of a person name, a company name, an account name, and a place name.
  • the name data may include at least one name string.
  • the name string may be a person name, a company name, an account name, a place name, and the like, and the present application is not specifically limited herein.
  • the name of the person may be the real name of the user, or may be a nickname set by the user.
  • the name data may be of a type including only one name, for example, only a person name; Include multiple name categories, such as including person name, company name, account name, and so on.
  • the name data includes a plurality of name types at the same time, if there is a correspondence between the plurality of name types, the plurality of name types may be divided, for example, a name of one of the plurality of name types may be set.
  • the string is the primary name, and other names in the multiple name categories are set as the dependent names.
  • only the primary name may be displayed in the display interface, and other dependent names corresponding to the primary name may be stored corresponding to the primary name.
  • the slave name can be displayed when the location where the master name is located detects an electrical signal triggered by the click.
  • operations such as copying, editing, and the like may be performed on the subordinate name.
  • the specific content of the name data is “Company A Xiaowang China Bank Account XX”, the name of the person is set as the master name, and the names of other companies and accounts are subordinate names.
  • the main name “Xiaowang” is displayed in the display interface, and other subordinate names “Company” and “Bank of China Account XX” corresponding to “Xiaowang” may be “small” with the main name.
  • Wang corresponds to storage.
  • the subordinate names "Company A” and “Bank of China Account XX” may be displayed correspondingly. Further, the subordinate names may be copied, edited, and the like.
  • the first predetermined character may be the first character of the name string.
  • the name string may be a person name, a company name, an account name, a place name, and the like.
  • the name string can include one or more characters.
  • the first character of the name string can typically represent the entire name string in many cases. For example, when the name string is a person name, the first character thereof may be a last name, and when the name string is a company name, the first character may also be the first word of the company name.
  • the first predetermined character is set as the first character of the name string, the subsequent list of names formed by the name string conforms to the usage habit of the user, and is helpful for the user to perform query based on the formed name list. .
  • the predetermined code segment corresponding to the Unicode of the first predetermined character may be determined by a collator in Java.
  • a sorter such as a collator in Java
  • the Unicode corresponding to each character after sorting is divided into at least one predetermined code segment. Specifically, when the sorter collator performs sorting, the order of the two characters can be compared according to the characteristics of the language, and the characters in the unified character set can be segmented according to the language features.
  • the first predetermined character of all "ah” (a) pronunciations can be ranked in front of the first predetermined characters of all "bar” pronunciations according to their pronunciation.
  • the first predetermined character of all "bar” pronunciations is placed in front of the first predetermined character of all "can” (ca) pronunciations.
  • the Chinese characters in the name data are divided into “a” to "z” 26 predetermined code segments according to the pronunciation of the first predetermined character.
  • the existing collator in Java is used to realize pronunciation, spelling, and the like according to the language.
  • Features are compared to sort.
  • the specific ordering manner is not limited to the foregoing description, and those skilled in the art may also make other manners, such as selecting other existing sorters or writing according to the sorting function, etc., inspired by the technical essence of the present application.
  • the functions and effects achieved by the same or similar to the present application are within the scope of the present application.
  • a name string is newly added, and when it is determined that the first predetermined character of the newly added name string specifically belongs to a corresponding predetermined code segment, the first The predetermined character is compared with a character corresponding to the Unicode of the beginning and/or the end of the predetermined code segment. If the order of the first predetermined characters is in an order after the first character of the predetermined code segment but before the end of the paragraph character, or an order before the first character of the next segment, it may be determined that the first predetermined character belongs to the Predetermine the code segment.
  • At least one of said predetermined code segments constitutes a language coded block; wherein different said language coded blocks have at least partially identical predetermined categories.
  • the step of dividing the predetermined code into the predetermined categories includes: dividing the name string into the same when the predetermined categories corresponding to the predetermined code segments in which the name strings belonging to the different language coded blocks are located are the same A preset data set of a predetermined category.
  • different language code blocks can be set for different language types in the Unicode.
  • it may include a Chinese-Japanese Korean code block, an English code block, a Russian code block, an Arabic code block, and the like.
  • different language coding blocks may have at least partially identical predetermined categories.
  • the collator in Java can be used to set a plurality of predetermined code segments according to the pronunciation and the like of the language.
  • the Unicode corresponding to the first predetermined character may be divided into "a" to "z” 26 predetermined code segments according to the pronunciation characteristics of the first predetermined character.
  • the "a" to "z” 26 predetermined code segments respectively correspond to predetermined categories of A to Z26 categories.
  • the English coding block please refer to Table 3, using the collator in Java to set a plurality of predetermined code segments according to the character spelling characteristics.
  • a plurality of predetermined code segments may be set according to the order in which the Unicode codes are themselves arranged.
  • the first predetermined character may be divided into A to Z26 predetermined code segments according to a spelling characteristic of the first predetermined character.
  • the predetermined category corresponding to the first predetermined code segment first Unicode 97 is the category A, and the name string whose initial letter is A is assigned to the predetermined category A.
  • the predetermined category corresponding to the first predetermined code segment first Unicode 98 is the B category, and the name string whose initial letter B is assigned to the predetermined category B.
  • a predetermined category Z is prepared.
  • the entire English coded block is divided into 26 code segments, corresponding to 26 predetermined categories.
  • the name string belonging to the same predetermined category may be divided into the predetermined category to form the pre- Set the data set to achieve mixed sorting between Chinese and English.
  • the segment header or the segment end of each corresponding predetermined code segment may be set with an index value, when the English code block When the corresponding index value is the same as the index value corresponding to the Chinese coded block, it may be determined that the name string with the same index value belongs to the same category.
  • the Chinese coded block and the English coded block may not be divided into the same predetermined category, that is, each is divided according to a respective predetermined category, and no shuffling is performed.
  • the English coding block and the noon coding block are classified and arranged according to the order of the Unicode, and may be classified and arranged in the order of the English coding block and the Chinese coding block as a whole.
  • the labels of the specific classifications of the different coding blocks may be the same or different, and the application is not specifically limited herein.
  • the type of the language coding block is not limited to the example Chinese language coding block and the English language coding block, and the number is not limited to the two language coding blocks.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • the following steps may be included in the step S14 of determining a predetermined code segment.
  • Step S140 Determine a language coding block corresponding to the Unicode of the first predetermined character in the name string.
  • Step S142 In the language coding block, determining a predetermined code segment in which the Unicode of the first predetermined character is located.
  • the corresponding language coded block when determining the predetermined code segment, may be first determined according to the Unicode of the first predetermined character in the name string. Based on the determined language coding block, a predetermined code segment in which the Unicode of the first predetermined character is located is further determined.
  • different language code blocks may be set in the Unicode for different language types.
  • the Unicode of the first character of each language code block and the number of characters are known.
  • the Unicode of the first predetermined character may be compared with the Unicode of each linguistic coding block.
  • the Unicode of the first predetermined character is greater than or equal to the first character Unicode of a certain linguistic coding block, and the difference between the Unicode of the first linguistic coding block and the Unicode code of the linguistic coding block is less than the number of the linguistic coding blocks, Then, it can be determined that the first predetermined character is a character in the language encoding block.
  • the Unicode of the first predetermined character may be searched in the determined language encoding block to quickly determine A predetermined code segment in which the Unicode of the first predetermined character is located.
  • the determined language coding block generally includes a plurality of language code segments, and each language code segment corresponds to at least one unified code.
  • the Unicode of the first predetermined character may be compared with the Unicode of the language code segment, and if the same, the uniformity of the first predetermined character may be determined. The predetermined code segment in which the code is located.
  • the English language coding block may include 26 language code segments.
  • each predetermined code segment corresponds to a Unicode code.
  • the Unicode of the predetermined character may be compared with the Unicode of each predetermined code segment, and if the same, the A predetermined code segment in which a Unicode of a predetermined character is located.
  • the method may further include: dividing the name string into a preset specified category when there is no language encoding block corresponding to the Unicode of the first predetermined character in the name string.
  • a preset specified category may be set in advance on the client for categorizing the name string in which the corresponding language coded block is not found.
  • the preset specified category may be a single category arranged after the predetermined category, which may be displayed with a special tag. For example, "#" can be used to indicate that the following includes a preset specified category that does not belong to the language coded block included in the Unicode.
  • step S16 of displaying the category identifier and the name string may further include:
  • the name strings belonging to the same predetermined category are arranged in the coding order of the Unicode of the first predetermined character.
  • the first predetermined character of the name string may correspond to a unique Unicode.
  • the name strings belonging to the same predetermined category are sorted, they may be arranged in the coding order of the Unicode of the first predetermined characters.
  • the encoding of the Unicode of the first predetermined character since the encoding of the Unicode of the first predetermined character is itself a hexadecimal number, it may be arranged in order of the size of the numbers. For example, the name strings under the same predetermined category may be arranged in a small to large order according to the Unicode of the first predetermined character.
  • a first predetermined character such as "Chen Yi”, “Cheng Er”, “Cai San”, according to its first predetermined character “Cai”,
  • the characteristics of the order of "Chen” and “Cheng” are from small to large, and the name strings under the C category are arranged as “Cai San”, “Chen Yi” and “Cheng Er”, which realize the same category. , the sorting of the name string.
  • the method may further include the following steps.
  • Step S161 Acquire a second predetermined character in the name string.
  • Step S163 Arranging the name strings belonging to the same predetermined category according to the coding order of the Unicode of the second predetermined character.
  • the name string may include the same first predetermined character.
  • the second predetermined character may be a combination of one or more characters in the name string except the first predetermined character.
  • the order may be arranged according to the size of the numbers.
  • the name strings in the same predetermined category may be arranged in a small to large order according to the Unicode of the second predetermined character.
  • the name string when the name string is a name, it can represent a name containing multiple first names.
  • the name string under the W category, the name string includes: “Wang Yi”, “Wang Er”, “Wang San”.
  • the Unicode of the first predetermined character is obtained separately, it is found that the corresponding Unicode codes are the same.
  • a Unicode of the second character other than the first predetermined character may be further acquired.
  • the second predetermined characters that can be specifically obtained are the unified codes of “one”, “two”, and “three”, respectively.
  • the order of the Unicode of the second predetermined characters “one", “two", and “three” is “two", "three", and “one”.
  • the name string corresponding to the second predetermined character may be arranged as “Wang Er”, “Wang San”, “Wang Yi”, thereby realizing that when the first predetermined characters are the same, the same category Sort the name string below.
  • the following steps may be further included in the step S14 of displaying the category identifier and the name string.
  • Step S160 Acquire a unified code of the second predetermined character in the name string.
  • Step S162 Arranging the name strings belonging to the same predetermined category according to the coding order of the Unicode of the second predetermined character.
  • the second predetermined character may be a combination of one or more characters in the name string except the first predetermined character.
  • the encoding of the Unicode of the second predetermined character is itself a hexadecimal number, it may be arranged in order of the size of the numbers.
  • the Unicode of the second predetermined character is arranged in the order of small to large for the name strings under the same predetermined category.
  • the name string when the name string is a name, for example, under the C category, the name string includes: “Chen Yi”, “Cheng Er”, "Cai San”.
  • the first predetermined characters of the above name string are “Cai”, “Chen”, “Cheng”; the second predetermined characters are "One”, “Two”, “Three”.
  • the Unicode of the second predetermined character corresponding to "one", “two", and “three” is “two” from small to large. , "three", "one”, therefore, the corresponding name string under the C category is sorted as “Cheng Er", "Cai San", “Chen Yi”.
  • the step S10 of acquiring the name data may further include: receiving a first predetermined character specified by the user.
  • the first predetermined character specified by the user when the name data is acquired, the first predetermined character specified by the user may be received.
  • the specified first predetermined character may be one character or a combination of multiple characters in the name string.
  • the specified first predetermined character is used to sort the character strings belonging to the same predetermined category name after the name string is attributed to the predetermined category.
  • the specified first predetermined character may be a list of the last name of the name, may be a country list of the location, may be a bank category list of the account, etc., and the application is not specifically limited herein.
  • the last names in the last name list may be arranged in order.
  • the list of surnames includes “Brown”, “Smith”, “White”, etc., which are arranged in order of priority: “Brown”, "Smith”, “White”.
  • the name data includes a name string: "Alice Brown”, “Ada Smith”, “Alma White”.
  • the name string is arranged in the order of "Alice Brown”, “Ada Smith”, and "Alma White”.
  • the embodiment of the present application further provides a client 100, which may include: a data acquisition module 10, a Unicode acquisition module 12, a predetermined class division module 14, and a display module 16.
  • client 100 may include: a data acquisition module 10, a Unicode acquisition module 12, a predetermined class division module 14, and a display module 16.
  • the data obtaining module 10 is configured to obtain name data, where the name data includes at least one name string.
  • the client may be a communication device having a network communication function, such as a desktop computer, a notebook computer, a tablet computer, a smart phone, and a smart wearable device.
  • the client can also be software running in the above communication device.
  • the client can be used by the user, and can perform interactive activities such as instant messaging, and can perform various e-commerce activities such as online transaction and online electronic payment in the instant communication process.
  • the name data may be pre-stored in the client.
  • the pre-stored name data may be retrieved by the processor.
  • Name data It can also be obtained by downloading from the server side.
  • the name data can be downloaded from the server through the data communication module.
  • the name data may also be temporarily obtained by the user in an input manner.
  • the manner of obtaining the name data may be other methods, for example, accepting the name data sent by other clients, and the like, which is not specifically limited herein.
  • the name data may be a set of names that the user needs to use when performing interactive activities such as instant messaging and e-commerce.
  • the name may be specifically: name, nickname, company name, account name, location name, etc., and the application is not specifically limited herein.
  • the name data includes at least one name string.
  • the name string may be a specific name in the name, nickname, company name, account name, and place name.
  • the name string may be a specific name, such as “Wang Erxiao”.
  • the name string may also be a specific account name, such as a bank card number of a bank: "622848440637874213".
  • the form of the name string may be at least one word, word, phrase, or the like in a certain language, or may be a symbol, a number, a letter, a plurality of combinations, etc., of course, the name string may also be in other forms. This application is not specifically limited herein.
  • the Unicode obtaining module 12 is configured to obtain a Unicode of the first predetermined character in the name string.
  • the Unicode is used to set a uniform and unique encoding for characters in each language to meet the requirements of text conversion and processing across languages and platforms.
  • the Unicode has a character set that can be mapped with a predetermined number of characters such that characters in the character set of the Unicode have a one-to-one correspondence with characters in each language. That is, a Unicode unique corresponds to one character.
  • the predetermined number of digits may be hexadecimal, binary, etc., which are not stated here.
  • the Unicode code is a typical Unicode code that uses hexadecimal digits 0-0x10FFFF to map characters in various languages.
  • Chinese, Japanese, and Korean CJK Choinese Japanese Korean three characters occupy the 0x3000 to 0x9FFF part of Unicode.
  • the UCS-2 specification code is commonly used in Unicode. Specifically, it encodes one character with two bytes. For example, the encoding of the Chinese character "jing" is 0x7ECF. Since character encoding is generally expressed in hexadecimal, in order to distinguish from decimal, hexadecimal starts with 0x, 0x7ECF is converted to decimal with 32463, UCS-2 encodes characters with two bytes, and two bytes are 16 Bit binary, the 16th power of 2 is equal to 65536, so UCS-2 can encode up to 65536 characters.
  • a unified character set may be stored in the client, and the unified character set may include a character and a Unicode corresponding to the character.
  • the unified character set may also be stored on the server end, and the client may obtain related characters in the unified character set and a unified code corresponding to the characters after establishing communication with the server end through the communication module.
  • Each of the first predetermined characters may uniquely correspond to a unified code.
  • the first predetermined character may be the first character of the name string, or may be in the name string
  • the combination of the first character and the second character may also be a second character alone, or may be some predetermined character in the name string pre-specified by the user.
  • the first predetermined character is not limited to the above description. Other changes may be made by those skilled in the art in light of the technical spirit of the present application. However, as long as the functions and effects thereof are the same or similar to the present application, they should be covered by the present application.
  • the first predetermined character may uniquely have a unified code, according to the first predetermined character, by using a lookup manner or by a mapping relationship between the first predetermined character and the Unicode, A Unicode of the first predetermined character in the name string.
  • the first predetermined character can be converted into a hexadecimal number, and then the search is performed in the unified character set. If the same hexadecimal number in the unified character set is found, the Unicode of the first predetermined character can be uniquely determined.
  • the manner of obtaining is not limited to the above description.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • a predetermined class division module 14 configured to determine a predetermined code segment corresponding to the Unicode of the first predetermined character, and divide the name string into a predetermined category corresponding to the predetermined code segment; wherein, the predetermined code segment The number is at least one, and each of the predetermined code segments corresponds to a predetermined category.
  • the language ordering mechanism in the programming language API can be used to divide the Unicode of the first predetermined character according to the pronunciation, spelling and the like of the language. At least one predetermined code segment is determined. Each of the predetermined code segments may be correspondingly set to a predetermined category, or a plurality of predetermined code segments may be correspondingly set to the same predetermined category, thereby classifying the characters of the plurality of languages into a limited category.
  • unified coding is not arranged according to the pronunciation of the language.
  • a sorter such as a collator in Java, may be used to sort the characters according to the characteristics of the pronunciation of the language.
  • the Unicode corresponding to the sorted characters is divided into at least one predetermined code segment. Specifically, when the sorter collator performs sorting, the order of the two characters can be compared according to the characteristics of the language, and the characters in the unified character set can be segmented according to the language features.
  • the first predetermined character of all "ah” (a) pronunciations can be ranked in front of the first predetermined characters of all "bar” pronunciations according to their pronunciation.
  • the first predetermined character of all "bar” pronunciations is placed in front of the first predetermined character of all "can” (ca) pronunciations.
  • the first predetermined character in the name data is divided into "a” to "z” 26 predetermined code segments according to the pronunciation of the characters.
  • the predetermined category is a set of a plurality of name strings.
  • the predetermined category may be implemented in a data structure by a queue, a data stack, an array, or the like.
  • the predetermined category of data structure is When the data structure mode of the array is used, the first predetermined character pronunciation name strings belonging to the same predetermined code segment may be stored in the same array to form a predetermined category. For example, “Chen Yi”, “Cheng Er”, “Cai San”, and “Cheng Si” are stored under the same array to form a predetermined category.
  • the manner in which the data structure of the predetermined category is formed is not limited to the above description. Other changes may be made by those skilled in the art in light of the technical spirit of the present application. However, as long as the functions and effects thereof are the same or similar to the present application, they should be covered by the present application.
  • the collator in Java may be used to sequentially arrange the Unicode codes corresponding to the first predetermined characters according to the pronunciation, spelling and the like of the language, and is divided into At least one predetermined code segment.
  • the first predetermined character may be compared with a character corresponding to the Unicode of the segment header and/or the segment end of the predetermined code segment. If the order of the first predetermined character is after the first character of the predetermined code segment, but before the end of the paragraph character, or in the order before the first character of the next segment, it may be determined that the first predetermined character belongs to the predetermined Code segment.
  • a correspondence relationship may be formed between the predetermined code segment header and/or the segment end Unicode and the predetermined class. Please refer to Table 1 above.
  • the predetermined category of the name string may be determined according to the correspondence between the predetermined code segment and the predetermined category.
  • the collator in the existing Java is utilized, and the sorting according to the pronunciation, spelling, and the like of the language is implemented.
  • the specific ordering manner is not limited to the foregoing description, and those skilled in the art may also make other manners, such as selecting other existing sorters or writing according to the sorting function, etc., inspired by the technical essence of the present application.
  • the functions and effects achieved by the same or similar to the present application are within the scope of the present application.
  • a unified character set may be integrated in the operating system, that is, a Unicode corresponding to each character may be obtained from an operating system.
  • the segment header Unicode of the predetermined code segment may be stored in the client, so that when the Unicode segment is segmented, the segment header Unicode may be used to distinguish different predetermined code segments; or the predetermined code segment is a segment header Unicode and a segment end segment Unicode of the predetermined code segment.
  • the client stores a Unicode of 26 predetermined code segments, a total of 26; or a segment header Unicode and a segment end of the predetermined code segment
  • the Unicode code a total of 52. Therefore, the storage space required is relatively small compared to the client.
  • the unified coding is not arranged according to the pronunciation and pronunciation of each language.
  • the Unicode code corresponding to the first predetermined character may be corresponding to an index value according to the pronunciation and other characteristics of the language, and may be sorted according to the index value, and then the index value may be sorted.
  • the latter Unicode code is divided into at least one predetermined code segment.
  • a corresponding index value corresponding to the Unicode of the first predetermined character may be given according to the pronunciation thereof.
  • the index value can be viewed as a number. For example, “ah” (a), its index value may be 1; and “bar” (ba), its index value may be 10, then all "ah” (a) to "bar” (ba) between the pronunciation of Chinese characters Its index value is between 1-10.
  • the predetermined index corresponding to the first index value of the encoded segment is [1, 10) is the A category, that is, the word whose initial letter a is pronounced, such as "ah", "an”, etc., is assigned to the category A.
  • the predetermined index value corresponding to the first index value of the code segment segment [10, 20) is the B category, that is, the word whose initial letter b is pronounced, such as "bar", "being", etc., is assigned to the B category.
  • the Z category is compiled.
  • the entire Chinese code block is divided into 26 code segments, corresponding to 26 predetermined categories.
  • the first reservation may be determined by performing a lookup comparison between the index value of the Unicode of the first predetermined character and the segment header index value and/or the segment tail index value of the preset code segment.
  • the predetermined code segment corresponding to the character. Since the predetermined code segment corresponds to a predetermined category, the name string may be divided into a predetermined category corresponding to the predetermined code segment.
  • the index value of the corresponding Unicode code is compared with a predetermined code segment header index value and/or a segment tail index value, thereby determining the first predetermined character.
  • the corresponding predetermined code segment and the predetermined category Please refer to Table 2 for details. For example, if the index value corresponding to the Unicode of the first predetermined character is greater than or equal to 1 and less than 10, the predetermined code segment corresponding to the first predetermined character and the predetermined category may be determined, and the The name string corresponding to a predetermined character is divided into the A category of the predetermined code segment.
  • a unified character set may be integrated in the operating system, that is, the unified code corresponding to each first predetermined character may be obtained from an operating system.
  • the segment header index value of the predetermined code segment that can be stored only in the client, so that when the Unicode segment is segmented, different predetermined code segments are distinguished by the segment header index value and/or the segment tail index value.
  • the client stores an index value of 26 predetermined code segments, for a total of 26; or a segment header index value and a segment tail of a predetermined code segment.
  • the index value is 52 in total. Therefore, the storage space required is relatively small compared to the client.
  • the manner of determining the predetermined code segment corresponding to the Unicode of the first predetermined character is not limited to the above description.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • the display module 16 is configured to display, in the display interface, a category identifier of the predetermined category, and a name string corresponding to the predetermined category.
  • the predetermined category may correspond to a category identifier, and when displayed in the display interface, Displaying the category identifier of the predetermined category together with the name string corresponding to the predetermined category to form a list of names that are convenient for the user to find.
  • the category identifier may be an identifier that is convenient for the user to identify according to the user's language habits. For example, for a user who uses Chinese, the category identifier may be the first letter of the first predetermined character pronunciation. For a user who uses English, the category identifier may be the first letter of the first predetermined character. Of course, for other language types, the category identifier may also be the phonetic symbol of the first predetermined character or the initials of the spelling, etc., and the examples are not repeated here.
  • the name string corresponding to the predetermined category may all display or display part of the name string or not display the name string.
  • the predetermined category includes a plurality of name strings.
  • a predetermined number of name strings can be displayed under the category mark corresponding to the predetermined category.
  • a maximum of three name strings can be set.
  • the client in the embodiment of the present application divides the name string into a predetermined category corresponding to the predetermined code segment based on a Unicode of the first predetermined character in the name string, and correspondingly sets a category identifier to the predetermined category.
  • Correspondingly displaying the category identifier of the predetermined category and the name string corresponding to the predetermined category in the display interface realizing the one-to-one correspondence between the Unicode and the name string, and the languages of the plurality of countries are in the same name Classified display in the list.
  • the user can quickly locate the specific name string in the name list by displaying the category identifier, and find the corresponding target name.
  • the traditional default dictionary is removed, and the existing Unicode code is used to sort and sort, which reduces a lot of manpower and saves human resource costs.
  • an embodiment of the present application further provides a client 110 , which may include: a processor 11 and a display 13 .
  • the processor 11 is configured to obtain name data, where the name data includes at least one name string; and is used to obtain a unified code of a first predetermined character in the name string; and determining the first predetermined character a predetermined code segment corresponding to the Unicode, dividing the name string into a predetermined category corresponding to the predetermined code segment; wherein the number of the predetermined code segments is at least one, and each of the predetermined code segments corresponds to a predetermined category Controlling the display 13 correspondingly displaying the category identifier of the predetermined category in the display interface, and a name string corresponding to the predetermined category.
  • processor 11 is implemented in any suitable manner.
  • processor 11 is a computer readable medium, logic gate, switch, application specific integrated circuit that employs, for example, a microprocessor or processor and computer readable program code (eg, software or firmware) executable by the (micro)processor.
  • ASIC Application Specific Integrated Circuit
  • programmable logic controller and embedded microcontroller form etc. This application is not Limited.
  • the client of the present application can be a hardware implementation manner of the method for displaying the name list of the present application, and can implement the implementation method of the name list display method of the present application and achieve the technical effect of the method implementation manner.
  • the name list processing method in the embodiment of the present invention may be executed by a server.
  • an embodiment of the present application provides a method for processing a name list, which may include the following steps.
  • Step S20 Acquire name data; wherein the name data includes at least one name string.
  • the server may be a device having data processing capabilities and network communication functions, such as a personal computer, a server computer, a handheld device or a portable device, a tablet device, a multi-processor device, and a distribution including any of the above devices or devices. Computing environment.
  • the server can also include software running on the above devices.
  • the name data may be pre-stored in the server.
  • the name data can be downloaded from the server through the data communication module.
  • the name data may also be temporarily obtained by the user in an input manner.
  • the manner of obtaining the name data may be other methods, for example, accepting the name data sent by other clients, and the like, which is not specifically limited herein.
  • the name data may be a set of names that the user needs to use when performing interactive activities such as instant messaging and e-commerce.
  • the name may be specifically: name, nickname, company name, account name, location name, etc., and the application is not specifically limited herein.
  • the name data includes at least one name string.
  • the name string may be a specific name in the name, nickname, company name, account name, and place name.
  • the name string may be a specific name, such as “Wang Erxiao”.
  • the name string may also be a specific account name, such as a bank card number of a bank: "622848440637874213".
  • the form of the name string may be at least one word, word, phrase, or the like in a certain language, or may be a symbol, a number, a letter, a plurality of combinations, etc., of course, the name string may also be in other forms. This application is not specifically limited herein.
  • Step S22 Acquire a unified code of the first predetermined character in the name string.
  • the Unicode is used to set a uniform and unique encoding for characters in each language to meet the requirements of text conversion and processing across languages and platforms.
  • the Unicode has a character set that can be mapped with a predetermined number of characters such that characters in the character set of the Unicode have a one-to-one correspondence with characters in each language. That is, a Unicode unique corresponds to one character.
  • the predetermined number of digits may be hexadecimal, binary, etc., which are not stated here.
  • the Unicode code is a typical Unicode code that uses hexadecimal digits 0-0x10FFFF to map characters in various languages.
  • Chinese, Japanese, and Korean CJK Choinese Japanese Korean three characters occupy the 0x3000 to 0x9FFF part of Unicode.
  • the UCS-2 specification code is commonly used in Unicode. Specifically, it encodes one character with two bytes. For example, the encoding of the Chinese character "jing" is 0x7ECF. Since character encoding is generally expressed in hexadecimal, in order to distinguish from decimal, hexadecimal starts with 0x, 0x7ECF is converted to decimal with 32463, UCS-2 encodes characters with two bytes, and two bytes are 16 Bit binary, the 16th power of 2 is equal to 65536, so UCS-2 can encode up to 65536 characters.
  • a unified character set may be stored in the client, and the unified character set may include a character and a Unicode corresponding to the character.
  • the unified character set may also be stored on the server end, and the client may obtain related characters in the unified character set and a unified code corresponding to the characters after establishing communication with the server end through the communication module.
  • Each of the first predetermined characters may uniquely correspond to a unified code.
  • the first predetermined character may be a first character of the name string, or may be a combination of a first character and a second character in the name string, or may be a second character alone, or may also be Is a predetermined character in the name string pre-specified by the user.
  • the first predetermined character is not limited to the above description.
  • Other changes may be made by those skilled in the art in light of the technical spirit of the present application.
  • the functions and effects thereof are the same or similar to the present application, they should be covered by the present application.
  • the first predetermined character may uniquely have a unified code, according to the first predetermined character, by using a lookup manner or by a mapping relationship between the first predetermined character and the Unicode, A Unicode of the first predetermined character in the name string.
  • the first predetermined character can be converted into a hexadecimal number, and then the search is performed in the unified character set. If the same hexadecimal number in the unified character set is found, the Unicode of the first predetermined character can be uniquely determined.
  • the manner of obtaining is not limited to the above description.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • Step S24 determining a predetermined code segment corresponding to the Unicode of the first predetermined character, and dividing the name string into a predetermined category corresponding to the predetermined code segment; wherein the number of the predetermined code segments is at least one, Each of the predetermined code segments corresponds to a predetermined category.
  • the language ordering mechanism in the programming language API can be used to divide the Unicode of the first predetermined character according to the pronunciation, spelling and the like of the language. At least one predetermined code segment is determined. Each of the predetermined code segments may be correspondingly set to a predetermined category, or a plurality of predetermined code segments may be correspondingly set to the same predetermined category, thereby classifying the characters of the plurality of languages into a limited category.
  • unified coding is not arranged according to the pronunciation of the language.
  • the segmentation is in accordance with the user's language habits, and can be sorted by using a sorter, such as a collator in Java, according to the characteristics of the pronunciation of the language.
  • the Unicode corresponding to the sorted characters is divided into at least one predetermined code segment. Specifically, when the sorter collator performs sorting, the order of the two characters can be compared according to the characteristics of the language, and the characters in the unified character set can be segmented according to the language features.
  • the first predetermined character of all "ah” (a) pronunciations can be ranked in front of the first predetermined characters of all "bar” pronunciations according to their pronunciation.
  • the first predetermined character of all "bar” pronunciations is placed in front of the first predetermined character of all "can” (ca) pronunciations.
  • the first predetermined character in the name data is divided into "a” to "z” 26 predetermined code segments according to the pronunciation of the characters.
  • the predetermined category is a set of a plurality of name strings.
  • the predetermined category may be implemented in a data structure by a queue, a data stack, an array, or the like.
  • the name strings of the first predetermined character pronunciation belonging to the same predetermined code segment may be stored under the same array to form a predetermined category.
  • "Chen Yi”, “Cheng Er”, “Cai San”, and “Cheng Si” are stored under the same array to form a predetermined category.
  • the manner in which the data structure of the predetermined category is formed is not limited to the above description. Other changes may be made by those skilled in the art in light of the technical spirit of the present application. However, as long as the functions and effects thereof are the same or similar to the present application, they should be covered by the present application.
  • the collator in Java may be used to sequentially arrange the Unicode codes corresponding to the first predetermined characters according to the pronunciation, spelling and the like of the language, and is divided into At least one predetermined code segment.
  • the first predetermined character may be compared with a character corresponding to the Unicode of the segment header and/or the segment end of the predetermined code segment. If the order of the first predetermined character is after the first character of the predetermined code segment, but before the end of the paragraph character, or in the order before the first character of the next segment, it may be determined that the first predetermined character belongs to the predetermined Code segment.
  • a correspondence relationship may be formed between the predetermined code segment header and/or the segment end Unicode and the predetermined class. Please refer to Table 1 above.
  • the predetermined category of the name string may be determined according to the correspondence between the predetermined code segment and the predetermined category.
  • the collator in the existing Java is utilized, and the sorting according to the pronunciation, spelling, and the like of the language is implemented.
  • the specific ordering manner is not limited to the foregoing description, and those skilled in the art may also make other manners, such as selecting other existing sorters or writing according to the sorting function, etc., inspired by the technical essence of the present application.
  • the functions and effects achieved by the same or similar to the present application are within the scope of the present application.
  • a unified character set may be integrated in the operating system, that is, a Unicode corresponding to each character may be used. Obtained from the operating system.
  • the segment header Unicode of the predetermined code segment may be stored in the client, so that when the Unicode segment is segmented, the segment header Unicode may be used to distinguish different predetermined code segments; or the predetermined code segment is a segment header Unicode and a segment end segment Unicode of the predetermined code segment.
  • the client stores a Unicode of 26 predetermined code segments, a total of 26; or a segment header Unicode and a segment end of the predetermined code segment
  • the Unicode code a total of 52. Therefore, the storage space required is relatively small compared to the client.
  • the unified coding is not arranged according to the pronunciation and pronunciation of each language.
  • the Unicode code corresponding to the first predetermined character may be corresponding to an index value according to the pronunciation and other characteristics of the language, and may be sorted according to the index value, and then the index value may be sorted.
  • the latter Unicode code is divided into at least one predetermined code segment.
  • a corresponding index value corresponding to the Unicode of the first predetermined character may be given according to the pronunciation thereof.
  • the index value can be viewed as a number. For example, “ah” (a), its index value may be 1; and “bar” (ba), its index value may be 10, then all "ah” (a) to "bar” (ba) between the pronunciation of Chinese characters Its index value is between 1-10.
  • the predetermined index corresponding to the first index value of the encoded segment is [1, 10) is the A category, that is, the word whose initial letter a is pronounced, such as "ah", "an”, etc., is assigned to the category A.
  • the predetermined index value corresponding to the first index value of the code segment segment [10, 20) is the B category, that is, the word whose initial letter b is pronounced, such as "bar", "being", etc., is assigned to the B category.
  • the Z category is compiled.
  • the entire Chinese code block is divided into 26 code segments, corresponding to 26 predetermined categories.
  • the first reservation may be determined by performing a lookup comparison between the index value of the Unicode of the first predetermined character and the segment header index value and/or the segment tail index value of the preset code segment.
  • the predetermined code segment corresponding to the character. Since the predetermined code segment corresponds to a predetermined category, the name string may be divided into a predetermined category corresponding to the predetermined code segment.
  • the index value of the corresponding Unicode code is compared with a predetermined code segment header index value and/or a segment tail index value, thereby determining the first predetermined character.
  • the corresponding predetermined code segment and the predetermined category Please refer to Table 2 for details. For example, if the index value corresponding to the Unicode of the first predetermined character is greater than or equal to 1 and less than 10, the predetermined code segment corresponding to the first predetermined character and the predetermined category may be determined, and the The name string corresponding to a predetermined character is divided into the A category of the predetermined code segment.
  • a unified character set may be integrated in the operating system, that is, the unified code corresponding to each first predetermined character may be obtained from an operating system.
  • the segment header index value of the predetermined code segment that can be stored only in the client, so that when the Unicode segment is segmented, the segmentation index value and/or the segment tail index value are used to distinguish different reservations.
  • Code segment Specifically, for example, when the predetermined code segment is 26, correspondingly, the client stores an index value of 26 predetermined code segments, for a total of 26; or a segment header index value and a segment tail of a predetermined code segment. The index value is 52 in total. Therefore, the storage space required is relatively small compared to the client.
  • the manner of determining the predetermined code segment corresponding to the Unicode of the first predetermined character is not limited to the above description.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • the name list processing method in the embodiment of the present application is based on the Unicode of the first predetermined character in the name string, and divides the name string into a predetermined category corresponding to the predetermined code segment, thereby realizing the use of the Unicode and the name string.
  • a one-to-one correspondence that classifies languages in multiple countries in the same name list.
  • the category identifier of the predetermined category and the name string corresponding to the predetermined category may be sent to the client.
  • the user can quickly locate the specific name string in the name list by displaying the category identifier, and find the corresponding target name.
  • the name list display method is freed from the traditional preset dictionary in the process of implementation, and uses the existing unified code to sort and sort, which reduces a large amount of human labor and saves human resource costs.
  • the Unicode stored on the server side serves as a basis for encoding and decoding the text data when flowing as data across regions or data platforms.
  • the Unicode by sorting and sorting the name list by using the Unicode, it is possible to implement the sorting of the name list by way of not implementing the additional dictionary introduced in the existing client. Specifically, for example, a pinyin4j dictionary may not be introduced, which saves storage space.
  • the name list processing method may further include: sending the category identifier of the predetermined category, and the name string corresponding to the predetermined category to the client for corresponding display.
  • the predetermined category may correspond to a category identifier.
  • Communication can be established between the server and the client. For example, when the server receives the sending request of the client, or after the server automatically updates, the server may send the category identifier of the predetermined category and the name string corresponding to the predetermined category to The corresponding client. After receiving the content, the client may display the category identifier of the predetermined category together with the name string corresponding to the predetermined category to form a name list that is convenient for the user to search.
  • the category identifier may be an identifier that is convenient for the user to identify according to the user's language habits. For example, for a user who uses Chinese, the category identifier may be the first letter of the first predetermined character pronunciation. For a user who uses English, the category identifier may be the first letter of the first predetermined character. Of course, for other language types, the category identifier may also be the phonetic symbol of the first predetermined character or the initials of the spelling, etc., and the examples are not repeated here.
  • the name string corresponding to the predetermined category may all display or display part of the name string or not display the name string.
  • the predetermined category includes a plurality of name strings, In this case, a predetermined number of name strings may be displayed under the category mark corresponding to the predetermined category. For example, a maximum of three name strings may be set.
  • the specific content of the name data may include at least one of a person name, a company name, an account name, and a place name.
  • the name data may include at least one name string.
  • the name string may be a person name, a company name, an account name, a place name, and the like, and the present application is not specifically limited herein.
  • the name of the person may be the real name of the user, or may be a nickname set by the user.
  • the name data may be of a type including only one name, for example, only a person name; or a plurality of name categories, for example, including a person name, a company name, an account name, and the like.
  • the name data includes a plurality of name types at the same time, if there is a correspondence between the plurality of name types, the plurality of name types may be divided, for example, a name of one of the plurality of name types may be set.
  • the string is the primary name, and other names in the multiple name categories are set as the dependent names.
  • only the primary name may be displayed in the display interface, and other dependent names corresponding to the primary name may be stored corresponding to the primary name.
  • the slave name can be displayed when the location where the master name is located detects an electrical signal triggered by the click.
  • operations such as copying, editing, and the like may be performed on the subordinate name.
  • the specific content of the name data is “Company A Xiaowang China Bank Account XX”, the name of the person is set as the master name, and the names of other companies and accounts are subordinate names.
  • the main name “Xiaowang” is displayed in the display interface, and other subordinate names “Company” and “Bank of China Account XX” corresponding to “Xiaowang” may be “small” with the main name.
  • Wang corresponds to storage.
  • the subordinate names "Company A” and “Bank of China Account XX” may be displayed correspondingly. Further, the subordinate names may be copied, edited, and the like.
  • the first predetermined character may be the first character of the name string.
  • the name string may be a person name, a company name, an account name, a place name, and the like.
  • the name string can include one or more characters.
  • the first character of the name string can typically represent the entire name string in many cases. For example, when the name string is a person name, the first character thereof may be a last name, and when the name string is a company name, the first character may also be the first word of the company name.
  • the first predetermined character is set as the first character of the name string, the subsequent list of names formed by the name string conforms to the usage habit of the user, and is helpful for the user to perform query based on the formed name list. .
  • the Unicode pair of the first predetermined character can be determined by a collator in Java The predetermined code segment should be.
  • a sorter such as a collator in Java
  • the Unicode corresponding to each character after sorting is divided into at least one predetermined code segment. Specifically, when the sorter collator performs sorting, the order of the two characters can be compared according to the characteristics of the language, and the characters in the unified character set can be segmented according to the language features.
  • the first predetermined character of all "ah” (a) pronunciations can be ranked in front of the first predetermined characters of all "bar” pronunciations according to their pronunciation.
  • the first predetermined character of all "bar” pronunciations is placed in front of the first predetermined character of all "can” (ca) pronunciations.
  • the Chinese characters in the name data are divided into “a” to "z” 26 predetermined code segments according to the pronunciation of the first predetermined character.
  • the collator in the existing Java is utilized, and the sorting according to the pronunciation, spelling, and the like of the language is implemented.
  • the specific ordering manner is not limited to the foregoing description, and those skilled in the art may also make other manners, such as selecting other existing sorters or writing according to the sorting function, etc., inspired by the technical essence of the present application.
  • the functions and effects achieved by the same or similar to the present application are within the scope of the present application.
  • a name string is newly added, and when it is determined that the first predetermined character of the newly added name string specifically belongs to a corresponding predetermined code segment, the first The predetermined character is compared with a character corresponding to the Unicode of the beginning and/or the end of the predetermined code segment. If the order of the first predetermined characters is in an order after the first character of the predetermined code segment but before the end of the paragraph character, or an order before the first character of the next segment, it may be determined that the first predetermined character belongs to the Predetermine the code segment.
  • At least one of said predetermined code segments constitutes a language coded block; wherein different said language coded blocks have at least partially identical predetermined categories.
  • the step of dividing the predetermined code into the predetermined categories includes: dividing the name string into the same when the predetermined categories corresponding to the predetermined code segments in which the name strings belonging to the different language coded blocks are located are the same A preset data set of a predetermined category.
  • different language code blocks can be set for different language types in the Unicode.
  • it may include a Chinese-Japanese Korean code block, an English code block, a Russian code block, an Arabic code block, and the like.
  • different language coding blocks may have at least partially identical predetermined categories.
  • the collator in Java can be used to set a plurality of predetermined code segments according to the pronunciation and the like of the language.
  • the Unicode corresponding to the first predetermined character may be divided into "a" to "z” 26 predetermined code segments according to the pronunciation characteristics of the first predetermined character.
  • the "a" to "z” 26 predetermined code segments respectively correspond to predetermined categories of A to Z26 categories.
  • the collator in Java can set multiple predetermined coding segments according to the character spelling characteristics.
  • a plurality of predetermined code segments may be set according to the order in which the Unicode codes are themselves arranged.
  • the first predetermined character may be divided into A to Z26 predetermined code segments according to a spelling characteristic of the first predetermined character.
  • the predetermined category corresponding to the first predetermined code segment first Unicode 97 is the category A, and the name string whose initial letter is A is assigned to the predetermined category A.
  • the predetermined category corresponding to the first predetermined code segment first Unicode 98 is the B category, and the name string whose initial letter B is assigned to the predetermined category B.
  • a predetermined category Z is prepared.
  • the entire English coded block is divided into 26 code segments, corresponding to 26 predetermined categories.
  • the name string belonging to the same predetermined category may be divided into the predetermined category to form the Preset the data set to achieve mixed sorting between Chinese and English.
  • the segment header or the segment end of each corresponding predetermined code segment may be set with an index value, when the English code block When the corresponding index value is the same as the index value corresponding to the Chinese coded block, it may be determined that the name string with the same index value belongs to the same category.
  • the Chinese coded block and the English coded block may not be divided into the same predetermined category, that is, each is divided according to a respective predetermined category, and is not mixed. row.
  • the English coding block and the noon coding block are classified and arranged according to the order of the Unicode, and may be classified and arranged in the order of the English coding block and the Chinese coding block as a whole.
  • the labels of the specific classifications of the different coding blocks may be the same or different, and the application is not specifically limited herein.
  • the type of the language coding block is not limited to the example Chinese language coding block and the English language coding block, and the number is not limited to the two language coding blocks.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • the following steps may be included in the step S24 of determining a predetermined code segment.
  • Step S240 Determine a language coding block corresponding to the Unicode of the first predetermined character in the name string.
  • Step S242 In the language coding block, determining a predetermined code segment in which the Unicode of the first predetermined character is located.
  • the corresponding language coded block when determining the predetermined code segment, may be first determined according to the Unicode of the first predetermined character in the name string. Based on the determined language coding block, further Determining a predetermined code segment in which the Unicode of the first predetermined character is located.
  • different language code blocks may be set in the Unicode for different language types.
  • the Unicode of the first character of each language code block and the number of characters are known.
  • the Unicode of the first predetermined character may be compared with the first character Unicode of each language coded block.
  • the Unicode of the first predetermined character is greater than or equal to the first character Unicode of a certain linguistic coding block, and the difference between the Unicode of the first linguistic coding block and the Unicode code of the linguistic coding block is less than the number of the linguistic coding blocks, Then, it can be determined that the first predetermined character is a character in the language encoding block.
  • the Unicode of the first predetermined character may be searched in the determined language encoding block to quickly determine A predetermined code segment in which the Unicode of the first predetermined character is located.
  • the determined language coding block generally includes a plurality of language code segments, and each language code segment corresponds to at least one unified code.
  • the Unicode of the first predetermined character may be compared with the Unicode of the language code segment, and if the same, the uniformity of the first predetermined character may be determined. The predetermined code segment in which the code is located.
  • the English language coding block may include 26 language code segments.
  • each predetermined code segment corresponds to a Unicode code.
  • the Unicode of the predetermined character may be compared with the Unicode of each predetermined code segment, and if the same, the A predetermined code segment in which a Unicode of a predetermined character is located.
  • the method may further include: dividing the name string into a preset specified category when there is no language encoding block corresponding to the Unicode of the first predetermined character in the name string.
  • a preset specified category may be set in advance on the client for categorizing the name string in which the corresponding language coded block is not found.
  • the preset specified category may be a single category arranged after the predetermined category, which may be displayed with a special tag. For example, "#" can be used to indicate that the following includes a preset specified category that does not belong to the language coded block included in the Unicode.
  • the client may further include: in the step of displaying the category identifier and the name string:
  • the name strings belonging to the same predetermined category are arranged in the coding order of the Unicode of the first predetermined character.
  • the first predetermined character of the name string may correspond to a unique Unicode.
  • the name strings belonging to the same predetermined category are sorted, they may be arranged in the coding order of the Unicode of the first predetermined characters.
  • the order may be arranged according to the size of the numbers.
  • the name strings under the same predetermined category may be arranged in a small to large order according to the Unicode of the first predetermined character.
  • a first predetermined character such as "Chen Yi”, “Cheng Er”, “Cai San”, according to its first predetermined character “Cai”,
  • the characteristics of the order of "Chen” and “Cheng” are from small to large, and the name strings under the C category are arranged as “Cai San”, “Chen Yi” and “Cheng Er”, which realize the same category. , the sorting of the name string.
  • the method may further include the following steps.
  • Step S261 Acquire a second predetermined character in the name string.
  • Step S263 Arranging the name strings belonging to the same predetermined category according to the coding order of the Unicode of the second predetermined character.
  • the name string may include the same first predetermined character.
  • the second predetermined character may be a combination of one or more characters in the name string except the first predetermined character.
  • the encoding of the Unicode of the second predetermined character is itself a hexadecimal number, it may be arranged in order of the size of the numbers.
  • the name strings in the same predetermined category may be arranged in a small to large order according to the Unicode of the second predetermined character.
  • the name string when the name string is a name, it can represent a name containing multiple first names.
  • the name string under the W category, the name string includes: “Wang Yi”, “Wang Er”, “Wang San”.
  • the Unicode of the first predetermined character is obtained separately, it is found that the corresponding Unicode codes are the same.
  • a Unicode of the second character other than the first predetermined character may be further acquired.
  • the second predetermined characters that can be specifically obtained are the unified codes of “one”, “two”, and “three”, respectively.
  • the order of the Unicode of the second predetermined characters “one", “two", and “three” is “two", "three", and “one”.
  • the name string corresponding to the second predetermined character may be arranged as “Wang Er”, “Wang San”, “Wang Yi”, thereby realizing that when the first predetermined characters are the same, the same category Sort the name string below.
  • the following steps may be further included in the step S26 of displaying the category identifier and the name string.
  • Step S260 Acquire a unified code of the second predetermined character in the name string.
  • Step S262 Arranging the name strings belonging to the same predetermined category according to the coding order of the Unicode of the second predetermined character.
  • the second predetermined character may be in the name string, except for the first predetermined A character or combination of characters other than a character.
  • the encoding of the Unicode of the second predetermined character is itself a hexadecimal number, it may be arranged in order of the size of the numbers.
  • the name strings in the same predetermined category may be arranged in a small to large order according to the Unicode of the second predetermined character.
  • the name string when the name string is a name, for example, under the C category, the name string includes: “Chen Yi”, “Cheng Er”, "Cai San”.
  • the first predetermined characters of the above name string are “Cai”, “Chen”, “Cheng”; the second predetermined characters are "One”, “Two”, “Three”.
  • the Unicode of the second predetermined character corresponding to "one", “two", and “three” is “two” from small to large. , "three", "one”, therefore, the corresponding name string under the C category is sorted as “Cheng Er", "Cai San", “Chen Yi”.
  • the step S20 of obtaining the name data may further include: receiving a first predetermined character specified by the user.
  • the first predetermined character specified by the user when the name data is acquired, the first predetermined character specified by the user may be received.
  • the specified first predetermined character may be one character or a combination of multiple characters in the name string.
  • the specified first predetermined character is used to sort the character strings belonging to the same predetermined category name after the name string is attributed to the predetermined category.
  • the specified first predetermined character may be a list of the last name of the name, may be a country list of the location, may be a bank category list of the account, etc., and the application is not specifically limited herein.
  • the last names in the last name list may be arranged in order.
  • the list of surnames includes “Brown”, “Smith”, “White”, etc., which are arranged in order of priority: “Brown”, "Smith”, “White”.
  • the name data includes a name string: "Alice Brown”, “Ada Smith”, “Alma White”.
  • the name string is arranged in the order of "Alice Brown”, “Ada Smith”, and "Alma White”.
  • an embodiment of the present application further provides a server 200, which may include: a data acquisition module 20, a Unicode acquisition module 22, and a predetermined category division module 24.
  • the data obtaining module 10 is configured to obtain name data, where the name data includes at least one name string.
  • the server may be a device having data processing capabilities and network communication functions, such as a personal computer, a server computer, a handheld device or a portable device, a tablet device, a multi-processor device, and a distribution including any of the above devices or devices. Computing environment.
  • the server can also include software running on the above devices.
  • the name data may be pre-stored in the server.
  • the name data can be downloaded from the server through the data communication module.
  • the name data may also be temporarily obtained by the user in an input manner.
  • the manner of obtaining the name data may be other methods, for example, accepting the name data sent by other clients, and the like, which is not specifically limited herein.
  • the name data may be a set of names that the user needs to use when performing interactive activities such as instant messaging and e-commerce.
  • the name may be specifically: name, nickname, company name, account name, location name, etc., and the application is not specifically limited herein.
  • the name data includes at least one name string.
  • the name string may be a specific name in the name, nickname, company name, account name, and place name.
  • the name string may be a specific name, such as “Wang Erxiao”.
  • the name string may also be a specific account name, such as a bank card number of a bank: "622848440637874213".
  • the form of the name string may be at least one word, word, phrase, or the like in a certain language, or may be a symbol, a number, a letter, a plurality of combinations, etc., of course, the name string may also be in other forms. This application is not specifically limited herein.
  • the Unicode obtaining module 12 is configured to obtain a Unicode of the first predetermined character in the name string.
  • the Unicode is used to set a uniform and unique encoding for characters in each language to meet the requirements of text conversion and processing across languages and platforms.
  • the Unicode has a character set that can be mapped with a predetermined number of characters such that characters in the character set of the Unicode have a one-to-one correspondence with characters in each language. That is, a Unicode unique corresponds to one character.
  • the predetermined number of digits may be hexadecimal, binary, etc., which are not stated here.
  • the Unicode code is a typical Unicode code that uses hexadecimal digits 0-0x10FFFF to map characters in various languages.
  • Chinese, Japanese, and Korean CJK Choinese Japanese Korean three characters occupy the 0x3000 to 0x9FFF part of Unicode.
  • the UCS-2 specification code is commonly used in Unicode. Specifically, it encodes one character with two bytes. For example, the encoding of the Chinese character "jing" is 0x7ECF. Since character encoding is generally expressed in hexadecimal, in order to distinguish from decimal, hexadecimal starts with 0x, 0x7ECF is converted to decimal with 32463, UCS-2 encodes characters with two bytes, and two bytes are 16 Bit binary, the 16th power of 2 is equal to 65536, so UCS-2 can encode up to 65536 characters.
  • a unified character set may be stored in the client, and the unified character set may include a character and a Unicode corresponding to the character.
  • the unified character set may also be stored on the server end, and the client may obtain related characters in the unified character set and a unified code corresponding to the characters after establishing communication with the server end through the communication module.
  • Each of the first predetermined characters may uniquely correspond to a unified code.
  • the first predetermined character may be the first character of the name string, or may be in the name string
  • the combination of the first character and the second character may also be a second character alone, or may be some predetermined character in the name string pre-specified by the user.
  • the first predetermined character is not limited to the above description. Other changes may be made by those skilled in the art in light of the technical spirit of the present application. However, as long as the functions and effects thereof are the same or similar to the present application, they should be covered by the present application.
  • the first predetermined character may uniquely have a unified code, according to the first predetermined character, by using a lookup manner or by a mapping relationship between the first predetermined character and the Unicode, A Unicode of the first predetermined character in the name string.
  • the first predetermined character can be converted into a hexadecimal number, and then the search is performed in the unified character set. If the same hexadecimal number in the unified character set is found, the Unicode of the first predetermined character can be uniquely determined.
  • the manner of obtaining is not limited to the above description.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • the predetermined category division module 14 is configured to: in the embodiment, use a language sorting mechanism in a programming language API (Application Programming Interface) to set the first reservation according to features such as pronunciation, spelling, and the like of the language.
  • the Unicode of the character is divided to set at least one predetermined code segment.
  • Each of the predetermined code segments may be correspondingly set to a predetermined category, or a plurality of predetermined code segments may be correspondingly set to the same predetermined category, thereby classifying the characters of the plurality of languages into a limited category.
  • unified coding is not arranged according to the pronunciation of the language.
  • a sorter such as a collator in Java, may be used to sort the characters according to the characteristics of the pronunciation of the language.
  • the Unicode corresponding to the sorted characters is divided into at least one predetermined code segment. Specifically, when the sorter collator performs sorting, the order of the two characters can be compared according to the characteristics of the language, and the characters in the unified character set can be segmented according to the language features.
  • the first predetermined character of all "ah” (a) pronunciations can be ranked in front of the first predetermined characters of all "bar” pronunciations according to their pronunciation.
  • the first predetermined character of all "bar” pronunciations is placed in front of the first predetermined character of all "can” (ca) pronunciations.
  • the first predetermined character in the name data is divided into "a” to "z” 26 predetermined code segments according to the pronunciation of the characters.
  • the predetermined category is a set of a plurality of name strings.
  • the predetermined category may be implemented in a data structure by a queue, a data stack, an array, or the like.
  • the data structure mode of the predetermined category is the data structure mode of the array
  • the name strings of the first predetermined character pronunciation belonging to the same predetermined code segment may be stored under the same array to form a predetermined category.
  • "Chen Yi”, “Cheng Er”, “Cai San”, and “Cheng Si” are stored under the same array to form a predetermined category.
  • forming a data knot of the predetermined category The configuration is not limited to the above description. Other changes may be made by those skilled in the art in light of the technical spirit of the present application. However, as long as the functions and effects thereof are the same or similar to the present application, they should be covered by the present application.
  • the collator in Java may be used to sequentially arrange the Unicode codes corresponding to the first predetermined characters according to the pronunciation, spelling and the like of the language, and is divided into At least one predetermined code segment.
  • the first predetermined character may be compared with a character corresponding to the Unicode of the segment header and/or the segment end of the predetermined code segment. If the order of the first predetermined character is after the first character of the predetermined code segment, but before the end of the paragraph character, or in the order before the first character of the next segment, it may be determined that the first predetermined character belongs to the predetermined Code segment.
  • a correspondence relationship may be formed between the predetermined code segment header and/or the segment end Unicode and the predetermined class. Please refer to Table 1 above.
  • the predetermined category of the name string may be determined according to the correspondence between the predetermined code segment and the predetermined category.
  • the collator in the existing Java is utilized, and the sorting according to the pronunciation, spelling, and the like of the language is implemented.
  • the specific ordering manner is not limited to the foregoing description, and those skilled in the art may also make other manners, such as selecting other existing sorters or writing according to the sorting function, etc., inspired by the technical essence of the present application.
  • the functions and effects achieved by the same or similar to the present application are within the scope of the present application.
  • a unified character set may be integrated in the operating system, that is, a Unicode corresponding to each character may be obtained from an operating system.
  • the segment header Unicode of the predetermined code segment may be stored in the client, so that when the Unicode segment is segmented, the segment header Unicode may be used to distinguish different predetermined code segments; or the predetermined code segment is a segment header Unicode and a segment end segment Unicode of the predetermined code segment.
  • the client stores a Unicode of 26 predetermined code segments, a total of 26; or a segment header Unicode and a segment end of the predetermined code segment
  • the Unicode code a total of 52. Therefore, the storage space required is relatively small compared to the client.
  • the unified coding is not arranged according to the pronunciation and pronunciation of each language.
  • the Unicode code corresponding to the first predetermined character may be corresponding to an index value according to the pronunciation and other characteristics of the language, and may be sorted according to the index value, and then the index value may be sorted.
  • the latter Unicode code is divided into at least one predetermined code segment.
  • a corresponding index value corresponding to the Unicode of the first predetermined character may be given according to the pronunciation thereof.
  • the index value can be viewed as a number. For example, “ah” (a), its index value may be 1; and “bar” (ba), its index value may be 10, then all "ah” (a) to "bar” (ba)
  • the Chinese character is pronounced between its index value between 1-10.
  • the predetermined index corresponding to the first index value of the encoded segment is [1, 10) is the A category, that is, the word whose initial letter a is pronounced, such as "ah", "an”, etc., is assigned to the category A.
  • the predetermined index value corresponding to the first index value of the code segment segment [10, 20) is the B category, that is, the word whose initial letter b is pronounced, such as "bar", "being", etc., is assigned to the B category.
  • the Z category is compiled.
  • the entire Chinese code block is divided into 26 code segments, corresponding to 26 predetermined categories.
  • the first reservation may be determined by performing a lookup comparison between the index value of the Unicode of the first predetermined character and the segment header index value and/or the segment tail index value of the preset code segment.
  • the predetermined code segment corresponding to the character. Since the predetermined code segment corresponds to a predetermined category, the name string may be divided into a predetermined category corresponding to the predetermined code segment.
  • the index value of the corresponding Unicode code is compared with a predetermined code segment header index value and/or a segment tail index value, thereby determining the first predetermined character.
  • the corresponding predetermined code segment and the predetermined category Please refer to Table 2 above for details. For example, if the index value corresponding to the Unicode of the first predetermined character is greater than or equal to 1 and less than 10, the predetermined code segment corresponding to the first predetermined character and the predetermined category may be determined, and the The name string corresponding to a predetermined character is divided into the A category of the predetermined code segment.
  • a unified character set may be integrated in the operating system, that is, the unified code corresponding to each first predetermined character may be obtained from an operating system.
  • the segment header index value of the predetermined code segment that can be stored only in the client, so that when the Unicode segment is segmented, different predetermined code segments are distinguished by the segment header index value and/or the segment tail index value.
  • the client stores an index value of 26 predetermined code segments, for a total of 26; or a segment header index value and a segment tail of a predetermined code segment.
  • the index value is 52 in total. Therefore, the storage space required is relatively small compared to the client.
  • the manner of determining the predetermined code segment corresponding to the Unicode of the first predetermined character is not limited to the above description.
  • a person skilled in the art may make other changes in the spirit of the technical application of the present application, but the functions and effects of the present invention are the same or similar to the present application, and should be covered by the present application.
  • the application server may be a hardware implementation manner of the method for processing the name list of the present application, and the implementation method of the name list processing method of the present application may be implemented and the technical effects of the method implementation manner may be achieved.
  • At least one of the predetermined code segments constitutes a language coded block; wherein different language coded blocks have at least partially identical predetermined categories; correspondingly, the predetermined class segmentation module is configured to: When the predetermined code segment corresponding to the predetermined code segment in which the name string of the language coding block is located is the same, the name string is equally divided into the predetermined category to implement the shuffling of the plurality of languages.
  • the predetermined class division module 24 includes a language coding block determining unit 240 and a predetermined code segment determining unit 242.
  • a language encoding block determining unit 240 configured to determine a language encoding block corresponding to the Unicode of the first predetermined character in the name string;
  • the predetermined code segment determining unit 242 is configured to determine, in the language coded block, a predetermined code segment in which the Unicode code of the first predetermined character is located.
  • the embodiment of the present application further provides a server, which may include: a processor.
  • the processor is configured to obtain the name data; wherein the name data includes at least one name string; and is used to obtain a Unicode of the first predetermined character in the name string; and determine a Unicode of the first predetermined character
  • the name string is divided into predetermined categories corresponding to the predetermined code segments; wherein the predetermined number of code segments is at least one, and each of the predetermined code segments corresponds to a predetermined category.
  • the processor can be implemented in any suitable manner.
  • a processor can employ, for example, a microprocessor or processor and a computer readable medium, logic gate, switch, or application-specific integrated circuit (such as software or firmware) that can be executed by the (micro)processor.
  • ASIC Application Specific Integrated Circuit
  • programmable logic controller programmable logic controller and embedded microcontroller form, etc. This application is not limited.
  • the application server may be a hardware implementation manner of the method for processing the name list of the present application, and the implementation method of the name list processing method of the present application may be implemented and the technical effects of the method implementation manner may be achieved.

Landscapes

  • Engineering & Computer Science (AREA)
  • Signal Processing (AREA)
  • Document Processing Apparatus (AREA)

Abstract

一种名称列表展示、处理方法及客户端、服务器。所述名称列表展示方法包括:获取名称数据;其中,所述名称数据包括有至少一个名称字符串(S10);获取所述名称字符串中第一预定字符的统一码(S12);确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别(S14);在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串(S16)。所述名称列表展示方法及客户端,能够有效适应多种语言生成名称列表。

Description

名称列表展示、处理方法及客户端、服务器 技术领域
本申请涉及数据处理领域,特别涉及一种名称列表展示、处理方法及客户端、服务器。
背景技术
互联网改变了人们的日常生活,使得人们可以在不见面的条件下进行沟通交流。这种沟通交流并不仅仅限于语言文字之间的对话,已经延伸到商业贸易,比如网络购物,金融管理等。
通常软件内会设置有联系人列表,以向用户提供联系人的信息。现有的软件内,会对联系人列表的数据进行分类排序,以使用户能够快速查询到目标联系人。具体的,例如,在软件中集成有pinyin4j字典,其中包括了每个汉字对应的拼音,使得输入的汉字可以通过查找字典查到对应的拼音。在生成联系人列表时,可以通过拼音的首字母进行分类排序。
随着全球经济一体化的发展,同一个软件可能会在多个国家进行推广使用。例如:手机淘宝、支付宝等。在一些情况下,一些软件中会存在联系人列表,以供用户查阅联系人,从而便于用户之间进行交互。例如,进行即时通讯,或者进行金融转账等等。这些软件在多个国家推广时,会面临不同国家的用户可能使用不同语言的情况,再者同一个国家内的用户,其存储的联系人名称也可能包括不同的语言。例如,一个中国用户有日本朋友、美国朋友、朝鲜朋友等等。在这样的情况下,涉及在生成联系人列表时,需要针对不用语言的联系人名称进行列表排序。如果针对每种语言都设置一个文字与发音的字典,制作字典以及后续更新针对字典进行更新,需要进行大量的人力劳动,费时费力。
发明内容
本申请实施方式提供一种能够有效适应多种语言生成名称列表的名称列表展示、处理方法及客户端、服务器。
为解决上述技术问题,本申请提供一种名称列表展示方法,其包括:获取名称数据;其中,所述名称数据包括有至少一个名称字符串;获取所述名称字符串中第一预定字符的统一码;确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分 至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别;在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串。
本申请实施方式还提供一种客户端,其包括:数据获取模块,用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;统一码获取模块,用于获取所述名称字符串中第一预定字符的统一码;预定类别划分模块,用于确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别;显示模块,用于在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串。
本申请实施方式还提供一种客户端,其包括:处理器、显示器,所述处理器用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;并用于获取所述名称字符串中第一预定字符的统一码;确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别;控制所述显示器在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串。
本申请实施方式还提供一种名称列表处理方法,其包括:获取名称数据;其中,所述名称数据包括有至少一个名称字符串;获取所述名称字符串中第一预定字符的统一码;确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
本申请实施方式还提供一种服务器,其包括:数据获取模块,用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;统一码获取模块,用于获取所述名称字符串中第一预定字符的统一码;预定类别划分模块,用于确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
本申请实施方式还提供一种服务器,其包括:处理器,所述处理器用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;并用于获取所述名称字符串中第一预定字符的统一码;确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
由以上本申请实施方式提供的技术方案可见,本申请实施方式基于名称字符串中第 一预定字符的统一码,将所述名称字符串划分至所述预定编码段对应的预定类别,并相应的对所述预定类别设置类别标识,在显示界面中对应展示所述预定类别的类别标识和与所述预定类别对应的名称字符串,实现了利用统一码与名称字符串的一一对应关系,将多个国家的语言在同一名称列表中进行分类展示。用户通过展示的类别标识可以快速定位至名称列表中的具体名称字符串,查找对对应的目标名称。所述名称列表展示方法在实现的过程中摆脱了传统的预设字典的方式,利用已有的统一码进行归类排序,减少了大量的人力劳动,节约了人力资源成本。
附图说明
为了更清楚地说明本申请实施方式或现有技术中的技术方案,下面将对实施方式或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请中记载的一些实施方式,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本申请一个实施方式提供的名称列表展示方法的流程图;
图2为本申请一个实施方式提供的名称列表展示方法的流程图;
图3为本申请一个实施方式提供的名称列表展示方法的流程图;
图4为本申请一个实施方式提供的名称列表展示方法的流程图;
图5为本申请一个实施方式提供的客户端的模块图;
图6为本申请一个实施方式提供的客户端的模块图;
图7为本申请一个实施方式提供的名称列表处理方法的流程图;
图8为本申请一个实施方式提供的名称列表处理方法的流程图;
图9为本申请一个实施方式提供的名称列表处理方法的流程图;
图10为本申请一个实施方式提供的名称列表处理方法的流程图;
图11为本申请一个实施方式提供的服务器的模块图;
图12为本申请一个实施方式提供的预定类别划分模块的示意图。
具体实施方式
为了使本技术领域的人员更好地理解本申请中的技术方案,下面将结合本申请实施方式中的附图,对本申请实施方式中的技术方案进行清楚、完整地描述,显然,所描述的实施方式仅仅是本申请一部分实施方式,而不是全部的实施方式。基于本申请中的实施方式,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施方式,都应当属于本申请保护的范围。
请参阅图1,本申请的一个实施方式提供一种名称列表展示方法可以包括如下步骤。
步骤S10:获取名称数据;其中,所述名称数据包括有至少一个名称字符串。
在本实施方式中,客户端可以是具有网络通讯功能的通信设备,例如台式电脑、笔记本电脑、平板电脑、智能手机和智能可穿戴设备等。当然,客户端也可以为运行于上述通信设备中的软件。所述客户端可以被用户使用,可以进行即时通讯等交互活动,在即时通讯过程中,可以进行网上交易、在线电子支付等各种电子商务活动。
在本实施方式中,所述名称数据可以预存于所述客户端中。当用户需要进行即时通讯、电子商务等交互活动时,可以通过处理器调取所述预存的名称数据。所述名称数据也可以通过从服务器端下载获得。当用户需要进行电子商务活动时,可以通过数据通讯模块从服务器端进行下载所述名称数据。或者,所述名称数据也可以由用户以输入的方式临时建立获得。此外,获取所述名称数据的方式还可以为其他方式,例如,接受其他客户端发送的名称数据等等,本申请在此并不作具体限定。
在本实施方式中,所述名称数据可以为用户在进行即时通讯、电子商务等交互活动时需要使用的名称集合。所述名称具体的可以为:姓名、昵称、公司名称、账户名称、地点名称等等,本申请在此并不作具体的限定。
在本实施方式中,所述名称数据包括有至少一个名称字符串。所述名称字符串可以为所述姓名、昵称、公司名称、账户名称、地点名称中的一个具体名字。具体的,所述名称字符串可以为具体的一个姓名,例如“王二小”。所述名称字符串也可以为具体的一个账户名称,例如某个银行的银行卡号:“622848440637874213”。所述名称字符串的形式具体可以为某种语言的至少一个字、词、词组等,或者可以为一种符号、数字、字母或者多种结合等,当然所述名称字符串还可以为其他形式,本申请在此并不作具体的限定。
步骤S12:获取所述名称字符串中第一预定字符的统一码。
在本实施方式中,统一码用于为每种语言中的字符设定了统一并且唯一编码,以满足跨语言、跨平台进行文本转换、处理的要求。所述统一码具有一个字符集,其可以用预定进制的数字来映射这些字符,使得在所述统一码的字符集中的字符与每种语言中的字符存在一一对应关系。即一个统一码唯一对应有一个字符。所述预定进制的数字可以为十六进制,二进制等等,此处不作一一陈述。具体的,例如Unicode码为一种典型的统一码,其用十六进制数字0-0x10FFFF来映射各种语言的字符。如,中、日、韩CJK(Chinese Japanese Korean)的三种文字占用了Unicode中0x3000到0x9FFF的部分。Unicode目前普遍采用的是UCS-2规范编码,具体的,其用两个字节来编码一个字符。比如汉字“经”的编码是0x7ECF。由于字符编码一般用十六进制来表示,为了与十进制 区分,十六进制以0x开头,0x7ECF转换成十进制就是32463,UCS-2用两个字节来编码字符,两个字节就是16位二进制,2的16次方等于65536,所以UCS-2最多能编码65536个字符。
在本实施方式中,在所述客户端可以存储有统一字符集,所述统一字符集中可以包括字符、以及与字符相对应的统一码。此外,所述统一字符集也可以存储于服务器端,客户端可以通过通信模块与所述服务器端建立通信后,获取所述统一字符集中的相关字符以及与字符相对应的统一码。所述每个第一预定字符可以唯一对应有一个统一码。具体的,所述第一预定字符可以为所述名称字符串的首字符,也可以为所述名称字符串中的首字符与第二字符的结合,也可以单独为第二字符,或者还可以是由用户预先指定的所述名称字符串中的某个预定字符。当然,所述第一预定字符并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在本实施方式中,由于所述第一预定字符可以唯一对应有一个统一码,根据所述第一预定字符,通过查找的方式或者通过所述第一预定字符与统一码的映射关系,可以获得所述名称字符串中第一预定字符的统一码。具体的,例如可以通过将所述第一预定字符转换为十六进制数字,然后在所述统一字符集中进行查找。若查找到在所述统一字符集中相同的十六进制数字,则可以唯一确定所述第一预定字符的统一码。或者,根据所述第一预定字符,通过字符与十六进制数字的映射关系,在所述统一字符集中进行查找,以获取所述名称字符串中第一预定字符的统一码。当然,所述获取的方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
步骤S14:确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
在本实施方式中,可以利用编程语言API(Application Programming Interface,应用程序编程接口)中的语言排序机制,根据语言的发音、拼写等特点,将所述第一预定字符的统一码进行划分,设定至少一个预定编码段。其中每个所述预定编码段可以对应设置为一个预定类别,或者多个预定编码段可以对应设置为同一个预定类别,进而将多种语言的文字进行归类至有限的类别。
由于一般情况下,统一编码并不按语言的发音等规律进行排列。为了使得预定编码段的划分符合用户语言习惯,可以利用排序器,例如Java中的collator,根据语言的发音等特点,将字符进行排序。进一步的,将排序后的字符对应的统一码划分为至少一个预 定的编码段。具体的,所述排序器collator在进行排序时,能够根据语言的特点,比较两个字符的先后顺序,进而能够将统一字符集中的字符按照语言特点进行分段排序。例如,在Java语言标准库里面,对于中文,可以根据其发音,所有“啊”(a)发音的第一预定字符排在所有“吧”(ba)发音的第一预定字符的前面。所有“吧”(ba)发音的第一预定字符排在所有“擦”(ca)发音的第一预定字符的前面。以此类推,根据字符的发音将名称数据中的第一预定字符分为“a”至“z”26个预定编码段。
在本实施方式中,所述预定类别为多个名称字符串的集合。所述预定类别在数据结构上可以通过队列、数据栈、数组等方式实现。例如,当所述预定类别的数据结构方式为数组的数据结构方式时,可以将第一预定字符发音属于同一预定编码段的名称字符串存储于同一个数组下,以形成一个预定类别。例如,将“陈一”、“成二”、“蔡三”、“程四”存储于同一个数组下,以形成一个预定类别。当然,形成所述预定类别的数据结构方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在一个具体的实施方式中,为了使得预定编码段的划分符合用户语言习惯,可以利用Java中的collator根据语言的发音、拼写等特点将第一预定字符对应的统一码按序排列,并分为至少一个预定编码段。当需要判断某个第一预定字符具体属于对应的某个预定编码段时,可以将所述第一预定字符与所述预定编码段的段首和/或段尾的统一码对应的字符比较。若该第一预定字符的顺序为位于所述预定编码段的段首字符以后,而位于段尾字符以前,或者位于下一个段首字符之前的顺序,则可以确定该第一预定字符属于此预定编码段。
此外,在划分预定编码段时,预定编码段段首和/或段尾统一码与预定类别之间可以形成有对应关系。请参看表1。当确定了预定编码段后,可以根据所述预定编码段与预定类别的对应关系,确定所述名称字符串的预定类别。
在本实施方式中,利用了现有的Java中的collator,实现了根据语言的发音、拼写等特点进行比较排序。当然,具体的排序方式并不限于上述描述,所属领域技术人员在本申请的技术精髓启发下,还可以作出其他方式的变更,例如选择现有的其他排序器或者根据排序功能自行进行编写等,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
表1
Figure PCTCN2016072307-appb-000001
在本实施方式中,操作系统中可以集成有统一字符集,即每个字符对应的统一码可以从操作系统中获得。所述客户端中可以仅存储所述预定编码段的段首统一码,如此可以实现在统一码分段时,通过段首统一码,区分不同的预定编码段;或者是所述预定编码段的段首统一码和所述预定编码段的段尾段尾统一码。具体的,例如当所述预定编码段为26个时,对应的,所述客户端存储的为26个预定编码段的统一码,共26个;或者预定编码段的段首统一码和段尾的统一码,共52个。因此,相对于客户端来说,所需占用的存储空间较小。
在另一个具体的实施方式中,由于一般情况下,统一编码并不按每种语言的发音等规律进行排列。为了使得预定编码段的划分符合用户语言习惯,可以将第一预定字符对应的统一码根据语言的发音等特点对应一个索引值,根据所述索引值进行排序,进而可以将所述根据索引值排序后的统一码分为至少一个预定的编码段。
例如,在Java语言标准库里面,对于中文,可以根据其发音,给所述第一预定字符对应的统一码一个对应的索引值。所述索引值可以看成一个数字。比如“啊”(a),其索引值可能是1;而“吧”(ba),其索引值可能是10,那么所有“啊”(a)至“吧“(ba)之间发音的汉字,它的索引值在1-10之间。
请参阅表2。针对每个预定编码段可以对应有一个预定类别。例如,编码段段首索引值为[1,10)对应的预定类别为A类别,即实现了将字符发音的首字母为a的字,例如“啊”、“安”等归至A类别中。编码段段首索引值为[10,20)对应的预定类别为B类别,即实现了将字符发音的首字母为b的字,例如“吧”、“被”等归至B类别中。以此类推,编制Z类别。从而将整个中文编码块划分至26个编码段,对应的26个预定类别。
表2
Figure PCTCN2016072307-appb-000002
在本实施方式中,通过将所述第一预定字符的统一码的索引值与预设编码段的段首索引值和/或段尾索引值进行查找比较的方式,可以确定所述第一预定字符对应的预定编码段。由于所述预定编码段对应有预定类别,因此可以将所述名称字符串划分至所述预定编码段对应的预定类别。
在一个具体的实施方式中,对于某个第一预定字符,将其对应的统一码的索引值和预定编码段首索引值和/或段尾索引值进行比较,进而确定所述第一预定字符所对应的预定编码段以及预定类别。请结合参阅表2。例如,如果某个第一预定字符的统一码对应的索引值大于等于所述1而小于10,则可以确定所述第一预定字符所对应的预定编码段和预定类别,进而可以将所述第一预定字符对应的所述名称字符串划分至所述预定编码段的A类别中。
在本实施方式中,操作系统中可以集成有统一字符集,即每个第一预定字符对应的统一码可以从操作系统中获得。所述客户端中可以仅存储的所述预定编码段的段首索引值,如此可以实现在统一码分段时,通过段首索引值和/或段尾索引值,区分不同的预定编码段。具体的,例如当所述预定编码段为26个时,对应的,所述客户端存储的为26个预定编码段的索引值,共26个;或者预定编码段的段首索引值和段尾的索引值,共52个。因此,相对于客户端来说,所需占用的存储空间较小。
当然,所述确定所述第一预定字符的统一码对应的预定编码段的方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
步骤S16:在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串。
在本实施方式中,所述预定类别可以对应有类别标识,在显示界面中展示时,可以 将所述预定类别的类别标识和与所述预定类别对应的名称字符串一同展示出来,形成便于用户查找的名称列表。
具体的,所述类别标识可以为根据用户语言习惯,便于用户识别的标识。例如,对于使用中文的用户,其类别标识可以为所述第一预定字符发音的首字母。对于使用英文的用户,其类别标识可以为所述第一预定字符的首字母。当然,对于其他语言种类而言,其类别标识也可以为所述第一预定字符发音的音标或者是拼写的首字母等等,此处不再重复举例。
在本实施方式中,所述预定类别对应的名称字符串可以全部显示或者显示部分所述名称字符串或者不显示所述名称字符串。例如,所述预定类别包括有很多名称字符串,此时,可设定在所述预定类别对应的类别标记下最多显示预定个数的名称字符串,例如可以设定最多显示3个名称字符串。此外,也可以设定在所述显示界面只显示所述类别标记。所述类别标记的相应位置检测到用户点击产生的信号时,可进一步向用户展示单个类别标记下的名称字符串。
本申请实施方式所述名称列表展示方法基于名称字符串中第一预定字符的统一码,将所述名称字符串划分至所述预定编码段对应的预定类别,并相应的对所述预定类别设置类别标识,在显示界面中对应展示所述预定类别的类别标识和与所述预定类别对应的名称字符串,实现了利用统一码与名称字符串的一一对应关系,将多个国家的语言在同一名称列表中进行分类展示。用户通过展示的类别标识可以快速定位至名称列表中的具体名称字符串,查找对对应的目标名称。所述名称列表展示方法在实现的过程中摆脱了传统的预设字典的方式,利用已有的统一码进行归类排序,减少了大量的人力劳动,节约了人力资源成本。
可以理解,在目前现有的操作系统中,通常都集成有统一码,以作为跨地域或数据平台数据流动时,对文字数据进行编码解码的依据。在本申请中,通过利用统一码对名称列表进行分类排序,可以实现没有在现有客户端中引入的额外字典的方式,进行名称列表分类排序。具体的,例如可以不引入类似pinyin4j字典,实现节省了存储空间。
在一个实施方式中,所述名称数据的具体内容可以包括,人名、公司名称、账户名称、地点名称中的至少一个。
在本实施方式中,所述名称数据可以包括至少一个名称字符串。所述名称字符串具体的可以为人名、公司名称、账户名称、地点名称等等,本申请在此并不作具体的限定。当所述名称数据的具体内容为人名时,所述人名可以为用户的真实姓名,也可以为用户给自己设定的昵称等。
所述名称数据的种类可以为只包括一个名称种类,例如只包括人名;也可以同时包 括多个名称种类,例如同时包括人名、公司名称、账户名称等等。当所述名称数据同时包括多个名称种类时,若多个名称种类之间有对应关系,则可以将所述多个名称种类进行划分,例如可以设定多个名称种类中的一种的名称字符串为主名称,设定多个名称种类中的其他名称作为从属名称。展示时,可以只将所述主名称在显示界面中进行显示,其他与主名称相对应的从属名称,可以与所述主名称对应存储。当主名称所在的位置检测到点击触发的电信号时,可以显示所述从属名称。另外,当所述从属名称在显示界面显示后,还可以对从属名称进行复制、编辑等操作。
例如,名称数据具体内容为“甲公司小王中国银行账户XX”,设定人名为主名称,其他公司、账户名称为从属名称。展示时,将所述主名称“小王”在显示界面中进行显示,其他与“小王”相对应的从属名称“甲公司”、“中国银行账户XX”,可以与所述主名称“小王”对应存储。当点击所述主名称“小王”时,可以相应地显示所述从属名称“甲公司”、“中国银行账户XX”,进一步地,可以对从属名称进行复制、编辑等操作。
在一个实施方式中,所述第一预定字符可以为所述名称字符串的首字符。
在本实施方式中,所述名称字符串具体的可以为人名、公司名称、账户名称、地点名称等等。所述名称字符串可以包括一个或者多个字符。而所述名称字符串的首字符很多情况下能够典型地代表整个名称字符串。例如,当所述名称字符串为人名时,其首字符可以为姓氏,当所述名称字符串为公司名称时,其首字符也可为公司名称第一个字。当将所述第一预定字符设定为所述名称字符串的首字符,后续以所述名称字符串进行分类形成的名称列表符合用户的使用习惯,有助于用户基于形成的名称列表进行查询。
在一个实施方式中,可以通过Java中的collator确定所述第一预定字符的统一码对应的预定编码段。
在本实施方式中,由于一般情况下,统一编码并不按语言的发音等规律进行排列。为了使得预定编码段的划分符合用户语言习惯,可以利用排序器,例如Java中的collator,根据语言的发音等特点,将字符进行排序。进一步的,将排序后的每个字符对应的统一码划分为至少一个预定的编码段。具体的,所述排序器collator在进行排序时,能够根据语言的特点,比较两个字符的先后顺序,进而能够将统一字符集中的字符按照语言特点进行分段排序。例如,在Java语言标准库里面,对于中文,可以根据其发音,所有“啊”(a)发音的第一预定字符排在所有“吧”(ba)发音的第一预定字符的前面。所有“吧”(ba)发音的第一预定字符排在所有“擦”(ca)发音的第一预定字符的前面。以此类推,根据第一预定字符的发音将名字数据中的中文字符分为“a”至“z”26个预定编码段。
在本实施方式中,利用了现有的Java中的collator,实现了根据语言的发音、拼写等 特点进行比较排序。当然,具体的排序方式并不限于上述描述,所属领域技术人员在本申请的技术精髓启发下,还可以作出其他方式的变更,例如选择现有的其他排序器或者根据排序功能自行进行编写等,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在一个具体的实施方式中,例如,新增加了一个名称字符串,需要判断所述新增加的名称字符串的第一预定字符具体属于对应的某个预定编码段时,可以将所述第一预定字符与所述预定编码段的段首和/或段尾的统一码对应的字符比较。若该第一预定字符的顺序为位于所述预定编码段的段首字符以后而位于段尾字符以前的顺序,或者位于下一个段首字符之前的顺序,则可以确定该第一预定字符属于此预定编码段。
在一个实施方式中,至少一个所述预定编码段组成一个语言编码块;其中,不同所述语言编码块具有至少部分相同的预定类别。相应的,在将预定编码划分预定类别的步骤中包括:在属于不同所述语言编码块的名称字符串所处的预定编码段对应的预定类别相同时,将所述名称字符串同划分至所述预定类别的预设数据集中。
在本实施方式中,在统一码中针对不同的语言种类,可以设定有不同的语言编码块。例如,可以包括中日韩文编码块、英文编码块、俄文编码块、阿拉伯文编码块等等。其中,不同所述语言编码块可以具有至少部分相同的预定类别。
在一个具体的实施方式中,请结合表1所示。对于中文编码块而言,利用Java中的collator可以根据语言的发音等特点对应设定多个预定编码段。例如,可以根据所述第一预定字符的发音特点,将所述第一预定字符对应的统一码分为“a”至“z”26个预定编码段。针对每个预定编码段可以对应有一个预定类别。例如所述“a”至“z”26个预定编码段分别对应的预定类别为A至Z26个类别。
对于英文编码块而言,请参看表3,利用Java中的collator可以根据字符拼写特点对应设置多个预定编码段。此外,当所述英文编码块的统一码本身排列顺序为按照A至Z的顺序排列时,也可以根据所述统一码本身排列顺序设置多个预定编码段。例如,可以根据第一预定字符的拼写特点,将所述第一预定字符分为A到Z26个预定编码段。针对每个预定编码段可以对应有一个预定类别。例如,第一预定编码段段首统一码97对应的预定类别为A类别,即将首字母为A的名称字符串归至预定类别A中。第二预定编码段段首统一码98对应的预定类别为B类别,即将首字母为B的名称字符串归至预定类别B中。依次类推,编制预定类别Z。从而将整个英文编码块划分至26个编码段,对应的26个预定类别。
表3
Figure PCTCN2016072307-appb-000003
请参阅表4,由于所述中文编码块的预定类别和英文编码块的预定类别从A至Z相同,因此可以将属于相同预定类别的名称字符串划分至所述预定类别中,形成所述预设数据集,进而实现中英文混合排序。
此外,所述英文编码块利用Java中的collator根据字符拼写特点对应设置多个预定编码段时,对应的每个预定编码段的段首或段尾可以设置有索引值,当所述英文编码块对应的索引值与所述中文编码块对应的索引值相同时,则可以确定所述索引值相同的名称字符串属于同一类别。
表4
Figure PCTCN2016072307-appb-000004
在另一个实施方式中,请参阅表5,所述中文编码块与所述英文编码块也可以不划分至相同的所述预定类别中,即各自按照各自的预定类别进行划分,不进行混排。例如,英文编码块、中午编码块按照统一码的先后顺序进行分类排列,可以整体上按照英文编码块在前,中文编码块在后的顺序进行分类和排列。不同的编码块具体的分类的标记可以相同,也可以不同,本申请在此并不作具体的限定。
表5
英文编码块 预定编码段段首统一码 预定类别
  //&#x97 A类别
  //&#x98 B类别
  //&#x99 C类别
  //&#x100 D类别
  //&#x101 E类别
  //&#x102 F类别
  //…
中文编码块 预定编码段段首统一码 预定类别
  //&#x554a A类别
  //&#x82ad B类别
  //&#x64e6 C类别
  //&#x642d D类别
  //&#x86fe E类别
  //&#x53d1 F类别
  //…
当然,所述语言编码块的种类并不限于举例的中文语言编码块、英文语言编码块,数量也不限于两种语言编码块。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
请参阅图2,在一个实施方式中,在确定预定编码段的步骤S14中可以包括如下步骤。
步骤S140:确定所述名称字符串中第一预定字符的统一码对应的语言编码块。
步骤S142:在所述语言编码块中,确定所述第一预定字符的统一码所在的预定编码段。
在本实施方式中,在确定所述预定编码段时,可以先根据所述名称字符串中第一预定字符的统一码确定其对应的语言编码块。在基于确定的语言编码块的基础上,进一步确定所述第一预定字符的统一码所在的预定编码段。
具体的,在统一码中针对不同的语言种类,可以设定有不同的语言编码块。每个语言编码块的首字符的统一码以及字符的个数都是已知的。在确定所述名称字符串中的第 一预定字符的统一码对应的语言编码块时,可将所述第一预定字符的统一码与每个语言编码块的首字符统一码进行比较。当所述第一预定字符的统一码大于等于某个语言编码块的首字符统一码,且其与某个语言编码块的首字符统一码的差值小于所述语言编码块的个数时,则可以确定所述第一预定字符为所述语言编码块中的字符。
当确定了所述名称字符串中第一预定字符的统一码对应的语言编码块后,可以将所述第一预定字符的统一码在所述确定的语言编码块内进行查找,以便快速确定所述第一预定字符的统一码所在的预定编码段。具体的,所述确定的语言编码块一般地包括有多个语言编码段,每个语言编码段至少对应有一个统一码。确定所述第一预定字符的统一码时,可以将所述第一预定字符的统一码与所述语言编码段的统一码进行对比,若有相同,则可以确定所述第一预定字符的统一码所在的预定编码段。
在一个具体的实施方式中,如表3所示,英文语言编码块可以包括26个语言编码段。在这26个语言编码段中,每个预定编码段对应有一个统一码。当所述预定编码段对应的统一码为按照大小顺序排列时,可以将所述一一预定字符的统一码与每个预定编码段的统一码进行比较,若有相同,则可以确定所述第一预定字符的统一码所在的预定编码段。
在一个实施方式中,所述方法还可以包括:在没有与所述名称字符串中第一预定字符的统一码对应的语言编码块时,将所述名称字符串划分至预设指定类别。
在本实施方式中,可以预先在客户端设置有预设指定类别,用于归类没有查找到对应语言编码块的名称字符串。所述预设指定类别可以为排列于所述预定类别后的一个单独的类别,其可以用一个特殊的标记进行展示。例如可以用“#”表示其下包括的是不属于统一码所包括的语言编码块的预设指定类别。
在本实施方式中,在没有与所述名称字符串中第一预定字符的统一码对应的语言编码块时,说明所述名称字符串中第一预定字符不属于统一码中已有的语言或符号。此时,可以将其划分至预设指定类别。
在一个实施方式中,在展示类别标识和名称字符串的步骤S16中还可以包括:
将属于同一所述预定类别的所述名称字符串按照所述第一预定字符的统一码的编码顺序排列。
在本实施方式中,所述名称字符串的第一预定字符可以对应有唯一的统一码。将属于同一预定类别的所述名称字符串进行排序时,可以按照所述第一预定字符的统一码的编码顺序排列。
具体的,由于所述第一预定字符的统一码的编码本身为一个十六进制的数字,可以依照数字的大小顺序进行排列。例如可以按照所述第一预定字符的统一码由小到大的顺序对所述同一预定类别下的名称字符串进行排列。
在一个具体的实施方式中,例如对于C类别下的名称字符串:第一预定字符,例如“陈一”、“成二”、“蔡三”,可以根据其第一预定字符“蔡”、“陈”、“成”统一码由小到大的顺序的特点,将所述C类别下的名称字符串排列为“蔡三”、“陈一”、“成二”,实现了同一类别下,名称字符串的排序。
请参阅图3,在一个实施方式中,在所述第一预定字符的统一码相同时,所述方法还可以包括如下步骤。
步骤S161:获取所述名称字符串中第二预定字符。
步骤S163:将属于同一所述预定类别的所述名称字符串按照所述第二预定字符的统一码的编码顺序排列。
在本实施方式中,所述第一预定字符的统一码相同的情况,具体可表示所述名称字符串包含有相同的第一预定字符。
在本实施方式中,所述第二预定字符可以为所述名称字符串中,除了所述第一预定字符之外的一个字符或多个字符的组合。具体的,由于所述第二预定字符的统一码的编码本身可为一个十六进制的数字,可以依照数字的大小顺序进行排列。例如可以按照所述第二预定字符的统一码由小到大的顺序对所述同一预定类别下的名称字符串进行排列。
在一个具体的实施方式中,当所述名称字符串都为姓名时,其可以表示包含多个姓氏相同的名字。例如,在W类别下,名称字符串包括:“王一”、“王二”、“王三”。分别获取所述第一预定字符的统一码时,发现其对应的统一码均相同。此时,进一步的可以获取除了所述第一预定字符之外第二字符的统一码。具体可获取的第二预定字符分别为“一”、“二”、“三”的统一码。在统一码字符集中,所述第二预定字符“一”、“二”、“三”的统一码的排序为“二”、“三”、“一”。对应的,可以将与所述第二预定字符相对应的名称字符串排列为“王二”、“王三”、“王一”,从而实现了,当第一预定字符相同时,对同一类别下的名称字符串排序。
请参阅图4,在一个实施方式中,在展示类别标识和名称字符串的步骤S14中还可以包括如下步骤。
步骤S160:获取所述名称字符串中第二预定字符的统一码。
步骤S162:将属于同一所述预定类别的所述名称字符串按照所述第二预定字符的统一码的编码顺序排列。
在本实施方式中,所述第二预定字符可以为所述名称字符串中,除了所述第一预定字符之外的一个字符或多个字符的组合。具体的,由于所述第二预定字符的统一码的编码本身为一个十六进制的数字,可以依照数字的大小顺序进行排列。例如可以按照所述 第二预定字符的统一码由小到大的顺序对所述同一预定类别下的名称字符串进行排列。
在一个具体的实施方式中,当所述名称字符串为姓名时,例如在C类别下,名称字符串包括:“陈一”、“成二”、“蔡三”。上述名称字符串的第一预定字符为“蔡”、“陈”、“成”;第二预定字符为“一”、“二”、“三”。当名称字符串直接按照所述第二预定字符的统一码进行排序时,由于第二预定字符为“一”、“二”、“三”对应的统一码由小到大的排序为“二”、“三”、“一”,因此,C类别下对应的名称字符串的排序为“成二”、“蔡三”、“陈一”。本实施方式通过直接根据第二预定字符的统一码对与其对应的名称字符串进行排序,对于第一预定字符相同的情况来说,减少了不必要的第一预定字符统一码比较,能够提高排序的效率。
在一个实施方式中,在获取名称数据的步骤S10中还可以包括:接收用户指定的第一预定字符。
在本实施方式中,在获取名称数据时,可以接收用户指定的第一预定字符。具体的,所述指定的第一预定字符可为所述名称字符串中的一个字符或多个字符的组合。所述指定的第一预定字符用于在所述名称字符串归至预定类别后,将属于同一预定类别名称字符串进行排序。所述指定的第一预定字符具体的可以为名称的姓氏列表,可以为地点的国别列表,可以为账户的银行类别列表等等,本申请在此并不作具体限定。
在一个具体的实施方式中,所述指定的第一预定字符为一个姓氏列表时,所述姓氏列表中的姓氏可以按照顺序排列好。例如姓氏列表中包括“Brown”、“Smith”、“White”等,其按照先后顺序排列为:“Brown”、“Smith”、“White”。当所述名称数据包括的名称字符串为:“Alice Brown”、“Ada Smith”、“Alma White”。根据所述指定排序字符中“Brown”、“Smith”、“White”的排列顺序,上述名称字符串排列先后顺序为“Alice Brown”、“Ada Smith”、“Alma White”。
请参阅图5,本申请实施方式还提供一种客户端100,其可以包括:数据获取模块10、统一码获取模块12、预定类别划分模块14、显示模块16。
数据获取模块10,用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串。
在本实施方式中,客户端可以是具有网络通讯功能的通信设备,例如台式电脑、笔记本电脑、平板电脑、智能手机和智能可穿戴设备等。当然,客户端也可以为运行于上述通信设备中的软件。所述客户端可以被用户使用,可以进行即时通讯等交互活动,在即时通讯过程中,可以进行网上交易、在线电子支付等各种电子商务活动。
在本实施方式中,所述名称数据可以预存于所述客户端中。当用户需要进行即时通讯、电子商务等交互活动时,可以通过处理器调取所述预存的名称数据。所述名称数据 也可以通过从服务器端下载获得。当用户需要进行电子商务活动时,可以通过数据通讯模块从服务器端进行下载所述名称数据。或者,所述名称数据也可以由用户以输入的方式临时建立获得。此外,获取所述名称数据的方式还可以为其他方式,例如,接受其他客户端发送的名称数据等等,本申请在此并不作具体限定。
在本实施方式中,所述名称数据可以为用户在进行即时通讯、电子商务等交互活动时需要使用的名称集合。所述名称具体的可以为:姓名、昵称、公司名称、账户名称、地点名称等等,本申请在此并不作具体的限定。
在本实施方式中,所述名称数据包括有至少一个名称字符串。所述名称字符串可以为所述姓名、昵称、公司名称、账户名称、地点名称中的一个具体名字。具体的,所述名称字符串可以为具体的一个姓名,例如“王二小”。所述名称字符串也可以为具体的一个账户名称,例如某个银行的银行卡号:“622848440637874213”。所述名称字符串的形式具体可以为某种语言的至少一个字、词、词组等,或者可以为一种符号、数字、字母或者多种结合等,当然所述名称字符串还可以为其他形式,本申请在此并不作具体的限定。
统一码获取模块12,用于获取所述名称字符串中第一预定字符的统一码。
在本实施方式中,统一码用于为每种语言中的字符设定了统一并且唯一编码,以满足跨语言、跨平台进行文本转换、处理的要求。所述统一码具有一个字符集,其可以用预定进制的数字来映射这些字符,使得在所述统一码的字符集中的字符与每种语言中的字符存在一一对应关系。即一个统一码唯一对应有一个字符。所述预定进制的数字可以为十六进制,二进制等等,此处不作一一陈述。具体的,例如Unicode码为一种典型的统一码,其用十六进制数字0-0x10FFFF来映射各种语言的字符。如,中、日、韩CJK(Chinese Japanese Korean)的三种文字占用了Unicode中0x3000到0x9FFF的部分。Unicode目前普遍采用的是UCS-2规范编码,具体的,其用两个字节来编码一个字符。比如汉字“经”的编码是0x7ECF。由于字符编码一般用十六进制来表示,为了与十进制区分,十六进制以0x开头,0x7ECF转换成十进制就是32463,UCS-2用两个字节来编码字符,两个字节就是16位二进制,2的16次方等于65536,所以UCS-2最多能编码65536个字符。
在本实施方式中,在所述客户端可以存储有统一字符集,所述统一字符集中可以包括字符、以及与字符相对应的统一码。此外,所述统一字符集也可以存储于服务器端,客户端可以通过通信模块与所述服务器端建立通信后,获取所述统一字符集中的相关字符以及与字符相对应的统一码。所述每个第一预定字符可以唯一对应有一个统一码。具体的,所述第一预定字符可以为所述名称字符串的首字符,也可以为所述名称字符串中 的首字符与第二字符的结合,也可以单独为第二字符,或者还可以是由用户预先指定的所述名称字符串中的某个预定字符。当然,所述第一预定字符并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在本实施方式中,由于所述第一预定字符可以唯一对应有一个统一码,根据所述第一预定字符,通过查找的方式或者通过所述第一预定字符与统一码的映射关系,可以获得所述名称字符串中第一预定字符的统一码。具体的,例如可以通过将所述第一预定字符转换为十六进制数字,然后在所述统一字符集中进行查找。若查找到在所述统一字符集中相同的十六进制数字,则可以唯一确定所述第一预定字符的统一码。或者,根据所述第一预定字符,通过字符与十六进制数字的映射关系,在所述统一字符集中进行查找,以获取所述名称字符串中第一预定字符的统一码。当然,所述获取的方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
预定类别划分模块14,用于确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
在本实施方式中,可以利用编程语言API(Application Programming Interface,应用程序编程接口)中的语言排序机制,根据语言的发音、拼写等特点,将所述第一预定字符的统一码进行划分,设定至少一个预定编码段。其中每个所述预定编码段可以对应设置为一个预定类别,或者多个预定编码段可以对应设置为同一个预定类别,进而将多种语言的文字进行归类至有限的类别。
由于一般情况下,统一编码并不按语言的发音等规律进行排列。为了使得预定编码段的划分符合用户语言习惯,可以利用排序器,例如Java中的collator,根据语言的发音等特点,将字符进行排序。进一步的,将排序后的字符对应的统一码划分为至少一个预定的编码段。具体的,所述排序器collator在进行排序时,能够根据语言的特点,比较两个字符的先后顺序,进而能够将统一字符集中的字符按照语言特点进行分段排序。例如,在Java语言标准库里面,对于中文,可以根据其发音,所有“啊”(a)发音的第一预定字符排在所有“吧”(ba)发音的第一预定字符的前面。所有“吧”(ba)发音的第一预定字符排在所有“擦”(ca)发音的第一预定字符的前面。以此类推,根据字符的发音将名称数据中的第一预定字符分为“a”至“z”26个预定编码段。
在本实施方式中,所述预定类别为多个名称字符串的集合。所述预定类别在数据结构上可以通过队列、数据栈、数组等方式实现。例如,当所述预定类别的数据结构方式 为数组的数据结构方式时,可以将第一预定字符发音属于同一预定编码段的名称字符串存储于同一个数组下,以形成一个预定类别。例如,将“陈一”、“成二”、“蔡三”、“程四”存储于同一个数组下,以形成一个预定类别。当然,形成所述预定类别的数据结构方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在一个具体的实施方式中,为了使得预定编码段的划分符合用户语言习惯,可以利用Java中的collator根据语言的发音、拼写等特点将第一预定字符对应的统一码按序排列,并分为至少一个预定编码段。当需要判断某个第一预定字符具体属于对应的某个预定编码段时,可以将所述第一预定字符与所述预定编码段的段首和/或段尾的统一码对应的字符比较。若该第一预定字符的顺序为位于所述预定编码段的段首字符以后,而位于段尾字符以前,或者位于下一个段首字符之前的顺序,则可以确定该第一预定字符属于此预定编码段。
此外,在划分预定编码段时,预定编码段段首和/或段尾统一码与预定类别之间可以形成有对应关系。请参看上述表1。当确定了预定编码段后,可以根据所述预定编码段与预定类别的对应关系,确定所述名称字符串的预定类别。
在本实施方式中,利用了现有的Java中的collator,实现了根据语言的发音、拼写等特点进行比较排序。当然,具体的排序方式并不限于上述描述,所属领域技术人员在本申请的技术精髓启发下,还可以作出其他方式的变更,例如选择现有的其他排序器或者根据排序功能自行进行编写等,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在本实施方式中,操作系统中可以集成有统一字符集,即每个字符对应的统一码可以从操作系统中获得。所述客户端中可以仅存储所述预定编码段的段首统一码,如此可以实现在统一码分段时,通过段首统一码,区分不同的预定编码段;或者是所述预定编码段的段首统一码和所述预定编码段的段尾段尾统一码。具体的,例如当所述预定编码段为26个时,对应的,所述客户端存储的为26个预定编码段的统一码,共26个;或者预定编码段的段首统一码和段尾的统一码,共52个。因此,相对于客户端来说,所需占用的存储空间较小。
在另一个具体的实施方式中,由于一般情况下,统一编码并不按每种语言的发音等规律进行排列。为了使得预定编码段的划分符合用户语言习惯,可以将第一预定字符对应的统一码根据语言的发音等特点对应一个索引值,根据所述索引值进行排序,进而可以将所述根据索引值排序后的统一码分为至少一个预定的编码段。
例如,在Java语言标准库里面,对于中文,可以根据其发音,给所述第一预定字符对应的统一码一个对应的索引值。所述索引值可以看成一个数字。比如“啊”(a),其索引值可能是1;而“吧”(ba),其索引值可能是10,那么所有“啊”(a)至“吧“(ba)之间发音的汉字,它的索引值在1-10之间。
请参阅上述表2。针对每个预定编码段可以对应有一个预定类别。例如,编码段段首索引值为[1,10)对应的预定类别为A类别,即实现了将字符发音的首字母为a的字,例如“啊”、“安”等归至A类别中。编码段段首索引值为[10,20)对应的预定类别为B类别,即实现了将字符发音的首字母为b的字,例如“吧”、“被”等归至B类别中。以此类推,编制Z类别。从而将整个中文编码块划分至26个编码段,对应的26个预定类别。
在本实施方式中,通过将所述第一预定字符的统一码的索引值与预设编码段的段首索引值和/或段尾索引值进行查找比较的方式,可以确定所述第一预定字符对应的预定编码段。由于所述预定编码段对应有预定类别,因此可以将所述名称字符串划分至所述预定编码段对应的预定类别。
在一个具体的实施方式中,对于某个第一预定字符,将其对应的统一码的索引值和预定编码段首索引值和/或段尾索引值进行比较,进而确定所述第一预定字符所对应的预定编码段以及预定类别。请结合参阅表2。例如,如果某个第一预定字符的统一码对应的索引值大于等于所述1而小于10,则可以确定所述第一预定字符所对应的预定编码段和预定类别,进而可以将所述第一预定字符对应的所述名称字符串划分至所述预定编码段的A类别中。
在本实施方式中,操作系统中可以集成有统一字符集,即每个第一预定字符对应的统一码可以从操作系统中获得。所述客户端中可以仅存储的所述预定编码段的段首索引值,如此可以实现在统一码分段时,通过段首索引值和/或段尾索引值,区分不同的预定编码段。具体的,例如当所述预定编码段为26个时,对应的,所述客户端存储的为26个预定编码段的索引值,共26个;或者预定编码段的段首索引值和段尾的索引值,共52个。因此,相对于客户端来说,所需占用的存储空间较小。
当然,所述确定所述第一预定字符的统一码对应的预定编码段的方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
显示模块16,用于在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串。
在本实施方式中,所述预定类别可以对应有类别标识,在显示界面中展示时,可以 将所述预定类别的类别标识和与所述预定类别对应的名称字符串一同展示出来,形成便于用户查找的名称列表。
具体的,所述类别标识可以为根据用户语言习惯,便于用户识别的标识。例如,对于使用中文的用户,其类别标识可以为所述第一预定字符发音的首字母。对于使用英文的用户,其类别标识可以为所述第一预定字符的首字母。当然,对于其他语言种类而言,其类别标识也可以为所述第一预定字符发音的音标或者是拼写的首字母等等,此处不再重复举例。
在本实施方式中,所述预定类别对应的名称字符串可以全部显示或者显示部分所述名称字符串或者不显示所述名称字符串。例如,所述预定类别包括有很多名称字符串,此时,可设定在所述预定类别对应的类别标记下最多显示预定个数的名称字符串,例如可以设定最多显示3个名称字符串。此外,也可以设定在所述显示界面只显示所述类别标记。所述类别标记的相应位置检测到用户点击产生的信号时,可进一步向用户展示单个类别标记下的名称字符串。
本申请实施方式所述客户端基于名称字符串中第一预定字符的统一码,将所述名称字符串划分至所述预定编码段对应的预定类别,并相应的对所述预定类别设置类别标识,在显示界面中对应展示所述预定类别的类别标识和与所述预定类别对应的名称字符串,实现了利用统一码与名称字符串的一一对应关系,将多个国家的语言在同一名称列表中进行分类展示。用户通过展示的类别标识可以快速定位至名称列表中的具体名称字符串,查找对对应的目标名称。在实现的过程中摆脱了传统的预设字典的方式,利用已有的统一码进行归类排序,减少了大量的人力劳动,节约了人力资源成本。
请参阅图6,本申请实施方式还提供一种客户端110,其可以包括:处理器11、显示器13。
所述处理器11用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;并用于获取所述名称字符串中第一预定字符的统一码;确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别;控制所述显示器13在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串。
在本实施方式中,所述处理器11以按任何适当的方式实现。例如,处理器11以采取例如微处理器或处理器以及存储可由该(微)处理器执行的计算机可读程序代码(例如软件或固件)的计算机可读介质、逻辑门、开关、专用集成电路(Application Specific Integrated Circuit,ASIC)、可编程逻辑控制器和嵌入微控制器的形式等等。本申请并不 作限定。
本申请客户端可以是本申请名称列表展示方法的一种硬件实施方式,可以实现本申请名称列表展示方法实施方式并达到方法实施方式的技术效果。
本发明实施方式中的名称列表处理方法可以由服务器来执行。
请参阅图7,本申请的一个实施方式提供一种名称列表处理方法可以包括如下步骤。
步骤S20:获取名称数据;其中,所述名称数据包括有至少一个名称字符串。
在本实施方式中,服务器可以是具有数据处理能力和网络通讯功能的设备,例如个人计算机、服务器计算机、手持设备或便携式设备、平板型设备、多处理器装置、包括以上任何装置或设备的分布式计算环境。当然,服务器也可以包括运行于上述设备中的软件。
在本实施方式中,所述名称数据可以预存于所述服务器中。当用户需要进行电子商务活动时,可以通过数据通讯模块从服务器端进行下载所述名称数据。或者,所述名称数据也可以由用户以输入的方式临时建立获得。此外,获取所述名称数据的方式还可以为其他方式,例如,接受其他客户端发送的名称数据等等,本申请在此并不作具体限定。
在本实施方式中,所述名称数据可以为用户在进行即时通讯、电子商务等交互活动时需要使用的名称集合。所述名称具体的可以为:姓名、昵称、公司名称、账户名称、地点名称等等,本申请在此并不作具体的限定。
在本实施方式中,所述名称数据包括有至少一个名称字符串。所述名称字符串可以为所述姓名、昵称、公司名称、账户名称、地点名称中的一个具体名字。具体的,所述名称字符串可以为具体的一个姓名,例如“王二小”。所述名称字符串也可以为具体的一个账户名称,例如某个银行的银行卡号:“622848440637874213”。所述名称字符串的形式具体可以为某种语言的至少一个字、词、词组等,或者可以为一种符号、数字、字母或者多种结合等,当然所述名称字符串还可以为其他形式,本申请在此并不作具体的限定。
步骤S22:获取所述名称字符串中第一预定字符的统一码。
在本实施方式中,统一码用于为每种语言中的字符设定了统一并且唯一编码,以满足跨语言、跨平台进行文本转换、处理的要求。所述统一码具有一个字符集,其可以用预定进制的数字来映射这些字符,使得在所述统一码的字符集中的字符与每种语言中的字符存在一一对应关系。即一个统一码唯一对应有一个字符。所述预定进制的数字可以为十六进制,二进制等等,此处不作一一陈述。具体的,例如Unicode码为一种典型的统一码,其用十六进制数字0-0x10FFFF来映射各种语言的字符。如,中、日、韩CJK(Chinese Japanese Korean)的三种文字占用了Unicode中0x3000到0x9FFF的部分。 Unicode目前普遍采用的是UCS-2规范编码,具体的,其用两个字节来编码一个字符。比如汉字“经”的编码是0x7ECF。由于字符编码一般用十六进制来表示,为了与十进制区分,十六进制以0x开头,0x7ECF转换成十进制就是32463,UCS-2用两个字节来编码字符,两个字节就是16位二进制,2的16次方等于65536,所以UCS-2最多能编码65536个字符。
在本实施方式中,在所述客户端可以存储有统一字符集,所述统一字符集中可以包括字符、以及与字符相对应的统一码。此外,所述统一字符集也可以存储于服务器端,客户端可以通过通信模块与所述服务器端建立通信后,获取所述统一字符集中的相关字符以及与字符相对应的统一码。所述每个第一预定字符可以唯一对应有一个统一码。具体的,所述第一预定字符可以为所述名称字符串的首字符,也可以为所述名称字符串中的首字符与第二字符的结合,也可以单独为第二字符,或者还可以是由用户预先指定的所述名称字符串中的某个预定字符。当然,所述第一预定字符并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在本实施方式中,由于所述第一预定字符可以唯一对应有一个统一码,根据所述第一预定字符,通过查找的方式或者通过所述第一预定字符与统一码的映射关系,可以获得所述名称字符串中第一预定字符的统一码。具体的,例如可以通过将所述第一预定字符转换为十六进制数字,然后在所述统一字符集中进行查找。若查找到在所述统一字符集中相同的十六进制数字,则可以唯一确定所述第一预定字符的统一码。或者,根据所述第一预定字符,通过字符与十六进制数字的映射关系,在所述统一字符集中进行查找,以获取所述名称字符串中第一预定字符的统一码。当然,所述获取的方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
步骤S24:确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
在本实施方式中,可以利用编程语言API(Application Programming Interface,应用程序编程接口)中的语言排序机制,根据语言的发音、拼写等特点,将所述第一预定字符的统一码进行划分,设定至少一个预定编码段。其中每个所述预定编码段可以对应设置为一个预定类别,或者多个预定编码段可以对应设置为同一个预定类别,进而将多种语言的文字进行归类至有限的类别。
由于一般情况下,统一编码并不按语言的发音等规律进行排列。为了使得预定编码 段的划分符合用户语言习惯,可以利用排序器,例如Java中的collator,根据语言的发音等特点,将字符进行排序。进一步的,将排序后的字符对应的统一码划分为至少一个预定的编码段。具体的,所述排序器collator在进行排序时,能够根据语言的特点,比较两个字符的先后顺序,进而能够将统一字符集中的字符按照语言特点进行分段排序。例如,在Java语言标准库里面,对于中文,可以根据其发音,所有“啊”(a)发音的第一预定字符排在所有“吧”(ba)发音的第一预定字符的前面。所有“吧”(ba)发音的第一预定字符排在所有“擦”(ca)发音的第一预定字符的前面。以此类推,根据字符的发音将名称数据中的第一预定字符分为“a”至“z”26个预定编码段。
在本实施方式中,所述预定类别为多个名称字符串的集合。所述预定类别在数据结构上可以通过队列、数据栈、数组等方式实现。例如,当所述预定类别的数据结构方式为数组的数据结构方式时,可以将第一预定字符发音属于同一预定编码段的名称字符串存储于同一个数组下,以形成一个预定类别。例如,将“陈一”、“成二”、“蔡三”、“程四”存储于同一个数组下,以形成一个预定类别。当然,形成所述预定类别的数据结构方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在一个具体的实施方式中,为了使得预定编码段的划分符合用户语言习惯,可以利用Java中的collator根据语言的发音、拼写等特点将第一预定字符对应的统一码按序排列,并分为至少一个预定编码段。当需要判断某个第一预定字符具体属于对应的某个预定编码段时,可以将所述第一预定字符与所述预定编码段的段首和/或段尾的统一码对应的字符比较。若该第一预定字符的顺序为位于所述预定编码段的段首字符以后,而位于段尾字符以前,或者位于下一个段首字符之前的顺序,则可以确定该第一预定字符属于此预定编码段。
此外,在划分预定编码段时,预定编码段段首和/或段尾统一码与预定类别之间可以形成有对应关系。请参看上述表1。当确定了预定编码段后,可以根据所述预定编码段与预定类别的对应关系,确定所述名称字符串的预定类别。
在本实施方式中,利用了现有的Java中的collator,实现了根据语言的发音、拼写等特点进行比较排序。当然,具体的排序方式并不限于上述描述,所属领域技术人员在本申请的技术精髓启发下,还可以作出其他方式的变更,例如选择现有的其他排序器或者根据排序功能自行进行编写等,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在本实施方式中,操作系统中可以集成有统一字符集,即每个字符对应的统一码可 以从操作系统中获得。所述客户端中可以仅存储所述预定编码段的段首统一码,如此可以实现在统一码分段时,通过段首统一码,区分不同的预定编码段;或者是所述预定编码段的段首统一码和所述预定编码段的段尾段尾统一码。具体的,例如当所述预定编码段为26个时,对应的,所述客户端存储的为26个预定编码段的统一码,共26个;或者预定编码段的段首统一码和段尾的统一码,共52个。因此,相对于客户端来说,所需占用的存储空间较小。
在另一个具体的实施方式中,由于一般情况下,统一编码并不按每种语言的发音等规律进行排列。为了使得预定编码段的划分符合用户语言习惯,可以将第一预定字符对应的统一码根据语言的发音等特点对应一个索引值,根据所述索引值进行排序,进而可以将所述根据索引值排序后的统一码分为至少一个预定的编码段。
例如,在Java语言标准库里面,对于中文,可以根据其发音,给所述第一预定字符对应的统一码一个对应的索引值。所述索引值可以看成一个数字。比如“啊”(a),其索引值可能是1;而“吧”(ba),其索引值可能是10,那么所有“啊”(a)至“吧“(ba)之间发音的汉字,它的索引值在1-10之间。
请参阅上述表2。针对每个预定编码段可以对应有一个预定类别。例如,编码段段首索引值为[1,10)对应的预定类别为A类别,即实现了将字符发音的首字母为a的字,例如“啊”、“安”等归至A类别中。编码段段首索引值为[10,20)对应的预定类别为B类别,即实现了将字符发音的首字母为b的字,例如“吧”、“被”等归至B类别中。以此类推,编制Z类别。从而将整个中文编码块划分至26个编码段,对应的26个预定类别。在本实施方式中,通过将所述第一预定字符的统一码的索引值与预设编码段的段首索引值和/或段尾索引值进行查找比较的方式,可以确定所述第一预定字符对应的预定编码段。由于所述预定编码段对应有预定类别,因此可以将所述名称字符串划分至所述预定编码段对应的预定类别。
在一个具体的实施方式中,对于某个第一预定字符,将其对应的统一码的索引值和预定编码段首索引值和/或段尾索引值进行比较,进而确定所述第一预定字符所对应的预定编码段以及预定类别。请结合参阅表2。例如,如果某个第一预定字符的统一码对应的索引值大于等于所述1而小于10,则可以确定所述第一预定字符所对应的预定编码段和预定类别,进而可以将所述第一预定字符对应的所述名称字符串划分至所述预定编码段的A类别中。
在本实施方式中,操作系统中可以集成有统一字符集,即每个第一预定字符对应的统一码可以从操作系统中获得。所述客户端中可以仅存储的所述预定编码段的段首索引值,如此可以实现在统一码分段时,通过段首索引值和/或段尾索引值,区分不同的预定 编码段。具体的,例如当所述预定编码段为26个时,对应的,所述客户端存储的为26个预定编码段的索引值,共26个;或者预定编码段的段首索引值和段尾的索引值,共52个。因此,相对于客户端来说,所需占用的存储空间较小。
当然,所述确定所述第一预定字符的统一码对应的预定编码段的方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
本申请实施方式所述名称列表处理方法基于名称字符串中第一预定字符的统一码,将所述名称字符串划分至所述预定编码段对应的预定类别,实现了利用统一码与名称字符串的一一对应关系,将多个国家的语言在同一名称列表中进行分类。使用时,可将所述预定类别的类别标识和与所述预定类别对应的名称字符串发送至客户端。当在所述客户端的显示界面中对应展示时,用户通过展示的类别标识可以快速定位至名称列表中的具体名称字符串,查找对对应的目标名称。所述名称列表展示方法在实现的过程中摆脱了传统的预设字典的方式,利用已有的统一码进行归类排序,减少了大量的人力劳动,节约了人力资源成本。
可以理解,服务器端存储的统一码,以作为跨地域或数据平台数据流动时,对文字数据进行编码解码的依据。在本申请中,通过利用统一码对名称列表进行分类排序,可以实现没有在现有客户端中引入的额外字典的方式,进行名称列表分类排序。具体的,例如可以不引入类似pinyin4j字典,实现节省了存储空间。
在一个实施方式中,所述名称列表处理方法还可以包括:将所述预定类别的类别标识,和与所述预定类别对应的名称字符串发送至客户端进行对应展示。
在本实施方式中,所述预定类别可以对应有类别标识。服务器与客户端之间可以件建立通信。例如,在服务器端接收到客户端的发送请求时,或者是服务器端自动更新后等情况下,所述服务器端可以将所述预定类别的类别标识和与所述预定类别对应的名称字符串发送给相应的客户端。所述客户端接收到后,可以将所述预定类别的类别标识和与所述预定类别对应的名称字符串一同展示出来,形成便于用户查找的名称列表。
具体的,所述类别标识可以为根据用户语言习惯,便于用户识别的标识。例如,对于使用中文的用户,其类别标识可以为所述第一预定字符发音的首字母。对于使用英文的用户,其类别标识可以为所述第一预定字符的首字母。当然,对于其他语言种类而言,其类别标识也可以为所述第一预定字符发音的音标或者是拼写的首字母等等,此处不再重复举例。
在本实施方式中,所述预定类别对应的名称字符串可以全部显示或者显示部分所述名称字符串或者不显示所述名称字符串。例如,所述预定类别包括有很多名称字符串, 此时,可设定在所述预定类别对应的类别标记下最多显示预定个数的名称字符串,例如可以设定最多显示3个名称字符串。此外,也可以设定在所述显示界面只显示所述类别标记。所述类别标记的相应位置检测到用户点击产生的信号时,可进一步向用户展示单个类别标记下的名称字符串。
在一个实施方式中,所述名称数据的具体内容可以包括,人名、公司名称、账户名称、地点名称中的至少一个。
在本实施方式中,所述名称数据可以包括至少一个名称字符串。所述名称字符串具体的可以为人名、公司名称、账户名称、地点名称等等,本申请在此并不作具体的限定。当所述名称数据的具体内容为人名时,所述人名可以为用户的真实姓名,也可以为用户给自己设定的昵称等。
所述名称数据的种类可以为只包括一个名称种类,例如只包括人名;也可以同时包括多个名称种类,例如同时包括人名、公司名称、账户名称等等。当所述名称数据同时包括多个名称种类时,若多个名称种类之间有对应关系,则可以将所述多个名称种类进行划分,例如可以设定多个名称种类中的一种的名称字符串为主名称,设定多个名称种类中的其他名称作为从属名称。展示时,可以只将所述主名称在显示界面中进行显示,其他与主名称相对应的从属名称,可以与所述主名称对应存储。当主名称所在的位置检测到点击触发的电信号时,可以显示所述从属名称。另外,当所述从属名称在显示界面显示后,还可以对从属名称进行复制、编辑等操作。
例如,名称数据具体内容为“甲公司小王中国银行账户XX”,设定人名为主名称,其他公司、账户名称为从属名称。展示时,将所述主名称“小王”在显示界面中进行显示,其他与“小王”相对应的从属名称“甲公司”、“中国银行账户XX”,可以与所述主名称“小王”对应存储。当点击所述主名称“小王”时,可以相应地显示所述从属名称“甲公司”、“中国银行账户XX”,进一步地,可以对从属名称进行复制、编辑等操作。
在一个实施方式中,所述第一预定字符可以为所述名称字符串的首字符。
在本实施方式中,所述名称字符串具体的可以为人名、公司名称、账户名称、地点名称等等。所述名称字符串可以包括一个或者多个字符。而所述名称字符串的首字符很多情况下能够典型地代表整个名称字符串。例如,当所述名称字符串为人名时,其首字符可以为姓氏,当所述名称字符串为公司名称时,其首字符也可为公司名称第一个字。当将所述第一预定字符设定为所述名称字符串的首字符,后续以所述名称字符串进行分类形成的名称列表符合用户的使用习惯,有助于用户基于形成的名称列表进行查询。
在一个实施方式中,可以通过Java中的collator确定所述第一预定字符的统一码对 应的预定编码段。
在本实施方式中,由于一般情况下,统一编码并不按语言的发音等规律进行排列。为了使得预定编码段的划分符合用户语言习惯,可以利用排序器,例如Java中的collator,根据语言的发音等特点,将字符进行排序。进一步的,将排序后的每个字符对应的统一码划分为至少一个预定的编码段。具体的,所述排序器collator在进行排序时,能够根据语言的特点,比较两个字符的先后顺序,进而能够将统一字符集中的字符按照语言特点进行分段排序。例如,在Java语言标准库里面,对于中文,可以根据其发音,所有“啊”(a)发音的第一预定字符排在所有“吧”(ba)发音的第一预定字符的前面。所有“吧”(ba)发音的第一预定字符排在所有“擦”(ca)发音的第一预定字符的前面。以此类推,根据第一预定字符的发音将名字数据中的中文字符分为“a”至“z”26个预定编码段。
在本实施方式中,利用了现有的Java中的collator,实现了根据语言的发音、拼写等特点进行比较排序。当然,具体的排序方式并不限于上述描述,所属领域技术人员在本申请的技术精髓启发下,还可以作出其他方式的变更,例如选择现有的其他排序器或者根据排序功能自行进行编写等,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在一个具体的实施方式中,例如,新增加了一个名称字符串,需要判断所述新增加的名称字符串的第一预定字符具体属于对应的某个预定编码段时,可以将所述第一预定字符与所述预定编码段的段首和/或段尾的统一码对应的字符比较。若该第一预定字符的顺序为位于所述预定编码段的段首字符以后而位于段尾字符以前的顺序,或者位于下一个段首字符之前的顺序,则可以确定该第一预定字符属于此预定编码段。
在一个实施方式中,至少一个所述预定编码段组成一个语言编码块;其中,不同所述语言编码块具有至少部分相同的预定类别。相应的,在将预定编码划分预定类别的步骤中包括:在属于不同所述语言编码块的名称字符串所处的预定编码段对应的预定类别相同时,将所述名称字符串同划分至所述预定类别的预设数据集中。
在本实施方式中,在统一码中针对不同的语言种类,可以设定有不同的语言编码块。例如,可以包括中日韩文编码块、英文编码块、俄文编码块、阿拉伯文编码块等等。其中,不同所述语言编码块可以具有至少部分相同的预定类别。
在一个具体的实施方式中,请结合表1所示。对于中文编码块而言,利用Java中的collator可以根据语言的发音等特点对应设定多个预定编码段。例如,可以根据所述第一预定字符的发音特点,将所述第一预定字符对应的统一码分为“a”至“z”26个预定编码段。针对每个预定编码段可以对应有一个预定类别。例如所述“a”至“z”26个预定编码段分别对应的预定类别为A至Z26个类别。
对于英文编码块而言,请参看上述表3,利用Java中的collator可以根据字符拼写特点对应设置多个预定编码段。此外,当所述英文编码块的统一码本身排列顺序为按照A至Z的顺序排列时,也可以根据所述统一码本身排列顺序设置多个预定编码段。例如,可以根据第一预定字符的拼写特点,将所述第一预定字符分为A到Z26个预定编码段。针对每个预定编码段可以对应有一个预定类别。例如,第一预定编码段段首统一码97对应的预定类别为A类别,即将首字母为A的名称字符串归至预定类别A中。第二预定编码段段首统一码98对应的预定类别为B类别,即将首字母为B的名称字符串归至预定类别B中。依次类推,编制预定类别Z。从而将整个英文编码块划分至26个编码段,对应的26个预定类别。
请参阅上述表4,由于所述中文编码块的预定类别和英文编码块的预定类别从A至Z相同,因此可以将属于相同预定类别的名称字符串划分至所述预定类别中,形成所述预设数据集,进而实现中英文混合排序。
此外,所述英文编码块利用Java中的collator根据字符拼写特点对应设置多个预定编码段时,对应的每个预定编码段的段首或段尾可以设置有索引值,当所述英文编码块对应的索引值与所述中文编码块对应的索引值相同时,则可以确定所述索引值相同的名称字符串属于同一类别。
在另一个实施方式中,请参阅上述表5,所述中文编码块与所述英文编码块也可以不划分至相同的所述预定类别中,即各自按照各自的预定类别进行划分,不进行混排。例如,英文编码块、中午编码块按照统一码的先后顺序进行分类排列,可以整体上按照英文编码块在前,中文编码块在后的顺序进行分类和排列。不同的编码块具体的分类的标记可以相同,也可以不同,本申请在此并不作具体的限定。
当然,所述语言编码块的种类并不限于举例的中文语言编码块、英文语言编码块,数量也不限于两种语言编码块。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
请参阅图8,在一个实施方式中,在确定预定编码段的步骤S24中可以包括如下步骤。
步骤S240:确定所述名称字符串中第一预定字符的统一码对应的语言编码块。
步骤S242:在所述语言编码块中,确定所述第一预定字符的统一码所在的预定编码段。
在本实施方式中,在确定所述预定编码段时,可以先根据所述名称字符串中第一预定字符的统一码确定其对应的语言编码块。在基于确定的语言编码块的基础上,进一步 确定所述第一预定字符的统一码所在的预定编码段。
具体的,在统一码中针对不同的语言种类,可以设定有不同的语言编码块。每个语言编码块的首字符的统一码以及字符的个数都是已知的。在确定所述名称字符串中的第一预定字符的统一码对应的语言编码块时,可将所述第一预定字符的统一码与每个语言编码块的首字符统一码进行比较。当所述第一预定字符的统一码大于等于某个语言编码块的首字符统一码,且其与某个语言编码块的首字符统一码的差值小于所述语言编码块的个数时,则可以确定所述第一预定字符为所述语言编码块中的字符。
当确定了所述名称字符串中第一预定字符的统一码对应的语言编码块后,可以将所述第一预定字符的统一码在所述确定的语言编码块内进行查找,以便快速确定所述第一预定字符的统一码所在的预定编码段。具体的,所述确定的语言编码块一般地包括有多个语言编码段,每个语言编码段至少对应有一个统一码。确定所述第一预定字符的统一码时,可以将所述第一预定字符的统一码与所述语言编码段的统一码进行对比,若有相同,则可以确定所述第一预定字符的统一码所在的预定编码段。
在一个具体的实施方式中,如表3所示,英文语言编码块可以包括26个语言编码段。在这26个语言编码段中,每个预定编码段对应有一个统一码。当所述预定编码段对应的统一码为按照大小顺序排列时,可以将所述一一预定字符的统一码与每个预定编码段的统一码进行比较,若有相同,则可以确定所述第一预定字符的统一码所在的预定编码段。
在一个实施方式中,所述方法还可以包括:在没有与所述名称字符串中第一预定字符的统一码对应的语言编码块时,将所述名称字符串划分至预设指定类别。
在本实施方式中,可以预先在客户端设置有预设指定类别,用于归类没有查找到对应语言编码块的名称字符串。所述预设指定类别可以为排列于所述预定类别后的一个单独的类别,其可以用一个特殊的标记进行展示。例如可以用“#”表示其下包括的是不属于统一码所包括的语言编码块的预设指定类别。
在本实施方式中,在没有与所述名称字符串中第一预定字符的统一码对应的语言编码块时,说明所述名称字符串中第一预定字符不属于统一码中已有的语言或符号。此时,可以将其划分至预设指定类别。
在一个实施方式中,所述客户端在展示类别标识和名称字符串的步骤中还可以包括:
将属于同一所述预定类别的所述名称字符串按照所述第一预定字符的统一码的编码顺序排列。
在本实施方式中,所述名称字符串的第一预定字符可以对应有唯一的统一码。将属于同一预定类别的所述名称字符串进行排序时,可以按照所述第一预定字符的统一码的编码顺序排列。
具体的,由于所述第一预定字符的统一码的编码本身可为一个十六进制的数字,可以依照数字的大小顺序进行排列。例如可以按照所述第一预定字符的统一码由小到大的顺序对所述同一预定类别下的名称字符串进行排列。
在一个具体的实施方式中,例如对于C类别下的名称字符串:第一预定字符,例如“陈一”、“成二”、“蔡三”,可以根据其第一预定字符“蔡”、“陈”、“成”统一码由小到大的顺序的特点,将所述C类别下的名称字符串排列为“蔡三”、“陈一”、“成二”,实现了同一类别下,名称字符串的排序。
请参阅图9,在一个实施方式中,在所述第一预定字符的统一码相同时,所述方法还可以包括如下步骤。
步骤S261:获取所述名称字符串中第二预定字符。
步骤S263:将属于同一所述预定类别的所述名称字符串按照所述第二预定字符的统一码的编码顺序排列。
在本实施方式中,所述第一预定字符的统一码相同的情况,具体可表示所述名称字符串包含有相同的第一预定字符。
在本实施方式中,所述第二预定字符可以为所述名称字符串中,除了所述第一预定字符之外的一个字符或多个字符的组合。具体的,由于所述第二预定字符的统一码的编码本身为一个十六进制的数字,可以依照数字的大小顺序进行排列。例如可以按照所述第二预定字符的统一码由小到大的顺序对所述同一预定类别下的名称字符串进行排列。
在一个具体的实施方式中,当所述名称字符串都为姓名时,其可以表示包含多个姓氏相同的名字。例如,在W类别下,名称字符串包括:“王一”、“王二”、“王三”。分别获取所述第一预定字符的统一码时,发现其对应的统一码均相同。此时,进一步的可以获取除了所述第一预定字符之外第二字符的统一码。具体可获取的第二预定字符分别为“一”、“二”、“三”的统一码。在统一码字符集中,所述第二预定字符“一”、“二”、“三”的统一码的排序为“二”、“三”、“一”。对应的,可以将与所述第二预定字符相对应的名称字符串排列为“王二”、“王三”、“王一”,从而实现了,当第一预定字符相同时,对同一类别下的名称字符串排序。
请参阅图10,在一个实施方式中,在展示类别标识和名称字符串的步骤S26中还可以包括如下步骤。
步骤S260:获取所述名称字符串中第二预定字符的统一码。
步骤S262:将属于同一所述预定类别的所述名称字符串按照所述第二预定字符的统一码的编码顺序排列。
在本实施方式中,所述第二预定字符可以为所述名称字符串中,除了所述第一预定 字符之外的一个字符或多个字符的组合。具体的,由于所述第二预定字符的统一码的编码本身为一个十六进制的数字,可以依照数字的大小顺序进行排列。例如可以按照所述第二预定字符的统一码由小到大的顺序对所述同一预定类别下的名称字符串进行排列。
在一个具体的实施方式中,当所述名称字符串为姓名时,例如在C类别下,名称字符串包括:“陈一”、“成二”、“蔡三”。上述名称字符串的第一预定字符为“蔡”、“陈”、“成”;第二预定字符为“一”、“二”、“三”。当名称字符串直接按照所述第二预定字符的统一码进行排序时,由于第二预定字符为“一”、“二”、“三”对应的统一码由小到大的排序为“二”、“三”、“一”,因此,C类别下对应的名称字符串的排序为“成二”、“蔡三”、“陈一”。本实施方式通过直接根据第二预定字符的统一码对与其对应的名称字符串进行排序,对于第一预定字符相同的情况来说,减少了不必要的第一预定字符统一码比较,能够提高排序的效率。
在一个实施方式中,在获取名称数据的步骤S20中还可以包括:接收用户指定的第一预定字符。
在本实施方式中,在获取名称数据时,可以接收用户指定的第一预定字符。具体的,所述指定的第一预定字符可为所述名称字符串中的一个字符或多个字符的组合。所述指定的第一预定字符用于在所述名称字符串归至预定类别后,将属于同一预定类别名称字符串进行排序。所述指定的第一预定字符具体的可以为名称的姓氏列表,可以为地点的国别列表,可以为账户的银行类别列表等等,本申请在此并不作具体限定。
在一个具体的实施方式中,所述指定的第一预定字符为一个姓氏列表时,所述姓氏列表中的姓氏可以按照顺序排列好。例如姓氏列表中包括“Brown”、“Smith”、“White”等,其按照先后顺序排列为:“Brown”、“Smith”、“White”。当所述名称数据包括的名称字符串为:“Alice Brown”、“Ada Smith”、“Alma White”。根据所述指定排序字符中“Brown”、“Smith”、“White”的排列顺序,上述名称字符串排列先后顺序为“Alice Brown”、“Ada Smith”、“Alma White”。
请参阅图11,本申请实施方式还提供一种服务器200,其可以包括:数据获取模块20、统一码获取模块22、预定类别划分模块24。
数据获取模块10,用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串。
在本实施方式中,服务器可以是具有数据处理能力和网络通讯功能的设备,例如个人计算机、服务器计算机、手持设备或便携式设备、平板型设备、多处理器装置、包括以上任何装置或设备的分布式计算环境。当然,服务器也可以包括运行于上述设备中的软件。
在本实施方式中,所述名称数据可以预存于所述服务器中。当用户需要进行电子商务活动时,可以通过数据通讯模块从服务器端进行下载所述名称数据。或者,所述名称数据也可以由用户以输入的方式临时建立获得。此外,获取所述名称数据的方式还可以为其他方式,例如,接受其他客户端发送的名称数据等等,本申请在此并不作具体限定。
在本实施方式中,所述名称数据可以为用户在进行即时通讯、电子商务等交互活动时需要使用的名称集合。所述名称具体的可以为:姓名、昵称、公司名称、账户名称、地点名称等等,本申请在此并不作具体的限定。
在本实施方式中,所述名称数据包括有至少一个名称字符串。所述名称字符串可以为所述姓名、昵称、公司名称、账户名称、地点名称中的一个具体名字。具体的,所述名称字符串可以为具体的一个姓名,例如“王二小”。所述名称字符串也可以为具体的一个账户名称,例如某个银行的银行卡号:“622848440637874213”。所述名称字符串的形式具体可以为某种语言的至少一个字、词、词组等,或者可以为一种符号、数字、字母或者多种结合等,当然所述名称字符串还可以为其他形式,本申请在此并不作具体的限定。
统一码获取模块12,用于获取所述名称字符串中第一预定字符的统一码。
在本实施方式中,统一码用于为每种语言中的字符设定了统一并且唯一编码,以满足跨语言、跨平台进行文本转换、处理的要求。所述统一码具有一个字符集,其可以用预定进制的数字来映射这些字符,使得在所述统一码的字符集中的字符与每种语言中的字符存在一一对应关系。即一个统一码唯一对应有一个字符。所述预定进制的数字可以为十六进制,二进制等等,此处不作一一陈述。具体的,例如Unicode码为一种典型的统一码,其用十六进制数字0-0x10FFFF来映射各种语言的字符。如,中、日、韩CJK(Chinese Japanese Korean)的三种文字占用了Unicode中0x3000到0x9FFF的部分。Unicode目前普遍采用的是UCS-2规范编码,具体的,其用两个字节来编码一个字符。比如汉字“经”的编码是0x7ECF。由于字符编码一般用十六进制来表示,为了与十进制区分,十六进制以0x开头,0x7ECF转换成十进制就是32463,UCS-2用两个字节来编码字符,两个字节就是16位二进制,2的16次方等于65536,所以UCS-2最多能编码65536个字符。
在本实施方式中,在所述客户端可以存储有统一字符集,所述统一字符集中可以包括字符、以及与字符相对应的统一码。此外,所述统一字符集也可以存储于服务器端,客户端可以通过通信模块与所述服务器端建立通信后,获取所述统一字符集中的相关字符以及与字符相对应的统一码。所述每个第一预定字符可以唯一对应有一个统一码。具体的,所述第一预定字符可以为所述名称字符串的首字符,也可以为所述名称字符串中 的首字符与第二字符的结合,也可以单独为第二字符,或者还可以是由用户预先指定的所述名称字符串中的某个预定字符。当然,所述第一预定字符并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在本实施方式中,由于所述第一预定字符可以唯一对应有一个统一码,根据所述第一预定字符,通过查找的方式或者通过所述第一预定字符与统一码的映射关系,可以获得所述名称字符串中第一预定字符的统一码。具体的,例如可以通过将所述第一预定字符转换为十六进制数字,然后在所述统一字符集中进行查找。若查找到在所述统一字符集中相同的十六进制数字,则可以唯一确定所述第一预定字符的统一码。或者,根据所述第一预定字符,通过字符与十六进制数字的映射关系,在所述统一字符集中进行查找,以获取所述名称字符串中第一预定字符的统一码。当然,所述获取的方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
预定类别划分模块14,用于在本实施方式中,可以利用编程语言API(Application Programming Interface,应用程序编程接口)中的语言排序机制,根据语言的发音、拼写等特点,将所述第一预定字符的统一码进行划分,设定至少一个预定编码段。其中每个所述预定编码段可以对应设置为一个预定类别,或者多个预定编码段可以对应设置为同一个预定类别,进而将多种语言的文字进行归类至有限的类别。
由于一般情况下,统一编码并不按语言的发音等规律进行排列。为了使得预定编码段的划分符合用户语言习惯,可以利用排序器,例如Java中的collator,根据语言的发音等特点,将字符进行排序。进一步的,将排序后的字符对应的统一码划分为至少一个预定的编码段。具体的,所述排序器collator在进行排序时,能够根据语言的特点,比较两个字符的先后顺序,进而能够将统一字符集中的字符按照语言特点进行分段排序。例如,在Java语言标准库里面,对于中文,可以根据其发音,所有“啊”(a)发音的第一预定字符排在所有“吧”(ba)发音的第一预定字符的前面。所有“吧”(ba)发音的第一预定字符排在所有“擦”(ca)发音的第一预定字符的前面。以此类推,根据字符的发音将名称数据中的第一预定字符分为“a”至“z”26个预定编码段。
在本实施方式中,所述预定类别为多个名称字符串的集合。所述预定类别在数据结构上可以通过队列、数据栈、数组等方式实现。例如,当所述预定类别的数据结构方式为数组的数据结构方式时,可以将第一预定字符发音属于同一预定编码段的名称字符串存储于同一个数组下,以形成一个预定类别。例如,将“陈一”、“成二”、“蔡三”、“程四”存储于同一个数组下,以形成一个预定类别。当然,形成所述预定类别的数据结 构方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在一个具体的实施方式中,为了使得预定编码段的划分符合用户语言习惯,可以利用Java中的collator根据语言的发音、拼写等特点将第一预定字符对应的统一码按序排列,并分为至少一个预定编码段。当需要判断某个第一预定字符具体属于对应的某个预定编码段时,可以将所述第一预定字符与所述预定编码段的段首和/或段尾的统一码对应的字符比较。若该第一预定字符的顺序为位于所述预定编码段的段首字符以后,而位于段尾字符以前,或者位于下一个段首字符之前的顺序,则可以确定该第一预定字符属于此预定编码段。
此外,在划分预定编码段时,预定编码段段首和/或段尾统一码与预定类别之间可以形成有对应关系。请参看上述表1。当确定了预定编码段后,可以根据所述预定编码段与预定类别的对应关系,确定所述名称字符串的预定类别。
在本实施方式中,利用了现有的Java中的collator,实现了根据语言的发音、拼写等特点进行比较排序。当然,具体的排序方式并不限于上述描述,所属领域技术人员在本申请的技术精髓启发下,还可以作出其他方式的变更,例如选择现有的其他排序器或者根据排序功能自行进行编写等,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
在本实施方式中,操作系统中可以集成有统一字符集,即每个字符对应的统一码可以从操作系统中获得。所述客户端中可以仅存储所述预定编码段的段首统一码,如此可以实现在统一码分段时,通过段首统一码,区分不同的预定编码段;或者是所述预定编码段的段首统一码和所述预定编码段的段尾段尾统一码。具体的,例如当所述预定编码段为26个时,对应的,所述客户端存储的为26个预定编码段的统一码,共26个;或者预定编码段的段首统一码和段尾的统一码,共52个。因此,相对于客户端来说,所需占用的存储空间较小。
在另一个具体的实施方式中,由于一般情况下,统一编码并不按每种语言的发音等规律进行排列。为了使得预定编码段的划分符合用户语言习惯,可以将第一预定字符对应的统一码根据语言的发音等特点对应一个索引值,根据所述索引值进行排序,进而可以将所述根据索引值排序后的统一码分为至少一个预定的编码段。
例如,在Java语言标准库里面,对于中文,可以根据其发音,给所述第一预定字符对应的统一码一个对应的索引值。所述索引值可以看成一个数字。比如“啊”(a),其索引值可能是1;而“吧”(ba),其索引值可能是10,那么所有“啊”(a)至“吧“(ba) 之间发音的汉字,它的索引值在1-10之间。
请参阅上述表2。针对每个预定编码段可以对应有一个预定类别。例如,编码段段首索引值为[1,10)对应的预定类别为A类别,即实现了将字符发音的首字母为a的字,例如“啊”、“安”等归至A类别中。编码段段首索引值为[10,20)对应的预定类别为B类别,即实现了将字符发音的首字母为b的字,例如“吧”、“被”等归至B类别中。以此类推,编制Z类别。从而将整个中文编码块划分至26个编码段,对应的26个预定类别。
在本实施方式中,通过将所述第一预定字符的统一码的索引值与预设编码段的段首索引值和/或段尾索引值进行查找比较的方式,可以确定所述第一预定字符对应的预定编码段。由于所述预定编码段对应有预定类别,因此可以将所述名称字符串划分至所述预定编码段对应的预定类别。
在一个具体的实施方式中,对于某个第一预定字符,将其对应的统一码的索引值和预定编码段首索引值和/或段尾索引值进行比较,进而确定所述第一预定字符所对应的预定编码段以及预定类别。请结合参阅上述表2。例如,如果某个第一预定字符的统一码对应的索引值大于等于所述1而小于10,则可以确定所述第一预定字符所对应的预定编码段和预定类别,进而可以将所述第一预定字符对应的所述名称字符串划分至所述预定编码段的A类别中。
在本实施方式中,操作系统中可以集成有统一字符集,即每个第一预定字符对应的统一码可以从操作系统中获得。所述客户端中可以仅存储的所述预定编码段的段首索引值,如此可以实现在统一码分段时,通过段首索引值和/或段尾索引值,区分不同的预定编码段。具体的,例如当所述预定编码段为26个时,对应的,所述客户端存储的为26个预定编码段的索引值,共26个;或者预定编码段的段首索引值和段尾的索引值,共52个。因此,相对于客户端来说,所需占用的存储空间较小。
当然,所述确定所述第一预定字符的统一码对应的预定编码段的方式并不限于上述描述。所属领域技术人员在本申请的技术精髓启示下,还可能做出其他方式的变更,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请保护范围内。
本申请服务器可以是本申请名称列表处理方法的一种硬件实施方式,可以实现本申请名称列表处理方法实施方式并达到方法实施方式的技术效果。
在一个实施方式中,至少一个所述预定编码段组成一个语言编码块;其中,不同所述语言编码块具有至少部分相同的预定类别;相应的,所述预定类别划分模块用于:在属于不同所述语言编码块的名称字符串所处的预定编码段对应的预定类别相同时,将所述名称字符串同划分至所述预定类别中,以实现多种语言的混排。
请参阅图12,在一个实施方式中,所述预定类别划分模块24包括:语言编码块确定单元240、预定编码段确定单元242。
语言编码块确定单元240,用于确定所述名称字符串中第一预定字符的统一码对应的语言编码块;
预定编码段确定单元242,用于在所述语言编码块中,确定所述第一预定字符的统一码所在的预定编码段。
本申请实施方式还提供一种服务器,其可以包括:处理器。
所述处理器用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;并用于获取所述名称字符串中第一预定字符的统一码;确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
在本实施方式中,所述处理器可以按任何适当的方式实现。例如,处理器可以采取例如微处理器或处理器以及存储可由该(微)处理器执行的计算机可读程序代码(例如软件或固件)的计算机可读介质、逻辑门、开关、专用集成电路(Application Specific Integrated Circuit,ASIC)、可编程逻辑控制器和嵌入微控制器的形式等等。本申请并不作限定。
本申请服务器可以是本申请名称列表处理方法的一种硬件实施方式,可以实现本申请名称列表处理方法实施方式并达到方法实施方式的技术效果。
本说明书中的上述各个实施方式均采用递进的方式描述,各个实施方式之间相同相似部分相互参照即可,每个实施方式重点说明的都是与其他实施方式不同之处。尤其对于客户端实施方式而言,由于其处理器执行的工作基本相似于方法实施方式,所以描述的比较简单,相关之处参见方法实施方式部分说明即可。
虽然通过实施方式描绘了本申请,在本申请技术精髓启示下,本领域技术人员可能对上述多个实施方式之间进行组合,也可以对本申请的实施方式进行变化,但只要其实现的功能和效果与本申请相同或相似,均应涵盖于本申请的保护范围内。

Claims (23)

  1. 一种名称列表展示方法,其特征在于,其包括:
    获取名称数据;其中,所述名称数据包括有至少一个名称字符串;
    获取所述名称字符串中第一预定字符的统一码;
    确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别;
    在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串。
  2. 如权利要求1所述的方法,其特征在于,在展示类别标识和名称字符串的步骤中还包括:
    将属于同一所述预定类别的所述名称字符串按照所述第一预定字符的统一码的编码顺序排列。
  3. 如权利要求2所述的方法,其特征在于,在所述第一预定字符的统一码相同时,所述方法还包括:
    获取所述名称字符串中第二预定字符;
    将属于同一所述预定类别的所述名称字符串按照所述第二预定字符的统一码的编码顺序排列。
  4. 如权利要求1所述的方法,其特征在于,在展示类别标识和名称字符串的步骤中还包括:
    获取所述名称字符串中第二预定字符的统一码;
    将属于同一所述预定类别的所述名称字符串按照所述第二预定字符的统一码的编码顺序排列。
  5. 如权利要求1所述的方法,其特征在于,在获取名称数据的步骤中还包括:接收用户指定的第一预定字符。
  6. 一种客户端,其特征在于,其包括:
    数据获取模块,用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;
    统一码获取模块,用于获取所述名称字符串中第一预定字符的统一码;
    预定类别划分模块,用于确定所述第一预定字符的统一码对应的预定编码段,将所 述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别;
    显示模块,用于在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串。
  7. 一种客户端,其特征在于,其包括:处理器、显示器,
    所述处理器用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;并用于获取所述名称字符串中第一预定字符的统一码;确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别;控制所述显示器在显示界面中对应展示所述预定类别的类别标识,和与所述预定类别对应的名称字符串。
  8. 一种名称列表处理方法,其特征在于,其包括:
    获取名称数据;其中,所述名称数据包括有至少一个名称字符串;
    获取所述名称字符串中第一预定字符的统一码;
    确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
  9. 如权利要求8所述的方法,其特征在于,所述方法还包括:将所述预定类别的类别标识,和与所述预定类别对应的名称字符串发送至客户端进行对应展示。
  10. 如权利要求8所述的方法,其特征在于:所述名称数据包括:人名、公司名称、账户名称、地点名称中的至少一个。
  11. 如权利要求8所述的方法,其特征在于:所述第一预定字符为所述名称字符串的首字符。
  12. 如权利要求8所述的方法,其特征在于:通过Java中的collator确定所述第一预定字符的统一码对应的预定编码段。
  13. 如权利要求8所述的方法,其特征在于,至少一个所述预定编码段组成一个语言编码块;其中,不同所述语言编码块具有至少部分相同的预定类别;
    相应的,在将预定编码划分预定类别的步骤中包括:在属于不同所述语言编码块的名称字符串所处的预定编码段对应的预定类别相同时,将所述名称字符串同划分至所述预定类别中。
  14. 如权利要求13所述的方法,其特征在于,在确定预定编码段的步骤中包括:
    确定所述名称字符串中第一预定字符的统一码对应的语言编码块;
    在所述语言编码块中,确定所述第一预定字符的统一码所在的预定编码段。
  15. 如权利要求14所述的方法,其特征在于,所述方法还包括:
    在没有与所述名称字符串中第一预定字符的统一码对应的语言编码块时,将所述名称字符串划分至预设指定类别。
  16. 如权利要求9所述的方法,其特征在于,所述客户端在展示类别标识和名称字符串的步骤中还包括:
    将属于同一所述预定类别的所述名称字符串按照所述第一预定字符的统一码的编码顺序排列。
  17. 如权利要求16所述的方法,其特征在于,在所述第一预定字符的统一码相同时,所述方法还包括:
    获取所述名称字符串中第二预定字符;
    将属于同一所述预定类别的所述名称字符串按照所述第二预定字符的统一码的编码顺序排列。
  18. 如权利要求9所述的方法,其特征在于,在展示类别标识和名称字符串的步骤中还包括:
    获取所述名称字符串中第二预定字符的统一码;
    将属于同一所述预定类别的所述名称字符串按照所述第二预定字符的统一码的编码顺序排列。
  19. 如权利要求9所述的方法,其特征在于,在获取名称数据的步骤中还包括:接收用户的指定的第一预定字符。
  20. 一种服务器,其特征在于,其包括:
    数据获取模块,用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;
    统一码获取模块,用于获取所述名称字符串中第一预定字符的统一码;
    预定类别划分模块,用于确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
  21. 如权利要求20所述的服务器,其特征在于:至少一个所述预定编码段组成一 个语言编码块;其中,不同所述语言编码块具有至少部分相同的预定类别;
    相应的,所述预定类别划分模块用于:在属于不同所述语言编码块的名称字符串所处的预定编码段对应的预定类别相同时,将所述名称字符串同划分至所述预定类别中。
  22. 如权利要求21所述的服务器,其特征在于,所述预定类别划分模块包括:
    语言编码块确定单元,用于确定所述名称字符串中第一预定字符的统一码对应的语言编码块;
    预定编码段确定单元,用于在所述语言编码块中,确定所述第一预定字符的统一码所在的预定编码段。
  23. 一种服务器,其特征在于,其包括:处理器,
    所述处理器用于获取名称数据;其中,所述名称数据包括有至少一个名称字符串;并用于获取所述名称字符串中第一预定字符的统一码;确定所述第一预定字符的统一码对应的预定编码段,将所述名称字符串划分至所述预定编码段对应的预定类别;其中,所述预定编码段的数量为至少一个,每个所述预定编码段对应一个预定类别。
PCT/CN2016/072307 2016-01-27 2016-01-27 名称列表展示、处理方法及客户端、服务器 Ceased WO2017128101A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
PCT/CN2016/072307 WO2017128101A1 (zh) 2016-01-27 2016-01-27 名称列表展示、处理方法及客户端、服务器
CN201680080048.5A CN108702407A (zh) 2016-01-27 2016-01-27 名称列表展示、处理方法及客户端、服务器

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2016/072307 WO2017128101A1 (zh) 2016-01-27 2016-01-27 名称列表展示、处理方法及客户端、服务器

Publications (1)

Publication Number Publication Date
WO2017128101A1 true WO2017128101A1 (zh) 2017-08-03

Family

ID=59396968

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/072307 Ceased WO2017128101A1 (zh) 2016-01-27 2016-01-27 名称列表展示、处理方法及客户端、服务器

Country Status (2)

Country Link
CN (1) CN108702407A (zh)
WO (1) WO2017128101A1 (zh)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112241558B (zh) * 2020-09-03 2024-08-30 深圳市华阳国际工程设计股份有限公司 元素类型名的统一方法、装置以及计算机存储介质
CN112507000B (zh) * 2020-12-23 2021-12-14 深圳市普渡科技有限公司 机器人配置目标点的方法、装置、电子装置和存储介质
CN116011433A (zh) * 2021-10-22 2023-04-25 伊姆西Ip控股有限责任公司 应用测试的方法、设备和计算机程序产品

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20040069880A (ko) * 2003-01-30 2004-08-06 엘지전자 주식회사 휴대 단말기 소스 파일의 유니코드 변환 방법
CN101459712A (zh) * 2009-01-05 2009-06-17 深圳华为通信技术有限公司 一种电话本排序方法和手机设备
CN101686274A (zh) * 2008-09-22 2010-03-31 深圳富泰宏精密工业有限公司 联系人查找系统及方法
CN102281345A (zh) * 2011-06-10 2011-12-14 深圳桑菲消费通信有限公司 一种手机电话簿联系人的排序方法
CN103514160A (zh) * 2012-06-15 2014-01-15 华为终端有限公司 一种排序方法和移动设备
CN104902091A (zh) * 2015-05-27 2015-09-09 广东欧珀移动通信有限公司 一种通讯录排序方法及终端

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20040069880A (ko) * 2003-01-30 2004-08-06 엘지전자 주식회사 휴대 단말기 소스 파일의 유니코드 변환 방법
CN101686274A (zh) * 2008-09-22 2010-03-31 深圳富泰宏精密工业有限公司 联系人查找系统及方法
CN101459712A (zh) * 2009-01-05 2009-06-17 深圳华为通信技术有限公司 一种电话本排序方法和手机设备
CN102281345A (zh) * 2011-06-10 2011-12-14 深圳桑菲消费通信有限公司 一种手机电话簿联系人的排序方法
CN103514160A (zh) * 2012-06-15 2014-01-15 华为终端有限公司 一种排序方法和移动设备
CN104902091A (zh) * 2015-05-27 2015-09-09 广东欧珀移动通信有限公司 一种通讯录排序方法及终端

Also Published As

Publication number Publication date
CN108702407A (zh) 2018-10-23

Similar Documents

Publication Publication Date Title
US8095547B2 (en) Method and apparatus for detecting spam user created content
US8812300B2 (en) Identifying related names
CN104572685B (zh) 数据排序方法
CN104750666B (zh) 一种文本字符编码方式的识别方法及系统
CN103870001B (zh) 一种生成输入法候选项的方法及电子装置
CN111984589A (zh) 文档处理方法、文档处理装置和电子设备
CN112818692B (zh) 命名实体识别和处理方法、装置、设备及可读存储介质
US10496751B2 (en) Avoiding sentiment model overfitting in a machine language model
US11789586B2 (en) Method for displaying view and terminal device
WO2017128101A1 (zh) 名称列表展示、处理方法及客户端、服务器
CN111428230A (zh) 一种信息验证方法、装置、服务器及存储介质
CN112825089A (zh) 文章推荐方法、装置、设备及存储介质
CN108073655A (zh) 一种数据查询方法及装置
CN110134920B (zh) 绘文字兼容显示方法、装置、终端及计算机可读存储介质
CN110852078A (zh) 生成标题的方法和装置
US7366984B2 (en) Phonetic searching using multiple readings
CN113590756A (zh) 信息序列生成方法、装置、终端设备和计算机可读介质
CN115525728A (zh) 汉字排序、汉字检索和汉字插入的方法和装置
CN105653713B (zh) 一种确定设备识别码存在的方法及装置
CN118261152A (zh) 文本语料处理方法、装置、电子设备和存储介质
CN114661979B (zh) 信息处理方法及装置、设备、计算机可读存储介质
US7130470B1 (en) System and method of context-based sorting of character strings for use in data base applications
CN113378555B (zh) 个股的智能关联方法及相关产品
CN112989011B (zh) 数据查询方法、数据查询装置和电子设备
CN116740732A (zh) 一种影像文件的处理方法、装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16886988

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16886988

Country of ref document: EP

Kind code of ref document: A1