WO2014079290A1 - 对用户的长关系链数据的处理系统和方法 - Google Patents

对用户的长关系链数据的处理系统和方法 Download PDF

Info

Publication number
WO2014079290A1
WO2014079290A1 PCT/CN2013/085153 CN2013085153W WO2014079290A1 WO 2014079290 A1 WO2014079290 A1 WO 2014079290A1 CN 2013085153 W CN2013085153 W CN 2013085153W WO 2014079290 A1 WO2014079290 A1 WO 2014079290A1
Authority
WO
WIPO (PCT)
Prior art keywords
unit
log file
modification request
operation log
module
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2013/085153
Other languages
English (en)
French (fr)
Inventor
王辉
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Tencent Technology Shenzhen Co Ltd
Original Assignee
Tencent Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tencent Technology Shenzhen Co Ltd filed Critical Tencent Technology Shenzhen Co Ltd
Priority to US14/646,794 priority Critical patent/US9754006B2/en
Publication of WO2014079290A1 publication Critical patent/WO2014079290A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/27Replication, distribution or synchronisation of data between databases or within a distributed database system; Distributed database system architectures therefor
    • G06F16/275Synchronous replication
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/23Updating
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/95Retrieval from the web
    • G06F16/958Organisation or management of web site content, e.g. publishing, maintaining pages or automatic linking
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/10File systems; File servers
    • G06F16/17Details of further file system functions
    • G06F16/172Caching, prefetching or hoarding of files
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/10File systems; File servers
    • G06F16/17Details of further file system functions
    • G06F16/178Techniques for file synchronisation in file systems
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/28Databases characterised by their database models, e.g. relational or object models
    • G06F16/284Relational databases
    • G06F16/288Entity relationship models
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/40Business processes related to social networking or social networking services
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02DCLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • the present application relates to the field of computer and Internet data processing technologies, and in particular, to a system and method for processing long-chain data of a user. Background of the invention
  • UGC User Generated Content
  • SNS Social Network Service
  • the SNS is called the social network system and is an Internet application service system designed to help people build social networks.
  • the website system is expanding its social convenience and adding SNS features.
  • all the website systems with SNS features are called social network systems, such as: online community system, blog system, microblogging system. )Wait.
  • each user is the publisher of the information, and a large number of UGCs are produced almost all the time.
  • each user has its own relationship chain, and the user relationship chain mainly includes a user group that can interact with the user in the SNS, and the user relationship chain data includes Information about the identity, attributes, etc. of each user in this group, and the relationship between each user and the primary user.
  • some users have a large number of users in the relationship chain.
  • This relationship chain is called a long relationship chain in the industry, and a user with a long relationship chain is called a long relationship chain user.
  • MicroBlog is an SNS system for information sharing, dissemination, and retrieval based on user relationships. Users can access Weibo through wired communication networks or wireless communication networks, and various clients. A specified number of text and/or other multimedia information updates, and instant sharing.
  • each user can listen to (or pay attention to) other users, that is, the microblog information (ie UGC) published by the user who is listening (or paying attention) by the user can be transmitted to the user's Weibo in time.
  • the listener is the "listener" of the listener (some Weibo systems are also called "fans". In this article, the audience is taken as an example).
  • FIG. 1 is a processing system for long relationship chain user data in the prior art.
  • the system mainly includes a cache module and an inbound module.
  • the database stores the full length relationship chain data of the long relationship chain user.
  • the microblog system is a full audience list, and since the long relationship chain user reads the microblog, the full audience list is not required, and The front end responds very quickly to the read request of the long relationship chain user to the listener list, so a part of the listener list of each long relationship chain user is saved in the cache module of the memory according to the update time, and the cache module is used to respond to the front end.
  • the operation request for the long relationship chain data of the long relationship chain user because of the rapid memory operation, the response request of the long relationship chain user to the relational chain data can be responded to very quickly.
  • the modification request For the write operation request, that is, the modification request for the inbound modification of the corresponding database, the modification request needs to be synchronized to the inbound module, and the inbound module modifies the data in the database according to the modification request.
  • the cache module Since the cache module is a pure memory operation, the inbound module operates directly on the underlying database, and the speed of operating the database and pure memory operations is not an order of magnitude, and the speed difference is too large.
  • the cache module In order to solve the inconsistency between the cache module and the inbound module, the cache module must save the operation log of the long relationship chain user for a long time, until the inbound module completes the operation of the database to release the space occupied by the operation log, so the prior art is for the user.
  • the storage of the long relational chain data needs to rely heavily on the cache module, which not only occupies a large amount of memory space of the cache module, but also clears the memory if the cache module is abnormally restarted, and then loses a large number of modification requests, resulting in a long relationship in the database.
  • the chain data is seriously inconsistent with the front-end operation and the data error rate is high.
  • the main purpose of the present application is to provide a system and method for processing long-chain data of a user, so as to reduce the probability of losing the modification request and reduce the data error rate of the long-chain data in the database.
  • a processing system for long relationship chain data of a user including a cache module and a receiving module And a receiving module, wherein:
  • the cache module is disposed in the memory, and is configured to synchronously send the modification request in the operation request to the receiving module in response to an operation request of the front end to the long relationship chain data of the user;
  • the receiving module is configured to receive a modification request from the cache module, and save the modification request to an operation log file of a non-memory storage device;
  • the warehousing module is configured to read a modification request in an operation log file of the receiving module, and modify long relationship chain data in the database according to the read modification request.
  • a method for processing long relationship chain data of a user including:
  • the cache module buffers the operation request of the front end to the long relationship chain data of the user, and synchronously sends the modification request in the operation request to the receiving module described later;
  • the warehousing module reads the modification request in the operation log file of the receiving module, and modifies the long relationship chain data in the database according to the read modification request.
  • the modification request in the operation log file is read, and the long relationship chain data in the database is modified according to the read modification request.
  • an embodiment of the present application introduces a receiving module, and the cache module will be
  • the modification request of the user long relationship chain data is synchronized to the receiving module, and the receiving module synchronously saves the modification request in an operation log file of the non-memory storage device, and the inbound module does not acquire the modification of the front end from the cache module.
  • the request is obtained by reading the operation log file of the receiving module to obtain a modification request, and modifying the long relationship chain data in the database according to the modification request. Therefore, the application successfully decouples the cache module and the inbound module, and coordinates a fast modification request and a slow database operation by adding a receiving module, and the synchronous synchronization request from the cache module to the receiving module is faster.
  • the operation log file in the receiving module can be saved for the storage module for a long time. Therefore, the cache module does not have to save the operation log of the long relationship chain user for a long time, and reduces the occupation of the memory space by the cache module, and also reduces the occupation. Due to the abnormal restart of the cache module, the probability of losing the modification request is lost, thereby ensuring the consistency of the long relational chain data in the database with the front-end operation, and reducing the data error rate of the long relational chain data in the database.
  • FIG. 1 is a processing system for long relationship chain data of a user in the prior art
  • FIG. 2 is a schematic diagram of a composition processing system for long relationship chain data of a user according to the present application;
  • FIG. 3 is a flow chart of a method for processing long relationship chain data of a user according to the present application
  • FIG. 5 is a flow chart of the embodiment of FIG. 4 when the receiving module is restarted due to operation and maintenance operations, machine restart, and module abnormality;
  • FIG. 6 is a flow chart of the embodiment of FIG. 4 when the warehousing module is restarted due to operation and maintenance operations, machine restart, and module abnormality. Mode for carrying out the invention
  • FIG. 2 is a schematic diagram of a composition of a system for processing long-chain data of a user according to the present application.
  • the processing system includes a cache module 201, a receiving module 202, and an inbound module 203.
  • the cache module 201 is disposed in the memory, and the cached module 201 caches a part of the long relationship chain data of the latest update of the long relationship chain user, and is used to respond to the front end (such as the client, the web front end, ie, the user operation end).
  • the operation request of the user's long relationship chain data for most of the read requests, can be directly read from the cache module 201 and returned to the front end, thereby realizing the relationship chain data of the long relationship chain user.
  • the speed of the read response For the modification request of the front end to the long relationship chain data of the user, for example, adding a listener request in the microblog system, deleting the listener request, and modifying the listener request, the modification requests are synchronously sent to the receiving module 202.
  • the receiving module 202 is configured to receive a modification request from the cache module 201, and save the modification request as a log record to an operation log file (binlog) of the non-memory storage device, that is, the operation log file is not saved in the operation log file.
  • an operation log file (binlog) of the non-memory storage device
  • the operation log file is not saved in the operation log file.
  • memory it is stored in a storage device such as a hard disk. Since the request for modification of the long relationship chain data of the long relationship chain user is large, a new operation log file can be added to save the modification request after an operation log file is full.
  • the operation log file is written much faster than the database operation, and is similar to the speed of reading the memory, so the modification request received in the cache module 201 can be quickly synchronized to the receiving module 203.
  • the warehousing module 203 is configured to operate a database (DB), specifically for reading a log record of a strip in the operation log file of the receiving module 202, that is, a modification request, and modifying the database according to the read modification request.
  • DB database
  • Long relationship chain data For example, if a request to add a listener to the long relationship chain user is added to the long relationship chain data of the user in the database The listener.
  • the warehousing module 203 needs to read the operation log file of the receiving module 202 for a long time, and the operation log file is saved in the non-memory storage device. In the process, even if the machine is shut down due to failure, maintenance, etc., these operation log files are not lost, which reduces the probability of loss of the modification request, thereby ensuring the consistency of the long relationship chain data in the database with the front-end operation, and reducing the length in the database.
  • the data error rate of relational chain data can quickly synchronize the modification request of the user's long relationship chain data to the receiving module 202.
  • the cache module 201 does not have to save the operation log of the long relationship chain user for a long time.
  • the occupation of the memory space by the cache module 202 is reduced.
  • the inbound module in the prior art only saves the long relationship chain data of the long relationship chain user into a database.
  • the system needs to be expanded.
  • you When expanding the capacity, you must re-create a set of inbound modules and a larger capacity database, and then migrate all the data in the original database to the new database. This full amount of data migration for each expansion causes difficulties in the operation and maintenance of system equipment.
  • the processing system for the long relationship chain data of the user further includes a division user module, configured to perform unit division on the user population, and notify the unit (Unit) information.
  • the cache module 201, the receiving module 202, and the inbound module 203 are provided.
  • the unit division of the user population is to group the user groups, and each group is called a unit, which is convenient for expansion.
  • the cache module 201 is further configured to: distinguish the unit to which the user that initiated the modification request belongs, and send the modification request and its unit information to the receiving module synchronously.
  • the receiving module 202 is further configured to: respectively establish at least one operation log file for different units, as shown in FIG. 2, and save the modification request to an operation log file of the modification request corresponding unit. For a unit, after an operation file is full, a new operation log file can be added to save the modification request corresponding to the unit.
  • the warehousing module is further configured to: establish different databases according to the unit information for different units, as shown in FIG. 2, and modify the database of the corresponding unit of the modification request according to the read modification request when modifying the database.
  • Long relationship chain data The present application forms a minimum processing unit by unit dividing the user group.
  • the receiving module 202 and the inbound module 203 process the modification request and the storage in units of units. Library. After the total number of users in the SNS system is increased, the newly added users can be divided into new units.
  • a new processing device such as a server
  • the receiving module and the inbound module and the database only need to copy the operation log file of the unit to be migrated into the newly added receiving module, and open the newly added receiving module and the inbound module and the database, and then in the cache module.
  • the routing information of the newly added receiving module is added, so that the cache module can synchronously send the modification request of the new unit to the new receiving module. If you want to expand the database, you only need to stop the operation of the warehousing module, copy the original database data to the destination or add the database corresponding to the new unit, modify the routing information of the newly added database in the warehousing module, and then open the warehousing module. can.
  • This application uses a sub-unit processing to facilitate the expansion and expansion of the receiving module, the warehousing module, and the database.
  • the expansion-related equipment can be expanded very conveniently, it is possible to perform timely device expansion and expansion processing on bursts of data bursts, thereby ensuring a faster response speed of the entire SNS system to front-end data requests.
  • the receiving module 202 is further configured to: in synchronization with each unit, record synchronization progress information of the modification request in a memory, and feed back the synchronization progress information to the Cache module 201; after the receiving module stops and restarts, scans the operation log file of the unit for each unit, restores the synchronization progress information of the unit according to the latest operation log file of the unit, and synchronizes the progress information of the unit Feedback to the cache module.
  • the cache module 201 is further configured to: synchronously send, according to the progress information of each unit fed back by the receiving module 202, a modification request after the progress of the corresponding unit to the receiving module.
  • the cache module 201 receives the progress feedback and then sends the 100th modification request of the i-th unit and its subsequent modification request. Therefore, it is possible to further ensure that the receiving module 202 can automatically and automatically synchronize the modification request that is not successfully synchronized during the shutdown of the receiving module 202 after the shutdown and restart caused by the fault, the operation and maintenance operation, etc., thereby reducing the difficulty of operation and maintenance.
  • the receiving module 202 specifically includes a synchronization progress recording module, configured to record the synchronization progress information, and synchronize the progress progress information of the operation log corresponding unit after each operation operation log file synchronously saves a modification request. Add 1 to achieve synchronization of the record to save the synchronization progress information of the modification request.
  • the warehousing module 203 may be further configured to: for each unit, record the read progress information of the modification request in the unit operation log file in the memory, and each time an operation is read The log file marks that the file has been read; after the storage module 203 is stopped and restarted, the operation log file of the unit is scanned for each unit, and the reading progress of the unit is resumed according to the unread operation log file with the longest storage time.
  • the information according to the reading progress information of the unit, successively reads the modification request in the operation log file of the unit, and modifies the long relationship chain data in the database according to the read modification request. This can further ensure that the warehousing module 203 can quickly and automatically recover after a shutdown due to a failure, an operation and maintenance operation, and the like. The reading progress before the shutdown reduces the difficulty of operation and maintenance.
  • the warehousing module 203 specifically includes a read progress recording module, configured to record the read progress information, and read a record of the corresponding unit every time a record in the operation log file of one unit is read. Add 1 to record the read progress information of the modification request in the unit operation log file.
  • the present application also discloses a method of processing long relationship chain data of a user, which can be executed by the system.
  • FIG. 3 is a flow chart of a method for processing long relationship chain data of a user according to the present application. Referring to Figure 3, the method includes:
  • the cache module 201 sends the modification request to the receiving module 202, which will be described later, in response to the operation request of the front end to the long relationship chain data of the user.
  • the receiving module 202 receives the modification request from the cache module 201, and saves the modification request to an operation log file of the non-memory storage device.
  • the warehousing module 203 reads the modification request in the operation log file of the receiving module 202, and modifies the long relationship chain data in the database according to the read modification request.
  • the method further includes:
  • the user group is divided into units, and the unit information is notified to the cache module 201, the receiving module 202, and the inbound module 203.
  • the unit division of the user population is to group the user groups, each group being called a unit.
  • the specific method for unit dividing the user population may include: setting a specified unit size, sequentially numbering users in the system, and using the unit size to take a modulo value and a user having the same modulus value The same unit, or the number of the user is rounded up by the unit size, and the users with the same integer belong to the same unit.
  • the cache module 201 further distinguishes the single page to which the user who initiated the modification request belongs. And sending, by the receiving module 202, the modification request and the unit information to the receiving module 202; the receiving module 202 further establishing, according to different units, at least one operation log file, and saving the modification request to the modification request Corresponding to the operation log file of the unit; the warehousing module 203 further establishes different databases according to the unit information for different units, and when modifying the database, modifies the database of the corresponding unit of the modification request according to the read modification request Long relationship chain data.
  • the method of the present application may further include:
  • the receiving module 202 records synchronization progress information for synchronously saving the modification request in the memory for each unit, and feeds back the synchronization progress information to the cache module 201.
  • the specific manner of recording the synchronization progress information of the modification request in the memory includes: adding, after each operation operation log file synchronization to save a modification request, adding the synchronization progress information of the corresponding unit of the operation document to the file .
  • the receiving module 202 After the receiving module 202 stops and restarts, for each unit, scans the operation file of the unit, restores the synchronization progress information of the unit according to the latest operation log file of the unit, and feeds back the synchronization progress information of the unit to the unit.
  • the cache module 201; the cache module 201 synchronously transmits the modification request after the progress of the corresponding unit to the receiving module 202 according to the progress information of each unit fed back by the receiving module 202.
  • the restoring the synchronization progress information according to the latest operation log file of the unit specifically comprising: determining the number m of operation log files of the unit before the latest operation log file, and multiplying the number by an operation log file record Capacity n, using m X n as the synchronization progress information of the unit.
  • the method of the present application may further include: the inbound module 203, for each unit, recording, in the memory, the read progress information of the modification request in the unit operation file. , each time an operation log file is read, the file is marked as read. The reading of the modification request in the unit operation log file is recorded in the memory
  • the specific way of taking the progress information includes: Each time a record in the operation log file of one unit is read, the reading progress of the corresponding unit is incremented by one.
  • the specific method for the read progress information of the unread operation log file recovery unit according to the longest retention time includes: determining the number M of operation log files before the unread operation log file of the unit having the longest storage time, the number Multiply the record capacity n of an operation log file, and use MX n as the read progress information of the unit.
  • H ⁇ pre-divisions all users in the system according to the specified unit size (such as 4999).
  • the process includes:
  • Step 401 The Cache module 201 (ie, the cache module) synchronously sends a modification request for the user long relationship chain data from the front end to the receiving module 202 by using the specified unit size as a unit.
  • Step 402 The receiving module 202 receives the modification request sent by the cache module 201 synchronously, distinguishes the unit of the modification request, and records the modification request as an operation log into an operation log corresponding to the unit, to form an operation log file.
  • Step 403 After each operation operation log file synchronously saves a modification request, the synchronization progress information of the corresponding unit of the operation log file is incremented by 1, and the Cache module 201 is notified to synchronously send the next modification request of the unit.
  • Step 404 the warehousing module 203 scans and reads the receiving module 202 in units of units.
  • the operation log file in .
  • Step 405 The warehousing module 203 performs a modification operation on the underlying database of the corresponding unit according to the modification request recorded in the operation log file corresponding to each unit.
  • Step 406 The warehousing module 203 increments the reading progress of the corresponding unit by one for each record in the operation log file of one unit.
  • Step 407 Each time the operation log file is read by the warehousing module 203, the operation file mark is read.
  • FIG. 5 is a flow chart of the embodiment of FIG. 4 when the receiving module 202 is restarted due to operation and maintenance operations, machine restart, and module abnormality. Referring to FIG. 5, the process specifically includes:
  • Step 410 The receiving module 202 restarts.
  • Step 411 The receiving module 202 scans the operation log file of the unit for each unit in units of the unit.
  • Step 412 Restore the synchronization progress information of the unit according to the latest operation log file of each unit, and feed back synchronization progress information of the unit to the cache module 201. For example, it is: Determine the number of operation log files m of the unit before the latest operation log file, and multiply the number by the recording capacity n of an operation log file, and use m x n as the synchronization progress information of the unit.
  • Step 413 The cache module 201 synchronously sends the modification request of the corresponding unit to the receiving module 202 according to the progress information of each unit fed back by the receiving module 202.
  • Figure 6 is a flow chart of the embodiment of Figure 4 when the warehousing module 203 is restarted due to operation and maintenance operations, machine restarts, and module anomalies. Referring to FIG. 6, the process specifically includes:
  • Step 420 the warehousing module 203 is restarted.
  • Step 421 The warehousing module 203 scans the operation log file of the unit for each unit in the unit.
  • Step 422 The warehousing module 203 restores the reading progress information of the unit according to the unread operation log file with the longest storage time. For example, it is specifically: determining the number M of operation log files before the unread operation log file whose storage time is the longest, and multiplying the number by the recording capacity n of an operation file, and reading MX n as the unit Progress information.
  • Step 423 The warehousing module 203 successively reads the modification request in the operation log file of the unit according to the read progress information of each unit, for example, reads the modification request from the M+1 operation log file, and reads the modification request according to the read Modify the request to modify the long relationship chain data in the database.
  • the methods and systems provided herein can be implemented by hardware, or computer readable instructions, or a combination of hardware and computer readable instructions.
  • Computer readable instructions for use in the present application are stored by a plurality of processors in a readable storage medium, such as a hard disk, a CD-ROM, a DVD, an optical disk, a floppy disk, a magnetic tape, a RAM, a ROM, or other suitable storage device.
  • a readable storage medium such as a hard disk, a CD-ROM, a DVD, an optical disk, a floppy disk, a magnetic tape, a RAM, a ROM, or other suitable storage device.
  • the application provides a computer readable storage medium for storing instructions for causing a system or device to perform the methods described herein.
  • the system or device provided by the present application has a storage medium in which computer readable program code is stored for implementing the functions of any of the above embodiments, and these systems or devices (or CPUs or MPUs) can read and execute Program code stored on a storage medium.
  • the program code read from the storage medium can implement any of the above embodiments, and thus the program code and the storage medium storing the program code are part of the technical solution.
  • the storage medium for providing the program code includes a floppy disk, a hard disk, a magneto-optical disk, and an optical disk (for example) Such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), disk, flash card, ROM, etc.
  • the program code can also be downloaded from the server via the communication network.
  • an operation at least partially implemented by the program code may be implemented by an operating system running on a computer, thereby implementing the technical solution of any of the above embodiments, wherein the computer is executed based on the program code. instruction.
  • the program code in the storage medium is written to the memory, wherein the memory is located in an expansion board inserted in the computer or in an expansion unit connected to the computer.
  • the CPU in the expansion board or expansion unit performs at least part of the operation based on the program code according to the instructions, thereby implementing the technical solution of any of the above embodiments.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Databases & Information Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Computing Systems (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

本申请公开了一种对用户的长关系链数据的处理系统和方法,包括:缓存模块响应前端对所述用户的长关系链数据的操作请求,将操作请求中的修改请求同步发送给后述的接收模块;接收模块接收来自所述缓存模块的修改请求,将所述修改请求同步保存到非内存存储设备的操作日志文件中;入库模块读取所述接收模块的操作日志文件中的修改请求,并按照读取的修改请求修改数据库中的长关系链数据。利用本申请,可以以降低丟失修改请求的几率,降低数据库中的长关系链数据的数据错误率。

Description

对用户的长关系链数据的处理系统和方法 本申请要求于 2012 年 11 月 26 日提交中国专利局、 申请号为 201210483647.5、 发明名称为 "对用户的长关系链数据的处理系统和方 法" 的中国专利申请的优先权, 其全部内容通过引用结合在本申请中。 技术领域
本申请涉及计算机和互联网数据处理技术领域, 尤其涉及一种对用 户的长关系链数据的处理系统和方法。 发明背景
目前, 随着互联网技术的发展, 网络逐渐成为人们获取信息的重要 来源, 特别是在互联网进入 Web2.0 时代后, 用户既是网站内容的浏览 者, 也是网站内容的制造者。 用户参与创造的内容被称为用户生成内容 ( UGC, User Generated Content ),如用户发布的日志、照片等。在 Web2.0 时代, 由于 UGC的大量涌现, 网络信息量呈几何级快速增长。
目前, 最活跃的网络通信系统之一就是社交网络服务系统(SNS, Social Network Service )„ SNS筒称为社交网络系统, 是旨在帮助人们建 立社会性网络的互联网应用服务系统。 目前, 几乎所有的网站系统都在 扩展其社交便利性, 为其增加 SNS特性, 本文中将所有具有 SNS特性 的网站系统称为社交网络系统, 例如: 网上社区系统、 博客系统、 微博 客系统(筒称微博)等。
在 SNS中, 每个用户都是信息的发布者, 几乎时时刻刻都在生产出 大量的 UGC。 而且每个用户都有其自身的关系链,所述用户关系链主要 包括在 SNS中能和该用户进行互动的用户群体,用户关系链数据中包括 这个群体中的每一个用户的标识、 属性等信息, 以及每一个用户与主用 户的关系。 其中, 有些用户的关系链中的用户数量巨大, 这种关系链在 业界被称为长关系链, 拥有长关系链的用户被称为长关系链用户。
例如, 微博客( MicroBlog ), 筒称微博, 是一个基于用户关系的信 息分享、传播以及获取的 SNS系统, 用户可以通过有线通信网络或无线 通信网络、 以及各种客户端访问微博, 以指定数目的文字和 /或其它多媒 体信息更新信息, 并实现即时分享。 在微博系统中, 每一个用户都可以 收听(或关注)其它用户, 即被该用户收听(或关注) 的用户所发布的 微博信息(即 UGC )可以及时地传输到该用户的微博中, 收听者就是被 收听者的 "听众" (有些微博系统中也叫 "粉丝", 本文中以听众为例进 行说明)。 当然, 所有的用户也可以被其它用户收听(或关注)。 当某一 用户的听众的数量超过一定数目之后, 则该用户就变成了长关系链用 户, 例如微博中的一些明星用户, 其听众的数量往往有几百万甚至上千 万。
在生产 UGC的 SNS中, 由于数据是用户产生的, 海量的用户催生 出海量数据, 最终带来更大量级的数据读写请求。 特别是长关系链用户 的数据处理, 由于其长关系链中包括百万级甚至千万级数量的听众, 添 加或删除一个听众也要对被收听用户的长关系链进行相应的数据修改, 因此, 针对长关系链数据的请求的数量巨大, 触发频繁, 导致对相应的 数据库的操作量也巨大和频繁。 因此, 对用户的长关系链数据需要特殊 的处理。
图 1为现有技术中的一种针对长关系链用户数据的处理系统。 参见 图 1 , 该系统中主要包括緩存(cache )模块和入库模块。 所述数据库中 保存了长关系链用户的全量长关系链数据, 例如微博系统中是全量的听 众列表, 而由于长关系链用户读取微博并不需要全量听众列表, 并为了 向前端极速的响应这类长关系链用户的对听众列表的读取请求, 因此将 每一长关系链用户的一部分听众列表按更新时间保存在内存的緩存模 块中, 该緩存模块用于响应前端(如客户端, 网页前端, 即用户操作端) 对所述长关系链用户的长关系链数据的操作请求, 由于内存操作迅速, 因此可以极速响应长关系链用户对关系链数据的读取请求。 对于写操作 请求, 即需要对相应数据库进行入库修改的修改请求, 则需要将这些修 改请求同步给所述入库模块, 由入库模块根据这些修改请求修改数据库 中的数据。
但是上述现有技术具有如下缺点:
由于緩存模块是纯内存操作, 入库模块是直接对底层数据库进行操 作, 而操作数据库和纯内存操作的速度不是一个数量级, 速度相差太悬 殊。 为了解决 Cache模块和入库模块速度的不一致, 緩存模块必须长时 间保存对长关系链用户的操作日志, 直到入库模块完成对数据库的操作 才能释放操作日志占用的空间, 因此现有技术对用户的长关系链数据的 入库存储需要严重依赖緩存模块, 不但占用了緩存模块的大量内存空 间, 而且一旦緩存模块出现异常重启则会清空内存, 进而丟失大量的修 改请求, 导致数据库中的长关系链数据与前端操作严重不符, 数据错误 率高。 发明内容
有鉴于此, 本申请的主要目的在于提供一种对用户的长关系链数据 的处理系统和方法, 以降低丟失修改请求的几率, 降低数据库中的长关 系链数据的数据错误率。
本申请的技术方案是这样实现的:
一种对用户的长关系链数据的处理系统, 包括緩存模块、 接收模块 以及接收模块, 其中:
所述緩存模块设置在内存中, 用于响应前端对所述用户的长关系链 数据的操作请求, 将所述操作请求中的修改请求同步发送给所述接收模 块;
所述接收模块用于接收来自所述緩存模块的修改请求, 将所述修改 请求同步保存到非内存存储设备的操作日志文件中;
所述入库模块用于读取所述接收模块的操作日志文件中的修改请 求, 并按照读取的修改请求修改数据库中的长关系链数据。 一种对用户的长关系链数据的处理方法, 包括:
緩存模块緩存前端对所述用户的长关系链数据的操作请求, 将所述 操作请求中的修改请求同步发送给后述的接收模块;
接收模块接收来自所述緩存模块的修改请求, 将所述修改请求同步 保存到非内存存储设备的操作日志文件中;
入库模块读取所述接收模块的操作日志文件中的修改请求, 并按照 读取的修改请求修改数据库中的长关系链数据。 一种存储介质, 用于存储计算机可执行指令; 所述计算机可执行指 令用于控制处理器执行一种对用户的长关系链数据的处理方法, 所述方 法包括:
响应前端对用户的长关系链数据的操作请求, 并将所述操作请求中 的修改请求同步保存到非内存存储设备的操作日志文件中;
读取所述操作日志文件中的修改请求, 并按照读取的修改请求修改 数据库中的长关系链数据。
与现有技术相比, 本申请一实施例引入了接收模块, 緩存模块将前 端的对用户长关系链数据的修改请求同步给该接收模块, 由该接收模块 将所述修改请求同步保存在非内存存储设备的操作日志文件中, 而入库 模块不是从緩存模块获取前端的修改请求, 而是通过读取所述接收模块 的操作日志文件来获取修改请求, 并根据修改请求修改数据库中的长关 系链数据。 因此, 本申请将所述緩存模块和入库模块成功解耦, 通过增 加一个接收模块, 来协调快速的修改请求和慢速的数据库操作, 从緩存 模块同步修改请求到接收模块的同步速度较快, 而接收模块中的操作日 志文件又可以长时间保存供入库模块读取, 因此, 緩存模块不必长时间 保存对长关系链用户的操作日志, 降低緩存模块对内存空间的占用, 也 降低了由于緩存模块异常重启导致丟失修改请求的几率, 进而保证数据 库中的长关系链数据与前端操作的符合程度, 降低了数据库中的长关系 链数据的数据错误率。 附图简要说明
图 1为现有技术中的一种针对用户的长关系链数据的处理系统; 图 2为本申请所述对用户的长关系链数据的处理系统的一种组成示 意图;
图 3 为本申请所述对用户的长关系链数据的处理方法的一种流程 图;
图 4为本申请所述具体实施例的一种流程图;
图 5为图 4所述实施例中当接收模块因为运维动作、 机器重启、 模 块异常而重启时的流程图;
图 6为图 4所述实施例中当入库模块因为运维动作、 机器重启、 模 块异常而重启时的流程图。 实施本发明的方式
下面结合附图及具体实施例对本申请再作进一步详细的说明 图 2为本申请所述对用户的长关系链数据的处理系统的一种组成示 意图。 参见图 2, 该处理系统包括緩存(cache )模块 201、 接收模块 202 以及入库模块 203。
所述緩存模块 201设置在内存中, 在该緩存模块 201中緩存了长关 系链用户的最新更新的一部分长关系链数据,用于响应前端(如客户端, 网页前端, 即用户操作端)对所述用户的长关系链数据的操作请求, 针 对其中大部分的读取请求, 可以直接从该緩存模块 201中读取并将读取 结果返回给前端, 从而实现长关系链用户的关系链数据的极速读取响 应。 而对于前端对所述用户的长关系链数据的修改请求, 例如微博系统 中的添加听众请求、 删除听众请求、 修改听众请求, 则将这些修改请求 同步发送给所述接收模块 202。
所述接收模块 202用于接收来自所述緩存模块 201的修改请求, 将 所述修改请求作为日志记录同步保存到非内存存储设备的操作日志文 件( Binlog ) 中, 即该操作日志文件不是保存在内存中, 而是保存在诸 如硬盘等存储设备中。 由于对长关系链用户的长关系链数据的修改请求 规模巨大, 因此可以在一个操作日志文件写满后, 增加新的操作日志文 件来保存所述修改请求。 写操作日志文件的速度要比操作数据库的速度 快很多, 与读取内存的速度相差不多, 因此緩存模块 201中接收到的修 改请求可以快速地同步给接收模块 203。
所述入库模块 203用于操作数据库( DB ), 具体用于读取所述接收 模块 202的操作日志文件中的一条条的日志记录即修改请求, 并按照读 取的修改请求修改数据库中的长关系链数据。 例如, 如果是向该长关系 链用户添加听众的请求, 则在数据库中的该用户的长关系链数据中添加 所述听众。
由于操作底层数据库的操作与内存操作的速度相差很多, 因此, 入 库模块 203需要较长时间地读取所述接收模块 202的操作日志文件, 而 所述操作日志文件保存在非内存的存储设备中, 即使由于故障、 维修等 原因停机,这些操作日志文件也不会丟失,降低了修改请求的丟失几率, 进而保证数据库中的长关系链数据与前端操作的符合程度, 降低了数据 库中的长关系链数据的数据错误率。 同时, 所述内存中的緩存模块 201 能够快速地将对用户的长关系链数据的修改请求同步给所述接收模块 202, 因此, 緩存模块 201不必长时间保存对长关系链用户的操作日志, 降低緩存模块 202对内存空间的占用。 另外, 图 1所示的现有技术中还存在模块扩展难的缺陷, 即现有技 术中的入库模块只会把长关系链用户的长关系链数据保存到一个数据 库里。 当系统的数据量增大后, 需要进行系统扩容, 在扩容时必须重新 创建一套入库模块和容量更大的数据库, 然后将原有数据库中的数据全 部迁移到新的数据库中。 这种每扩展一次就要进行全量的数据迁移造成 了对系统设备运营维护的困难。
作为一种改进, 在本申请的一种实施例中, 所述对用户的长关系链 数据的处理系统还进一步包括划分用户模块, 用于对用户人群进行单元 划分, 将单元(Unit )信息通知给所述緩存模块 201、 接收模块 202和 入库模块 203。 所述对用户人群进行单元划分, 就是将用户人群进行分 组, 每一组称为一个单元, 这样方便扩展。
在该实施例中, 所述緩存模块 201进一步用于: 区分发起所述修改 请求的用户所属的单元, 将所述修改请求及其单元信息同步发送给所述 接收模块。 所述接收模块 202进一步用于: 为不同的单元分别对应建立至少一 个操作日志文件, 如图 2所示, 并将所述修改请求同步保存到该修改请 求对应单元的操作日志文件中。 对于某一个单元, 可以在一个操作曰志 文件写满后, 增加新的操作日志文件来保存该单元对应的修改请求。
所述入库模块进一步用于: 按照所述单元信息为不同的单元对应建 立不同的数据库, 如图 2所示, 并在修改数据库时, 按照读取的修改请 求修改该修改请求对应单元的数据库中的长关系链数据。 本申请通过对用户群体进行单元划分, 形成了最小的处理单元, 在 对用户的长关系链数据的处理过程中, 接收模块 202和入库模块 203都 是按照单元为单位处理修改请求和存储入库的。当 SNS系统中的用户总 量攀升后, 可以将新增的用户划分到新的单元中, 在需要扩容接收模块 202和入库模块 203时, 可以新增处理设备 (如服务器)放置新扩容的 接收模块和入库模块以及数据库, 只需要把要迁移的单元的操作日志文 件复制到新增的接收模块中, 并开启新增的接收模块和入库模块和数据 库, 然后在所述緩存模块中添加新增的接收模块的路由信息, 以使緩存 模块可以将该新单元的修改请求同步发送给新的接收模块。 如果要扩容 数据库, 只需求停止入库模块的运行, 把原数据库数据复制到目的地或 增加新增单元对应的数据库, 修改入库模块中的新增数据库的路由信 息, 然后开启入库模块即可。
本申请由于采用了分单元处理, 方便接收模块、 入库模块以及数据 库的扩展扩容, 因此运营维护筒单。 同时, 由于可以非常方便地扩展扩 容相关设备, 因此, 可以对突发的数据量暴增作出及时的设备扩展扩容 处理, 从而保证整个 SNS系统对前端数据请求的较快的响应速度。 在本申请的又一种实施例中, 所述接收模块 202进一步用于: 针对 每一单元, 在内存中记录同步保存所述修改请求的同步进度信息, 并将 所述同步进度信息反馈给所述緩存模块 201 ;在本接收模块停机重启后, 针对每一单元, 扫描该单元的操作日志文件, 根据该单元最新的操作日 志文件恢复该单元的同步进度信息, 并将该单元的同步进度信息反馈给 所述緩存模块。 所述緩存模块 201进一步用于: 根据所述接收模块 202 反馈的各个单元的进度信息同步发送对应单元的、 该进度之后的修改请 求给所述接收模块。
例如, 接收模块 202返回的第 i个单元的进度信息为第 1000条, 则 緩存模块 201收到该进度反馈后再发送该第 i个单元的第 1001条修改请 求及其后续的修改请求。 因此,可以进一步保证接收模块 202由于故障、 运维动作等导致的停机重启后, 可以快速自动同步在所述接收模块 202 停机期间未同步成功的修改请求, 降低了运营维护的难度。
具体的, 所述接收模块 202中具体包括同步进度记录模块, 用于记 录所述同步进度信息, 在每次操作操作日志文件同步保存一条修改请求 完毕, 就把该操作日志对应单元的同步进度信息加 1 , 从而实现记录同 步保存所述修改请求的同步进度信息。
在又一种实施例中, 所述入库模块 203也可以进一步用于: 针对每 一单元, 在内存中记录对该单元操作日志文件中的修改请求的读取进度 信息, 每读完一个操作日志文件则标记该文件已读; 在本入库模块 203 停机重启后, 针对每一单元, 扫描该单元的操作日志文件, 根据保存时 间最久的未读操作日志文件恢复该单元的读取进度信息, 按照该单元的 读取进度信息接续读取该单元的操作日志文件中的修改请求, 并按照读 取的修改请求修改数据库中的长关系链数据。 这样可以进一步保证入库 模块 203由于故障、 运维动作等导致的停机重启后, 可以快速自动恢复 到停机前的读取进度, 降低了运营维护的难度。
具体的, 所述入库模块 203中具体包括读取进度记录模块, 用于记 录所述读取进度信息, 每读取一单元的操作日志文件中的一条记录, 就 把对应单元的读取进度加 1 , 从而实现记录对该单元操作日志文件中的 修改请求的读取进度信息。 与上述系统对应, 本申请还公开了对用户的长关系链数据的处理方 法, 可由所述系统执行。 图 3为本申请所述对用户的长关系链数据的处 理方法的一种流程图。 参见图 3 , 该方法包括:
301、 緩存模块 201 响应前端对所述用户的长关系链数据的操作请 求, 将其中的修改请求同步发送给后述的接收模块 202。
302、接收模块 202接收来自所述緩存模块 201的修改请求,将所述 修改请求同步保存到非内存存储设备的操作日志文件中。
303、入库模块 203读取所述接收模块 202的操作日志文件中的修改 请求, 并按照读取的修改请求修改数据库中的长关系链数据。
为了进一步方便模块和数据的扩展扩容, 提升运营维护效率, 在进 一步的实施例中, 该方法进一步包括:
对用户人群进行单元划分, 将单元信息通知给所述緩存模块 201、 接收模块 202和入库模块 203。 所述对用户人群进行单元划分, 就是将 用户人群进行分组, 每一组称为一个单元。 所述对用户人群进行单元划 分的具体方法例如可以包括: 设置指定的单元大小, 对系统内的用户按 顺序编号, 对所述用户的编号利用所述单元大小取模值, 模值相同的用 户属于同一单元, 或者对所述用户的编号利用所述单元大小取整, 整数 相同的用户属于同一单元。
所述緩存模块 201 进一步区分发起所述修改请求的用户所属的单 元, 将所述修改请求及其单元信息同步发送给所述接收模块 202; 所述 接收模块 202 进一步为不同的单元分别对应建立至少一个操作日志文 件, 将所述修改请求同步保存到该修改请求对应单元的操作日志文件 中; 所述入库模块 203进一步按照所述单元信息为不同的单元对应建立 不同的数据库, 并在修改数据库时, 按照读取的修改请求修改该修改请 求对应单元的数据库中的长关系链数据。
在一种实施例中, 本申请的方法还可以进一步包括:
所述接收模块 202针对每一单元, 在内存中记录同步保存所述修改 请求的同步进度信息,并将所述同步进度信息反馈给所述緩存模块 201。 所述在内存中记录同步保存所述修改请求的同步进度信息的具体方式 包括: 每次操作操作日志文件同步保存一条修改请求完毕后, 就把该操 作曰志文件对应单元的同步进度信息加 1。
在本接收模块 202停机重启后, 针对每一单元, 扫描该单元的操作 曰志文件, 根据该单元最新的操作日志文件恢复该单元的同步进度信 息, 并将该单元的同步进度信息反馈给所述緩存模块 201 ; 所述緩存模 块 201根据所述接收模块 202反馈的各个单元的进度信息同步发送对应 单元的、 该进度之后的修改请求给所述接收模块 202。 所述根据单元的 最新的操作日志文件恢复同步进度信息, 具体包括: 确定该单元的保存 时间在该最新的操作日志文件之前的操作日志文件数目 m, 将该数目乘 以一个操作日志文件的记录容量 n, 将 m X n作为该单元的同步进度信 息。
在又一种实施例中, 本申请所述的方法还可以进一步包括: 所述入库模块 203针对每一单元, 在内存中记录对该单元操作曰志 文件中的修改请求的读取进度信息, 每读完一个操作日志文件则标记该 文件已读。 所述在内存中记录对该单元操作日志文件中的修改请求的读 取进度信息的具体方式包括: 每读取一单元的操作日志文件中的一条记 录, 就把对应单元的读取进度加 1。
在本入库模块 203停机重启后, 针对每一单元, 扫描该单元的操作 日志文件, 根据保存时间最久的未读操作日志文件恢复该单元的读取进 度信息, 按照该单元的读取进度信息接续读取该单元的操作日志文件中 的修改请求, 并按照读取的修改请求修改数据库中的长关系链数据。 所 述根据保存时间最久的未读操作日志文件恢复单元的读取进度信息的 具体方法包括: 确定该单元的保存时间最久的未读操作日志文件之前的 操作日志文件数目 M, 将该数目乘以一个操作日志文件的记录容量 n, 将 M X n作为该单元的读取进度信息。 下面以一个更为具体的实施例, 进一步描述本申请所述的方法。 图 4为本申请所述具体实施例的一种流程图。 参见图 2和图 4, H殳预先 根据指定单元大小(如 4999 )对系统内的所有用户进行了单元划分, 该 流程包括:
步骤 401、 Cache模块 201 (即緩存模块)以指定单元大小为单元单 位将来自前端的对用户长关系链数据的修改请求同步发送给接收模块 202。
步骤 402、接收模块 202接收 cache模块 201同步发送来的修改请求, 区分该修改请求的单元, 并将该修改请求作为操作日志记录到该单元对 应的操作日志中, 形成操作日志文件。
步骤 403、 每次操作操作日志文件同步保存一条修改请求完毕后, 就把该操作日志文件对应单元的同步进度信息加 1 , 并通知 Cache模块 201同步发送该单元的下一个修改请求。
步骤 404、入库模块 203以单元为单位扫描并读取所述接收模块 202 中的操作日志文件。
步骤 405、 入库模块 203根据所述每个单元对应的操作日志文件中 记录的修改请求来对相应单元的底层数据库进行修改操作。
步骤 406、 入库模块 203每读取一单元的操作日志文件中的一条记 录, 就把对应单元的读取进度加 1。
步骤 407、 入库模块 203每读取完一个操作日志文件, 就对这个操 作曰志文件标记已读。 图 5为图 4所述实施例中当接收模块 202因为运维动作、机器重启、 模块异常而重启时的流程图。 参见图 5 , 该流程具体包括:
步骤 410、 接收模块 202重启。
步骤 411、 接收模块 202以所述单元为单位, 针对每一单元, 扫描 该单元的操作日志文件。
步骤 412、 根据每一单元最新的操作日志文件恢复该单元的同步进 度信息, 并将该单元的同步进度信息反馈给所述緩存模块 201。 例如具 体为: 确定该单元的保存时间在该最新的操作日志文件之前的操作日志 文件数目 m, 将该数目乘以一个操作日志文件的记录容量 n, 将 m x n 作为该单元的同步进度信息。
步骤 413、 所述緩存模块 201根据所述接收模块 202反馈的各个单 元的进度信息同步发送对应单元的、 该进度之后的修改请求给所述接收 模块 202。 图 6为图 4所述实施例中当入库模块 203因为运维动作、机器重启、 模块异常而重启时的流程图。 参见图 6 , 该流程具体包括:
步骤 420、 入库模块 203重启。 步骤 421、 入库模块 203以所述单元为单元, 针对每一单元, 扫描 该单元的操作日志文件。
步骤 422、 入库模块 203根据保存时间最久的未读操作日志文件恢 复该单元的读取进度信息。 例如具体为: 确定该单元的保存时间最久的 未读操作日志文件之前的操作日志文件数目 M, 将该数目乘以一个操作 曰志文件的记录容量 n, 将 M X n作为该单元的读取进度信息。
步骤 423、 入库模块 203按照各单元的读取进度信息接续读取该单 元的操作日志文件中的修改请求,例如从第 M+1个操作日志文件开始读 取修改请求, 并按照读取的修改请求修改数据库中的长关系链数据。 本申请提供的方法和系统可以由硬件、 或计算机可读指令、 或者硬 件和计算机可读指令的结合来实现。 本申请中使用的计算机可读指令由 多个处理器存储在可读存储介质中, 例如硬盘、 CD-ROM, DVD, 光盘、 软盘、 磁带、 RAM、 ROM或其它合适的存储设备。 或者, 至少部分计 算机可读指令可以由具体硬件替换, 例如, 定制集成线路、 门阵列、 FPGA、 PLD和具体功能的计算机等等。
本申请提供了计算机可读存储介质, 用于存储指令使得系统或设备 执行本文所述的方法。 具体地, 本申请提供的系统或设备都具有存储介 质,其中存储了计算机可读程序代码,用于实现上述任意实施例的功能, 并且这些系统或设备 (或 CPU或 MPU ) 能够读取并且执行存储在存储 介质中的程序代码。
在这种情况下, 从存储介质中读取的程序代码可以实现上述任一实 施例, 因此该程序代码和存储该程序代码的存储介质是技术方案的一部 分。
用于提供程序代码的存储介质包括软盘、 硬盘、 磁光盘、 光盘(例 如 CD-ROM、 CD-R, CD-RW、 DVD-ROM、 DVD-RAM、 DVD-RW, DVD+RW ), 磁盘、 闪存卡、 ROM等等。 可选地, 程序代码也可以通过 通信网络从 务器上下载。
应该注意的是, 对于由计算机执行的程序代码, 至少部分由程序代 码实现的操作可以由运行在计算机上的操作系统实现, 从而实现上述任 一实施例的技术方案, 其中该计算机基于程序代码执行指令。
另外, 存储介质中的程序代码被写入存储器, 其中, 该存储器位于 插入在计算机中的扩展板中, 或者位于连接到计算机的扩展单元中。 在 一实施例中,扩展板或扩展单元中的 CPU根据指令,基于程序代码执行 至少部分操作, 从而实现上述任一实施例的技术方案。 以上所述仅为本申请的较佳实施例而已, 并不用以限制本申请, 凡 在本申请的精神和原则之内, 所做的任何修改、 等同替换、 改进等, 均 应包含在本申请保护的范围之内。

Claims

权利要求书
1、 一种对用户的长关系链数据的处理系统, 其特征在于, 包括緩存 模块、 接收模块以及接收模块, 其中:
所述緩存模块设置在内存中, 用于响应前端对所述用户的长关系链 数据的操作请求, 将所述操作请求中的修改请求同步发送给所述接收模 块;
所述接收模块用于接收来自所述緩存模块的修改请求, 将所述修改 请求同步保存到非内存存储设备的操作日志文件中;
所述入库模块用于读取所述接收模块的操作日志文件中的修改请 求, 并按照读取的修改请求修改数据库中的长关系链数据。
2、 根据权利要求 1所述的系统, 其特征在于, 该系统进一步包括: 划分用户模块, 用于对用户人群进行单元划分, 并将所述单元划分信息 通知给所述緩存模块、 接收模块和入库模块;
所述緩存模块进一步用于: 区分发起所述修改请求的用户所属的单 元, 将所述修改请求及其单元信息同步发送给所述接收模块;
所述接收模块进一步用于: 为不同的单元分别对应建立至少一个操 作曰志文件, 将所述修改请求同步保存到该修改请求对应单元的操作曰 志文件中;
所述入库模块进一步用于: 按照所述单元信息为不同的单元对应建 立不同的数据库, 并在修改数据库时, 按照读取的修改请求修改该修改 请求对应单元的数据库中的长关系链数据。
3、 根据权利要求 2所述的系统, 其特征在于,
所述接收模块进一步用于: 针对每一单元, 在所述内存中记录同步 保存所述修改请求的同步进度信息, 并将所述同步进度信息反馈给所述 緩存模块; 在本接收模块停机重启后, 针对每一单元, 扫描该单元的操 作曰志文件, 根据该单元最新的操作日志文件恢复该单元的同步进度信 息, 并将该单元的同步进度信息反馈给所述緩存模块;
所述緩存模块进一步用于: 根据所述接收模块反馈的各个单元的进 度信息同步发送对应单元的、 该进度之后的修改请求给所述接收模块。
4、根据权利要求 3所述的系统, 其特征在于, 所述接收模块中具体 包括同步进度记录模块, 用于记录所述同步进度信息, 在每次操作操作 曰志文件同步保存一条修改请求完毕, 就把该操作日志对应单元的同步 进度信息加 1。
5、 根据权利要求 2所述的系统, 其特征在于,
所述入库模块进一步用于: 针对每一单元, 在所述内存中记录对该 单元操作日志文件中的修改请求的读取进度信息, 每读完一个操作曰志 文件则标记该文件已读; 在本入库模块停机重启后, 针对每一单元, 扫 描该单元的操作日志文件, 根据保存时间最久的未读操作日志文件恢复 该单元的读取进度信息, 按照该单元的读取进度信息接续读取该单元的 操作日志文件中的修改请求, 并按照读取的修改请求修改数据库中的长 关系链数据。
6、根据权利要求 5所述的系统, 其特征在于, 所述入库模块包括读 取进度记录模块, 用于记录所述读取进度信息, 每读取一单元的操作日 志文件中的一条记录, 就把对应单元的读取进度加 1。
7、 一种对用户的长关系链数据的处理方法, 其特征在于, 包括: 緩存模块緩存前端对所述用户的长关系链数据的操作请求, 将所述 操作请求中的修改请求同步发送给后述的接收模块;
接收模块接收来自所述緩存模块的修改请求, 将所述修改请求同步 保存到非内存存储设备的操作日志文件中; 入库模块读取所述接收模块的操作日志文件中的修改请求, 并按照 读取的修改请求修改数据库中的长关系链数据。
8、 根据权利要求 7所述的方法, 其特征在于, 该方法进一步包括: 对用户人群进行单元划分, 将单元信息通知给所述緩存模块、 接收 模块和入库模块;
所述緩存模块进一步区分发起所述修改请求的用户所属的单元, 将 所述修改请求及其单元信息同步发送给所述接收模块;
所述接收模块进一步为不同的单元分别对应建立至少一个操作曰志 文件, 将所述修改请求同步保存到该修改请求对应单元的操作日志文件 中;
所述入库模块进一步按照所述单元信息为不同的单元对应建立不同 的数据库, 并在修改数据库时, 按照读取的修改请求修改该修改请求对 应单元的数据库中的长关系链数据。
9、根据权利要求 8所述的方法, 其特征在于, 该方法进一步包括: 所述接收模块针对每一单元, 在内存中记录同步保存所述修改请求 的同步进度信息, 并将所述同步进度信息反馈给所述緩存模块; 在本接 收模块停机重启后, 针对每一单元, 扫描该单元的操作日志文件, 根据 该单元最新的操作日志文件恢复该单元的同步进度信息, 并将该单元的 同步进度信息反馈给所述緩存模块;
所述緩存模块根据所述接收模块反馈的各个单元的进度信息同步发 送对应单元的、 该进度之后的修改请求给所述接收模块。
10、 根据权利要求 9所述的方法, 其特征在于, 所述在内存中记录 同步保存所述修改请求的同步进度信息, 包括: 每次操作操作日志文件 同步保存一条修改请求完毕后, 就把该操作日志文件对应单元的同步进 度信息加 1。
11、 根据权利要求 9所述的方法, 其特征在于, 所述根据单元的最 新的操作日志文件恢复同步进度信息, 包括: 确定该单元的保存时间在 该最新的操作日志文件之前的操作日志文件数目 m, 将该数目乘以一个 操作日志文件的记录容量 n, 将 m X n作为该单元的同步进度信息。
12、根据权利要求 8所述的方法, 其特征在于, 该方法进一步包括: 所述入库模块针对每一单元, 在所述内存中记录对该单元操作曰志 文件中的修改请求的读取进度信息, 每读完一个操作日志文件则标记该 文件已读; 在本入库模块停机重启后, 针对每一单元, 扫描该单元的操 作日志文件, 根据保存时间最久的未读操作日志文件恢复该单元的读取 进度信息, 按照该单元的读取进度信息接续读取该单元的操作日志文件 中的修改请求, 并按照读取的修改请求修改数据库中的长关系链数据。
13、根据权利要求 12所述的方法, 其特征在于, 所述在内存中记录 对该单元操作日志文件中的修改请求的读取进度信息, 具体包括: 每读 取一单元的操作日志文件中的一条记录,就把对应单元的读取进度加 1。
14、根据权利要求 12所述的方法, 其特征在于, 所述根据保存时间 最久的未读操作日志文件恢复单元的读取进度信息, 具体包括: 确定该 单元的保存时间最久的未读操作日志文件之前的操作日志文件数目 M, 将该数目乘以一个操作日志文件的记录容量 n, 将 M X n作为该单元的 读取进度信息。
15、 一种存储介质, 用于存储计算机可执行指令; 所述计算机可执 行指令用于控制处理器执行一种对用户的长关系链数据的处理方法, 所 述方法包括:
响应前端对用户的长关系链数据的操作请求, 并将所述操作请求中 的修改请求同步保存到非内存存储设备的操作日志文件中;
读取所述操作日志文件中的修改请求, 并按照读取的修改请求修改 数据库中的长关系链数据。
16、根据权利要求 15所述的存储介质, 其特征在于, 所述响应前端 对用户的长关系链数据的操作请求包括: 将所述用户最新更新的一部分 长关系链数据緩存在内存中。
PCT/CN2013/085153 2012-11-26 2013-10-14 对用户的长关系链数据的处理系统和方法 Ceased WO2014079290A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US14/646,794 US9754006B2 (en) 2012-11-26 2013-10-14 System and method for processing long relation chain data of user

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201210483647.5A CN103838757B (zh) 2012-11-26 2012-11-26 对用户的长关系链数据的处理系统和方法
CN201210483647.5 2012-11-26

Publications (1)

Publication Number Publication Date
WO2014079290A1 true WO2014079290A1 (zh) 2014-05-30

Family

ID=50775501

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2013/085153 Ceased WO2014079290A1 (zh) 2012-11-26 2013-10-14 对用户的长关系链数据的处理系统和方法

Country Status (3)

Country Link
US (1) US9754006B2 (zh)
CN (1) CN103838757B (zh)
WO (1) WO2014079290A1 (zh)

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10007721B1 (en) * 2015-07-02 2018-06-26 Collaboration. AI, LLC Computer systems, methods, and components for overcoming human biases in subdividing large social groups into collaborative teams
CN106470150B (zh) * 2015-08-21 2020-04-24 腾讯科技(深圳)有限公司 关系链存储方法及装置
CN107924362B (zh) * 2015-09-08 2022-02-15 株式会社东芝 数据库系统、服务器装置、计算机可读取的记录介质及信息处理方法
US12461938B2 (en) * 2024-01-11 2025-11-04 Microsoft Technology Licensing, Llc Exporting customer data using a compliant tenant shard

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101826073A (zh) * 2009-03-06 2010-09-08 华为技术有限公司 分布式数据库同步方法、设备及系统
CN101876996A (zh) * 2009-12-01 2010-11-03 广州从兴电子开发有限公司 一种内存数据库到文件数据库的数据同步方法及系统
CN102024040A (zh) * 2010-12-08 2011-04-20 北京握奇数据系统有限公司 数据库同步方法、装置和系统
CN102238178A (zh) * 2011-05-30 2011-11-09 李牧森 基于互联网技术应用的多层次传播信息的客户端桌面媒体

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7412569B2 (en) * 2003-04-10 2008-08-12 Intel Corporation System and method to track changes in memory
CN1925530B (zh) * 2006-09-06 2011-01-05 华为技术有限公司 记录话单的系统及方法
CN101364217B (zh) * 2007-08-08 2011-06-22 华为技术有限公司 数据库中数据维护方法、设备及其系统
CN101247417B (zh) * 2008-03-07 2011-07-27 中国科学院计算技术研究所 双层元数据处理系统及方法
US20090313244A1 (en) * 2008-06-16 2009-12-17 Serhii Sokolenko System and method for displaying context-related social content on web pages
WO2010050288A1 (ja) * 2008-10-30 2010-05-06 インターナショナル・ビジネス・マシーンズ・コーポレーション サーバシステム、サーバ装置、プログラム、および方法
CN101916298A (zh) * 2010-08-31 2010-12-15 深圳市赫迪威信息技术有限公司 数据库操作方法、设备及系统
CN102289469B (zh) * 2011-07-26 2013-01-30 国电南瑞科技股份有限公司 一种支持通用数据库基于物理隔离设备同步数据的方法

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101826073A (zh) * 2009-03-06 2010-09-08 华为技术有限公司 分布式数据库同步方法、设备及系统
CN101876996A (zh) * 2009-12-01 2010-11-03 广州从兴电子开发有限公司 一种内存数据库到文件数据库的数据同步方法及系统
CN102024040A (zh) * 2010-12-08 2011-04-20 北京握奇数据系统有限公司 数据库同步方法、装置和系统
CN102238178A (zh) * 2011-05-30 2011-11-09 李牧森 基于互联网技术应用的多层次传播信息的客户端桌面媒体

Also Published As

Publication number Publication date
US20150286696A1 (en) 2015-10-08
CN103838757A (zh) 2014-06-04
US9754006B2 (en) 2017-09-05
CN103838757B (zh) 2017-06-09

Similar Documents

Publication Publication Date Title
US9286298B1 (en) Methods for enhancing management of backup data sets and devices thereof
CN105493474B (zh) 用于支持用于同步分布式数据网格中的数据的分区级别日志的系统及方法
CN104050250B (zh) 一种分布式键-值查询方法和查询引擎系统
CN112084258A (zh) 一种数据同步方法和装置
CN102750317B (zh) 数据持久化处理方法、装置及数据库系统
CN107623703B (zh) 全局事务标识gtid的同步方法、装置及系统
JP5686034B2 (ja) クラスタシステム、同期制御方法、サーバ装置および同期制御プログラム
US10545988B2 (en) System and method for data synchronization using revision control
CN102591970A (zh) 一种分布式键-值查询方法和查询引擎系统
CN110837423B (zh) 一种自动导引运输车数据采集的方法和装置
CN103078945B (zh) 对浏览器崩溃数据进行处理的方法与系统
CN106909595B (zh) 一种数据迁移方法及装置
WO2015184925A1 (zh) 分布式文件系统的数据处理方法及分布式文件系统
CN110334145A (zh) 数据处理的方法和装置
WO2022135471A1 (zh) 多版本并发控制和日志清除方法、节点、设备和介质
WO2014079290A1 (zh) 对用户的长关系链数据的处理系统和方法
WO2024109253A1 (zh) 一种数据备份方法、系统和设备
CN116233111A (zh) 一种基于Minio的大文件上传方法
CN104580425A (zh) 一种客户端数据同步方法及系统
US10664349B2 (en) Method and device for file storage
CN111708835A (zh) 区块链数据存储方法及装置
CN106855869B (zh) 一种实现数据库高可用的方法、装置和系统
CN112579556A (zh) 日切数据卸载方法、装置、设备及介质
CN111797352A (zh) 封禁帐号的方法、装置及封禁系统
CN118643102A (zh) 数据下发方法、装置、计算机设备及介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13857320

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 14646794

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 5.10.2015)

122 Ep: pct application non-entry in european phase

Ref document number: 13857320

Country of ref document: EP

Kind code of ref document: A1