WO2015055002A1 - 一种电子书文档的处理方法、终端及电子设备 - Google Patents

一种电子书文档的处理方法、终端及电子设备 Download PDF

Info

Publication number
WO2015055002A1
WO2015055002A1 PCT/CN2014/077413 CN2014077413W WO2015055002A1 WO 2015055002 A1 WO2015055002 A1 WO 2015055002A1 CN 2014077413 W CN2014077413 W CN 2014077413W WO 2015055002 A1 WO2015055002 A1 WO 2015055002A1
Authority
WO
WIPO (PCT)
Prior art keywords
segment
information
data
layout
starting point
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2014/077413
Other languages
English (en)
French (fr)
Inventor
张家方
张磊
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Xiaomi Inc
Original Assignee
Xiaomi Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Xiaomi Inc filed Critical Xiaomi Inc
Priority to KR1020147021518A priority Critical patent/KR20150055600A/ko
Priority to RU2015122426A priority patent/RU2617388C2/ru
Priority to JP2015542159A priority patent/JP6118418B2/ja
Priority to MX2014009068A priority patent/MX2014009068A/es
Priority to BR112014018483A priority patent/BR112014018483A8/pt
Priority to US14/471,702 priority patent/US20150106697A1/en
Publication of WO2015055002A1 publication Critical patent/WO2015055002A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/189Automatic justification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/12Use of codes for handling textual entities
    • G06F40/131Fragmentation of text files, e.g. creating reusable text-blocks; Linking to fragments, e.g. using XInclude; Namespaces
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/103Formatting, i.e. changing of presentation of documents
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/103Formatting, i.e. changing of presentation of documents
    • G06F40/106Display of layout of documents; Previewing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/12Use of codes for handling textual entities
    • G06F40/14Tree-structured documents
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/12Use of codes for handling textual entities
    • G06F40/14Tree-structured documents
    • G06F40/143Markup, e.g. Standard Generalized Markup Language [SGML] or Document Type Definition [DTD]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/205Parsing
    • G06F40/221Parsing markup language streams

Definitions

  • the invention relates to a method for processing an e-book document, a terminal and an electronic device.
  • the application is based on a Chinese patent application with the application number of 201310485775.8 and the application date of 2013/10/16, and claims the priority of the Chinese patent application.
  • the entire disclosure is hereby incorporated by reference.
  • the public disclosure involves the field of data processing, and more specific, and the method of processing and dealing with the documents of electronic and electronic books. Terminal terminal and electrical and electronic sub-equipment equipment. .
  • HHTTMMLL ((HHyyppeertrteexxtt MMaarrkkuupp LLaanngguuaaggee,; super-text text mark language);)
  • HHTTM MLL For the electronic electronic book document file edited with HHTTM MLL, it can be referred to as HHTTMMLL document file.
  • the mobile terminal After the user opens the power-on electronic book by using the mobile terminal, the mobile terminal will read the HHTTMMLL document file of the electronic book.
  • the electronic book is converted into a user account, and the user can view it through the mobile terminal.
  • Page page image image .
  • the main process of the above-mentioned translation process includes the steps of step-by-step analysis, step-by-step steps, and steps to generate pages into objects. Generate a page page image image like steps and so on. .
  • the invention found that the above-mentioned method of dealing with the electronic book of the electrical and electronic books has at least rarely existed in the following questions.
  • Title: 2200 is usually used, and the storage and storage capacity and the ability to calculate and calculate the capacity of the terminal are one-to-one. .
  • the volume capacity of the electronic book is relatively large, then the data of the mobile terminal is read and read at the time when the electronic book is read.
  • the efficiency of the operation efficiency of the step is decreased, and the time period of the long-moving mobile terminal is read and read and the electronic book is read. .
  • the purpose of the design project disclosed in the present disclosure is to provide a method for processing a document file for a type of electronic electronic book,
  • the terminal terminal and the electronic electronic sub-setting 3300 are prepared, so that when the mobile terminal device is read and read, the short-reading and reading time is shortened. And reduce the use of low internal memory. .
  • the present disclosure provides a case for processing an electronic electronic book and a document file.
  • the method method, the method package includes the following steps:
  • the page image is generated based on the layout layout data.
  • the e-book document is divided into a plurality of segments, and each time the e-book operation is performed, only one segment in the e-book document is parsed, and the page layout image is generated by using the layout layout data generated by the segment, so each time
  • the amount of data that needs to be processed during the operation is only the amount of data of one segment, so when the mobile terminal reads the e-book document, the efficiency of processing the data is improved, thereby shortening the time for reading the e-book;
  • the batch processing of the e-book document is realized.
  • the mobile terminal processes a segment of the e-book document, whether the parsing operation or the subsequent operation of generating the page image, the amount of data to be processed is relatively small, so the e-book pair is reduced. The occupation of the mobile terminal memory.
  • the method further includes:
  • the step of segmenting the content of the e-book document in a predetermined segmentation manner includes the following sub-steps:
  • the content of the e-book document is split into a plurality of segments of the same size as the segmentation value.
  • This program implements a specific segmentation approach.
  • the method further includes:
  • This scheme avoids a complete information being split into two segments, and further accurately divides the start and end points of the segment.
  • the method further includes:
  • This scheme avoids a complete information being split into two segments, and further accurately divides the start and end points of the segment.
  • the method further includes:
  • the method further includes:
  • This scheme avoids a complete information being split into two segments, and further accurately divides the start and end points of the segment.
  • the method further includes:
  • the method further includes:
  • an embodiment of the present disclosure further provides a terminal, where the terminal includes: an acquiring module, configured to acquire an e-book document;
  • segmentation module configured to segment the content of the e-book document according to a preset segmentation manner to generate a plurality of segments
  • component module configured to form the plurality of segments into an ordered segment group
  • a selection module for selecting a segment in the segment group and using the segment as the current segment
  • a parsing module configured to parse the content of the current segment, and generate layout layout data
  • a generation module is configured to generate a page image according to the layout layout data.
  • the terminal further includes:
  • a recording module configured to record position information of the current segment in the segment group
  • a first determining module configured to determine whether the data volume of the layout typesetting data is lower than a preset value
  • the first execution module is configured to: when the data volume is lower than the preset value, select the next segment of the current segment according to the location information, and return the next segment as the current segment to the execution parsing module.
  • the segmentation module includes:
  • a segmentation value determining unit configured to determine a segmentation value
  • a splitting unit configured to split the content of the e-book document into a plurality of segments of the same size as the segmentation value.
  • the terminal further includes:
  • a second determining module configured to determine whether the information at the starting point of the segment is complete
  • the second execution module is configured to move forward at the starting point of the segment to determine the starting point of the label when the information at the starting point of the segment is incomplete, and use the starting point of the label as the starting point of the segment.
  • the terminal further includes:
  • a third determining module configured to determine whether the information at the starting point of the segment is complete
  • a third execution module configured to: when the information at the starting point of the segment is incomplete, to the segment at the starting point of the segment The end point moves to determine the end point of the information, and the end point of the information is taken as the starting point of the segment and the ending point of the previous segment.
  • the terminal further includes:
  • a fourth determining module configured to determine whether the information at the end point of the segment is complete
  • a fourth execution module configured to: when the information at the end point of the segment is incomplete, move the end point of the determination information to the next segment at the end point of the segment, and use the end point of the information as the end point of the segment and the next segment The starting point.
  • the terminal further includes:
  • a fifth determining module configured to determine whether the information at the end point of the segment is complete
  • a fifth execution module configured to: when the information at the end point of the segment is incomplete, move the start point of the determination information to the start point of the segment at the end point of the segment, and use the starting point of the information as the end point of the segment and the next The starting point of the fragment.
  • the terminal further includes:
  • a sixth determining module configured to determine whether the typesetting atomic data included in the layout typesetting data is used within a preset time
  • the sixth execution module is configured to delete the typesetting atomic data when the typesetting atomic data is not used within a preset time.
  • the terminal further includes:
  • a seventh determining module configured to determine whether the memory occupied by the typesetting atomic data included in the layout type data is greater than a preset value
  • the seventh execution module is configured to delete the typesetting atomic data when the memory occupied by the typesetting atomic data is greater than a preset value.
  • an embodiment of the present disclosure further provides an electronic device including a memory, and one or more programs, wherein one or more programs are stored in the memory, and Configuring to execute the one or more programs by one or more processors includes instructions for:
  • a page image is generated based on the layout layout data.
  • FIG. 1 is an exemplary flowchart of a method for processing an e-book document according to an embodiment of the present disclosure
  • FIG. 2 is an exemplary flowchart of a method for processing an e-book document according to an embodiment of the present disclosure
  • FIG. 3 is an exemplary flowchart of a method for processing an e-book document according to an embodiment of the present disclosure
  • FIG. 5 is a schematic block diagram of a terminal according to an embodiment of the present disclosure
  • FIG. 6 is a schematic block diagram of a segmentation module according to an embodiment of the present disclosure.
  • FIG. 7 is a schematic block diagram of another terminal according to an embodiment of the present disclosure.
  • FIG. 8 is a schematic block diagram of still another terminal according to an embodiment of the present disclosure.
  • FIG. 9 is a schematic block diagram of still another terminal according to an embodiment of the present disclosure.
  • FIG. 10 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure. detailed description
  • the process of reading the e-book document by the mobile terminal mainly includes the following steps: parsing step: the mobile terminal parses the e-book document into layout layout data; the full pagination step: full page processing on the layout layout data to obtain a page frame; generating a page object Step: Generate a corresponding page object by using the page frame and the layout layout data; Generate a page image step: Render the page object to generate a page image.
  • parsing step the mobile terminal parses the e-book document into layout layout data
  • the full pagination step full page processing on the layout layout data to obtain a page frame
  • generating a page object Step Generate a corresponding page object by using the page frame and the layout layout data
  • Generate a page image step Render the page object to generate a page image.
  • the embodiment of the present disclosure provides a method, a terminal, and an electronic device for processing the e-book, so that the time for reading the e-book by the mobile terminal is short, and the memory usage is low. Since the processing method of the above electronic document, the specific implementation of the terminal and the electronic device exist in various ways, the following detailed description is given by way of specific embodiments:
  • FIG. 1 illustrates a method for processing an e-book document, and the method may include the following steps:
  • Step 101 Obtain an e-book document
  • the mobile terminal After the user opens the e-book through the mobile terminal, the mobile terminal reads the e-book document into the internal memory. At this time, the mobile terminal has acquired the e-book document.
  • the e-book document may be a streaming e-book document.
  • the streaming e-book document described herein refers to the described text, and the image has no fixed layout position, and the layout parameter (such as the layout width). , font size, line spacing, etc.) When you change, you need to re-layout the layout to adapt to the new typesetting parameters of the e-book document.
  • the streaming e-book document includes a document such as an HTML document, and the HTML document is composed of tags, so that when the mobile terminal acquires the HTML document of the e-book, the tag constituting the HTML document is also acquired.
  • Step 102 Segment the content of the e-book document according to a preset segmentation manner to generate a plurality of segments.
  • the preset segmentation mode may have multiple implementation forms, for example: determining a segmentation value, and splitting the content of the e-book document into a plurality of segments having the same size as the segmentation value.
  • the segmentation value represents the size of each segment after the e-book document is split, so the segmentation value may be preset, or may be set according to the user of the mobile terminal, or may be calculated according to experiments. .
  • Step 103 Form a plurality of segments into an ordered segment group
  • the e-book document is divided into a plurality of segments.
  • Step 104 Select a segment in the segment group, and use the segment as the current segment;
  • any one of the segments can be selected according to the user's needs, so that the selected segment can be processed in a subsequent step.
  • Step 105 Parse the content in the current segment to generate layout and layout data.
  • the content of a certain segment is only a part of the content of the complete e-book document, and the data to be parsed has much less content than the complete e-book document, so the parsing speed is obviously faster, and the amount of data is reduced, thereby The memory used is also greatly reduced.
  • the following steps may be further included: determining whether the typesetting atomic data included in the layout typesetting data is used within a preset time, and if not used, deleting the typesetting atomic data. . If the layout layout data has not been used for a long time, in order to save the memory occupied by the layout data, the layout atom data may be deleted first, and if necessary, the layout atom data may be newly generated.
  • the following steps may be further included: determining whether the memory occupied by the typesetting atomic data included in the layout type data is greater than a preset value, and if the value is greater than a preset value, deleting the typesetting atom data. If the layout and layout data is too much, the processing speed of the subsequent steps may be too slow, and all the layout layout data may be deleted, or a predetermined number of layout layout data may be deleted.
  • Step 106 Generate a page image according to the layout layout data.
  • the data volume of the layout typesetting data is generated by one segment, and the data volume of the layout typesetting data is relatively small. In the process of generating the page image, both the time occupation and the memory occupation are greatly reduced.
  • step 106 the following may further include the following three steps: 1. Performing paging processing on the layout layout data to generate a page layout framework; 2. generating a page object according to the page layout frame and the layout layout data; 3. rendering the page object Generate a page image.
  • the e-book document is divided into a plurality of segments, and each time the e-book operation is performed, only one segment in the e-book document is parsed, and the page layout image is generated by using the layout layout data generated by the segment.
  • FIG. 2 is a processing method of another e-book document, and the method includes the following steps: Step 201: Acquire an e-book document;
  • the mobile terminal After the user opens the e-book through the mobile terminal, the mobile terminal reads the e-book document into the internal memory. At this time, the mobile terminal has acquired the e-book document.
  • the e-book document may be an HTML document, and the HTML document is composed of tags. Therefore, when the mobile terminal acquires the HTML document of the e-book, the tag constituting the HTML document is also acquired.
  • Step 202 Segment the content of the e-book document according to a preset segmentation manner to generate a plurality of segments.
  • the preset segmentation mode may have multiple implementation forms.
  • the following describes a segmentation method, which is: determining a segmentation value, and splitting the content of the e-book document into a plurality of segments having the same size as the segmentation value.
  • the segmentation value represents the size of each segment after the e-book document is split, so the segmentation value may be preset, or may be set according to the user of the mobile terminal, or may be calculated according to experiments. .
  • Step 203 Combine multiple segments into an ordered segment group
  • the e-book document is divided into a plurality of segments.
  • Step 204 Select a segment in the segment group, and use the segment as the current segment;
  • any one of the segments can be selected according to the user's needs, so that the selected segment can be processed in a subsequent step.
  • Step 205 Parse the content in the current segment to generate layout layout data.
  • the content of a certain segment is only a part of the content of the complete e-book document, and the data to be parsed has much less content than the complete e-book document, so the parsing speed is obviously faster, and the amount of data is reduced, thereby The memory used is also greatly reduced.
  • Step 206 Record location information of the current segment in the segment group.
  • the location information may be recorded in bytes, the location information may be recorded by using an identifier, and the like.
  • the position information of the segment can be recorded in a byte offset manner, and the byte offset is in bytes.
  • the position information is not specifically limited herein, and it is within the protection scope of the present disclosure as long as the position of the segment in the segment group can be recorded.
  • Step 207 determining whether the data volume of the layout typesetting data is lower than a preset value, and if so, executing step 208, otherwise, performing step 209;
  • the layout layout data is generated by the current segment analysis, and the data amount of the layout layout data generated by each segment analysis is certain. If the data volume of the layout layout data that has been generated is sufficient to generate the page image, then the layout layout is utilized. The data generates a page image; if the amount of layout layout data that has been generated is insufficient to generate a page map For example, then it is necessary to parse the next segment to generate layout layout data, and then merge the previous layout layout data with the layout layout data generated by the next segment to generate a page image.
  • Step 208 Select the next segment of the current segment according to the location information, and use the next segment as the current segment, and return to step 205;
  • the location information of the current segment has been recorded in step 206, so the location information of the next segment can be determined by the location information of the current segment, so that the content in the next segment can be selected.
  • the next segment is then taken as the current pending segment, and the process returns to step 205.
  • Step 209 Generate a page image according to the layout layout data.
  • the data volume of the layout typesetting data is generated by one segment, and the data volume of the layout typesetting data is relatively small. In the process of generating the page image, both the time occupation and the memory occupation are greatly reduced.
  • the difference from the embodiment shown in FIG. 1 is that it is judged whether the data amount of the layout layout data that has been generated is enough to generate a page image, and if so, the layout is based on the existing layout.
  • the data generates a page image; otherwise, the next segment is parsed, and the page image is generated based on the existing layout layout data and the layout layout data generated by the next segment.
  • the starting point and the ending point of the segment can be further accurately divided.
  • the e-book document is specifically an HTML document
  • the starting point of the segment is the starting position of the segment
  • the ending point of the segment is the ending position of the segment
  • the HTML document is mainly composed of tags
  • the tag is composed of two angle brackets, that is, the tag is composed of The left angle bracket " ⁇ ” and the right angle bracket ">” are composed, and the contents of the label are between the two angle brackets.
  • the label must include the left angle bracket " ⁇ ” and the right angle bracket “>”, so that the label is complete, therefore, after segmentation, there is no need for a fragment containing only the left angle bracket or only the right tip
  • the non-correspondence of the parentheses indicates that a label is split into two fragments, so that the label cannot be parsed later.
  • FIG. 3 is a processing method of another e-book document, and the method includes the following steps: Step 301: Acquire an e-book document;
  • the mobile terminal After the user opens the e-book through the mobile terminal, the mobile terminal reads the e-book document into the internal memory. At this time, the mobile terminal has acquired the e-book document.
  • the e-book document may be an HTML document, and the HTML document is composed of tags. Therefore, when the mobile terminal acquires the HTML document of the e-book, the tag constituting the HTML document is also acquired.
  • Step 302 Segment the content of the e-book document according to a preset segmentation manner to generate multiple segments.
  • the preset segmentation mode may have multiple implementation forms.
  • the following describes a segmentation method, which is: determining a segmentation value, and splitting the content of the e-book document into a plurality of segments having the same size as the segmentation value.
  • the segmentation value represents the size of each segment after the e-book document is split, so the segmentation value may be preset, or may be set according to the user of the mobile terminal, or may be calculated according to experiments. .
  • Step 303 Determine whether the information at the starting point of the segment is complete, and if yes, perform step 305; otherwise, execute Step 304;
  • Step 304 Move a start point of the determination information to a segment at a starting point of the segment, and use a starting point of the information as a starting point of the segment and an ending point of the previous segment;
  • the first half of the information is divided into the previous segment, and the second half of the information is left at the starting point in the segment.
  • the adjustment method of step 304 in order to delete the first half of the information in the previous segment, and add the first half of the information to the segment.
  • the starting point of the segment is the same point as the ending point of the previous segment.
  • the second implementation is specifically: moving the end point of the segment to the end point of the segment to determine the end point of the information, and using the end point of the information as the starting point of the segment and the ending point of the previous segment.
  • the second way is to delete the second half of the information in the segment and add the second half of the message to the end point in the previous segment.
  • step 304 in order to ensure the integrity of the information in the segment, not only the method of step 304 can be adopted, but also the starting point of other information can be determined as the starting point of the segment and the ending point of the previous segment, although the starting point or ending of the segment in this way The point offset is excessive, but the problem of information integrity can also be solved. Therefore, the embodiment provided by the present disclosure is not limited to the method of step 304. Other methods similar to the principle of step 304 can be used as long as the information integrity can be achieved. There are no specific restrictions here.
  • Step 305 Combine multiple segments into an ordered segment group
  • the e-book document is divided into a plurality of segments.
  • Step 306 Select a segment in the segment group, and use the segment as the current segment;
  • any one of the segments can be selected according to the user's needs, so that the selected segment can be processed in a subsequent step.
  • Step 307 Parse the content in the current segment to generate layout layout data.
  • the content of a certain segment is only a part of the content of the complete e-book document, and the data to be parsed has much less content than the complete e-book document, so the parsing speed is obviously faster, and the amount of data is reduced, thereby The memory used is also greatly reduced.
  • Step 308 Generate a page image according to the layout layout data.
  • the data volume of the layout typesetting data is generated by one segment, and the data volume of the layout typesetting data is relatively small. In the process of generating the page image, both the time occupation and the memory occupation are greatly reduced.
  • the first method is to judge whether the information at the end point of the segment is complete. If it is not complete, move to the next segment at the end point of the segment to determine the end point of the information, and the information. The end point is the end point of the segment and the starting point of the next segment.
  • the second method is: determining whether the information at the end point of the segment is complete, if not complete, moving the starting point of the determining information to the starting point of the segment at the end point of the segment, and using the starting point of the information as the ending point of the segment and The starting point of the next segment.
  • the principle is similar to the steps 303 and 304. For details, refer to step 303 and step 304, and details are not described herein.
  • FIG. 4 shows a specific application example provided by the present disclosure.
  • This application example describes a specific processing method when the e-book document is an HTML document 1.
  • the processing method of the HTML document 1 includes the following steps: Step 1: Obtain an HTML document 1;
  • Step 2 segment the content of the HTML document 1 according to a preset segmentation manner to generate a plurality of HTML segments.
  • the maximum feature of the HTML document 1 parsing process is to parse the HTML fragment and parse a small HTML fragment. The time is within the acceptable range; if only one small HTML fragment is processed per parsing process, the single parsing time can be shortened; therefore, the HTML is segmented and parsed here.
  • the size value m of the segment for example, the manner of artificial setting or the manner of calculation.
  • the following focuses on how to determine the size of the fragment m:
  • HTML document 1 Since the parsing time of the HTML document 1 is directly determined by the number of HTML nodes, as the HTML document 1 is gradually increased, the number of HTML nodes is also increased, and the depth traversal of the nodes is gradually increased; thus, the HTML document is known. 1
  • the resolution curve with the size of the HTML document 1 should be a positive incremental curve. In general, when HTML document 1 is small, the number of HTML nodes is small. For modern computers, HTML document 1 parsing can complete deep traversal in a short time (such as facing 10K and 50K HTML documents1, The parsing time does not make much difference.) When the size of HTML document 1 reaches a certain level, the number of HTML nodes will increase non-linearly, which leads to the non-linear growth of HTML document 1 parsing time.
  • the size of the HTML document 1 corresponding to the HTML document 1 parsing performance inflection point can be determined within the interval [M, N].
  • the interval [M, N) there are two criteria to consider at this time: a, HTML document 1 segment should not be too much, otherwise it will increase the segment management complexity, which makes the segment size value m b, can not be too small; b, the parsing time of a single HTML fragment should not be too long, otherwise there will be shortcomings such as the user waiting time is too long, so that the size value of the segment m can not be too large; according to the above two criteria, the interval can be removed first [M , the minimum value M and the maximum value N in N] get a new interval, and then remove the maximum and minimum values of the new interval according to the above two criteria, and then cycle down. If the last value remains, the value is used as performance.
  • the size of the HTML document 1 corresponding to the inflection point that is, the size value m of the fragment, if the last two numbers are left, the intermediate value of the two numbers is taken as the size of the HTML document 1 corresponding to the performance inflection point, that is, the size value of the fragment! ⁇
  • Step three forming a plurality of segments into an ordered segment group
  • the biggest drawback of the parsing process is that it cannot be backtracked, ie the parsing must be performed in the physical order of the HTML document 1. If the parsing starts from the non-starting point of HTML document 1, there will be a problem that the HTML document 1 has a poor context, which causes the label node to be lost. Therefore, the parsing state interval 2 of HTML document 1 is used here to overcome this problem.
  • the parsing state of HTML document 1 will record the byte offset of the HTML document 1 in it, and the segmentation result of HTML document 1 is also described by the byte offset, so according to this byte offset
  • the corresponding relationship can complete the one-to-one correspondence of the HTML fragment to the parsing interval. Therefore, the first HTML fragment corresponds to the first parsing interval, and the parsing start state (the byte offset is the beginning of the first fragment) takes the initial parsing state, and the parsing end state is the completion of the first HTML.
  • the parsing state of the fragment (the byte offset is the end of the first segment).
  • the parsing state inheritance here is the cloning process of the parsing state, which is the byte offset of the HTML document 1 in which the parsing state is located (used to record the parsing state in the current position of the HTML document 1), and the stack information of the parent HTML tag node.
  • the context of the parsing (used for remembering Recording the parsing state falls within the text node of the HTML document 1 and so on.)
  • the original copy is copied to another parsing state.
  • each HTML fragment inherits the parsing state from the last HTML fragment in the process of participating in the parsing of the HTML document 1, and inherits the context of the previous HTML fragment, that is, overcomes the incomplete relationship between the upper and lower parts. problem.
  • Step 4 Select a segment in the segment group, and use the segment as the current segment;
  • Step 5 parsing the content of the current segment, and generating layout layout data 3;
  • Step 6 Generate a page image according to the layout type data 3 .
  • HTML full paging process
  • the full paging process is processed from the very beginning of the HTML physical space, and the layout driver process is called for each HTML fragment in turn from front to back; for each HTML fragment, when the paging process reaches the last page of the fragment, In the case of a half page, in order to ensure that the last page of the current segment is continuous with the next segment, the HTML data of the next segment is immediately parsed, and then the paging is started from the beginning of the last page of the current segment. Insufficient to use the layout of the next segment of the atom to avoid breakage between the segments.
  • the temporary parsing process is used to obtain the required page.
  • the full HTML context is not used, when the full paging process reaches the required page, it turns to the full page. Paging results to correct problems with parsing HTML from non-starting points.
  • the page number cannot be jumped because the page number may exceed the number of pages that have been paged in the full paging process. At this point, the percentage of the HTML physical space can be jumped.
  • the full page is completed, it can be passed. The page number is jumped.
  • the solution provided by the present disclosure consumes a smaller storage space in the large-page HTML document 1 in the full paging process, and reduces the acquisition time of the front partial full page result in the HTML document 1; using temporary parsing to make the HTML before the full page arrival The performance of the jump is improved, and problems that may exist in the temporary parsing are corrected after the full page is reached.
  • Full page-by-page results after full-page to request point, ensuring that the user does not destroy the book-style structure when reading the e-book with the large-size HTML document 1 as the main, and presents the content to the user in the form of the most respectful original book;
  • FIG. 5 is a terminal, where the terminal includes an obtaining module 11 for acquiring an e-book document, and a segmentation module 12 for using the content of the e-book document according to a preset segmentation manner. Segmentation, generating a plurality of segments; component module 13 for grouping a plurality of segments into an ordered segment group; and selecting module 14 for selecting a segment in the segment group as the current segment; parsing module 15, For parsing the content of the current segment, generating layout layout data; and generating a module 16 for generating a page image according to the layout layout data.
  • FIG. 6 shows a specific structure of the segmentation module 12, and the segmentation module 12 includes a segmentation value determining unit 121 and a splitting unit 122.
  • the segmentation value determining unit 121 is configured to determine the segmentation value
  • the segmentation unit 122 is configured to split the content of the e-book document into a plurality of segments having the same size as the segmentation value.
  • the e-book document is divided into a plurality of segments by the segmentation module 12, and each time the e-book is operated, the parsing module 15 parses only one segment in the e-book document, and the generation module 16 utilizes the The layout layout data generated by the segment generates a page image, so the amount of data that needs to be processed during the operation is only the data amount of one segment, so when the mobile terminal reads the e-book document, the efficiency of processing the data is improved, and further The time for reading the e-book is shortened; moreover, the segmentation operation realizes batch processing of the e-book document, and when the mobile terminal processes a segment of the e-book document, whether it is an analysis operation or a subsequent operation of generating a page image, The amount of data processed is relatively small, so the occupation of the memory of the mobile terminal by the e-book is reduced.
  • the segment value obtained by the segment value determining unit 121 represents the size of each segment after the e-book document is split, so the segment value may be preset or may be based on The mobile terminal user can also set it according to the experiment.
  • the splitting unit 122 then uses the segmentation value to split the e-book document.
  • FIG. 7 is another terminal, where the terminal includes an obtaining module 21 for acquiring an e-book document, and a segmentation module 22 for using the e-book document according to a preset segmentation manner.
  • Content segmentation generating a plurality of segments; component module 23, configured to form a plurality of segments into an ordered segment group; and selecting module 24, configured to select a segment in the segment group, using the segment as a current segment; parsing module 25
  • the recording module 26 is configured to record the position information of the current segment in the segment group;
  • the first determining module 27 is configured to determine whether the data volume of the layout typesetting data is lower than the pre-predetermined data.
  • the first execution module 28 is configured to: when the data volume is lower than the preset value, select the next segment of the current segment according to the location information, and return the next segment as the current segment to the execution parsing module 25; , used to generate a page image based on layout layout data.
  • FIG. 8 is another terminal, where the terminal includes an obtaining module 31 for acquiring an e-book document, and a segmentation module 32 for using the e-book document according to a preset segmentation manner.
  • Segmentation of the content generating a plurality of segments; a second determining module 33, configured to determine whether the information at the starting point of the segment is complete; and a second executing module 34, configured to: when the information at the starting point of the segment is incomplete, in the segment The starting point moves forward to determine the starting point of the label, and the starting point of the label is used as the starting point of the segment; the component module 35 is configured to form the plurality of segments into an ordered group of segments; the selecting module 36 is configured to A segment is selected from the group, and the segment is used as the current segment; a parsing module 37 is configured to parse the content of the current segment to generate layout layout data; and a generating module 38 is configured to generate the page image according to the layout layout data.
  • the difference from the embodiment shown in FIG. 5 is that, in order to prevent a complete information from being split into two segments, the starting point and the ending point of the segment are further performed. Precise division.
  • the third determining module replaces the second determining module 33, and replaces the second executing module 34 with the third executing module.
  • the third determining module is configured to determine whether the information at the starting point of the segment is complete; When it is judged that the information at the start point of the segment is incomplete, the end point of the determination information is moved to the end point of the segment at the start point of the segment, and the end point of the information is taken as the start point of the segment and the end point of the previous segment.
  • the second determining module 33 is replaced by the fourth determining module, and the second executing module 34 is replaced by the fourth executing module.
  • the fourth determining module is configured to determine whether the information at the end point of the segment is complete.
  • a fourth execution module configured to: when the information at the end point of the segment is incomplete, move the end point of the determination information to the next segment at the end point of the segment, and use the end point of the information as the end point of the segment and the next The starting point of the fragment.
  • the second determining module 33 is replaced by the fifth determining module, and the second executing module 34 is replaced by the fifth executing module.
  • the fifth determining module is configured to determine whether the information at the end point of the segment is complete.
  • a fifth execution module configured to: when the information at the end point of the segment is incomplete, move the start point of the determination information to the start point of the segment at the end point of the segment, and use the starting point of the information as the end point and the bottom of the segment The starting point of a fragment.
  • the terminal includes an obtaining module 41 for acquiring an e-book document, and a segmentation module 42 for using an e-book document according to a preset segmentation manner.
  • Content segmentation generating a plurality of segments; component module 43 for grouping a plurality of segments into an ordered segment group; and selecting module 44, for selecting a segment in the segment group, using the segment as a current segment; parsing module 45
  • the sixth determining module 46 is configured to determine whether the layout atomic data included in the layout layout data is used within a preset time; the sixth execution module 47, configured to When the typesetting atomic data is not used within the preset time, the typesetting atomic data is deleted.
  • the generating module 48 is configured to generate a page image according to the layout layout data.
  • the parsing module 45 after the parsing module 45 generates the layout layout data, in order to save the memory of the mobile terminal, it can be implemented by the sixth judging module 46 and the sixth executing module 47.
  • the typesetting atomic data may be deleted first, and if necessary, the typesetting atomic data may be newly generated.
  • the memory of the seventh determining module 46 can be replaced by the seventh determining module, and the sixth executing module 47 is replaced with the seventh executing module.
  • the seventh determining module is configured to determine whether the memory occupied by the typesetting atomic data included in the layout typesetting data is greater than a preset value; and the seventh executing module is configured to delete when the memory occupied by the typesetting atomic data is greater than a preset value Typesetting atomic data.
  • an embodiment of the present disclosure further provides an electronic device, which is used to implement the processing method of the e-book document provided in the foregoing embodiment, specifically:
  • the electronic device 1300 may include an RF (Radio Frequency) circuit 1310, a memory 1320 including one or more computer readable storage media, an input unit 1330, a display unit 1340, a sensor 1350, an audio circuit 1360, and a short-range wireless transmission module. 1370, including a processor 1380 having one or more processing cores, and a power supply 1390 and the like. It will be understood by those skilled in the art that the electronic device structure shown in Fig. 13 does not constitute a limitation on the electronic device, and may include more or less components than those illustrated, or some components may be combined, or different component arrangements. among them:
  • the RF circuit 1310 can be used for receiving and transmitting signals during and after receiving or transmitting information, in particular, receiving downlink information of the base station, and then processing it by one or more processors 1380; in addition, transmitting uplink data to the base station .
  • the RF circuit 1310 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, an LNA (Low Noise Amplifier). , duplexer, etc.
  • SIM Subscriber Identity Module
  • the RF circuit 1310 can also communicate with the network and other devices through wireless communication.
  • the wireless communication may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access). , Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), e-mail, SMS (Short Messaging Service), and the like.
  • GSM Global System of Mobile communication
  • GPRS General Packet Radio Service
  • CDMA Code Division Multiple Access
  • WCDMA Wideband Code Division Multiple Access
  • LTE Long Term Evolution
  • e-mail Short Messaging Service
  • Memory 1320 can be used to store software programs as well as modules.
  • the processor 1380 executes various functional applications and data processing by running software programs and modules stored in the memory 1320.
  • the storage 1320 may mainly include a storage program area and a storage data area, wherein the storage program area may store an operating system, an application required for at least one function (such as a sound playing function, an image playing function, etc.), and the like; the storage data area may be stored according to The data created by the use of the electronic device 1300 (such as audio data, phone book, etc.) and the like.
  • memory 1320 can include high speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid state storage device. Accordingly, memory 1320 can also include a memory controller to provide access to memory 1320 by processor 1380 and input unit 1330.
  • Input unit 1330 can be used to receive input numeric or character information, as well as to generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
  • the input unit 1330 can include A touch sensitive surface 1331 and other input devices 1332 are included.
  • Touch-sensitive surface 1331 also known as a touch display or touchpad, can collect touch operations on or near the user (eg, the user uses a finger, stylus, etc., on any touch-sensitive surface 1331 or The operation near the touch-sensitive surface 1331), and the corresponding connecting device is driven according to a preset program.
  • the touch-sensitive surface 1331 may include two parts of a touch detection device and a touch controller.
  • the touch detection device detects the touch orientation of the user, and detects a signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts the touch information into contact coordinates, and sends the touch information
  • the processor 1380 is provided and can receive commands from the processor 1380 and execute them.
  • the touch sensitive surface 1331 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic waves.
  • the input unit 1330 can also include other input devices 1332.
  • other input devices 1332 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, switch buttons, etc.), trackballs, mice, joysticks, and the like.
  • Display unit 1340 can be used to display information entered by the user or information provided to the user and various graphical user interfaces of electronic device 1300, which can be comprised of graphics, text, icons, video, and any combination thereof.
  • the display unit 1340 can include a display panel 1341.
  • the display panel 1341 can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like.
  • the touch-sensitive surface 1331 may be overlaid on the display panel 1341.
  • the touch-sensitive surface 1331 When the touch-sensitive surface 1331 detects a touch operation on or near it, it is transmitted to the processor 1380 to determine the type of the touch event, and then the processor 1380 is The type of touch event provides a corresponding visual output on display panel 1341.
  • the touch-sensitive surface 1331 and the display panel 1341 are implemented as two separate components to implement input and input functions, in some embodiments, the touch-sensitive surface 1331 can be integrated with the display panel 1341 to implement input. And output function.
  • the electronic device 1300 can also include at least one type of sensor 1350, such as a light sensor, motion sensor, and other sensors.
  • the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 1341 according to the brightness of the ambient light, and the proximity sensor may close the display panel 1341 when the electronic device 1300 moves to the ear. And / or backlight.
  • the gravity acceleration sensor can detect the magnitude of acceleration in all directions (usually three axes). When it is stationary, it can detect the magnitude and direction of gravity.
  • the electronic device 1300 can also be configured with gyroscopes, barometers, hygrometers, thermometers, infrared sensors and other sensors, here No longer.
  • An audio circuit 1360, a speaker 1361, and a microphone 1362 provide an audio interface between the user and the electronic device 1300.
  • the audio circuit 1360 can transmit the converted electrical data of the received audio data to the speaker 1361, and convert it into a sound signal output by the speaker 1361; on the other hand, the microphone 1362 converts the collected sound signal into an electrical signal, by the audio circuit 1360. After receiving, it is converted into audio data, and then processed by the audio data output processor 1380, sent to another terminal via the RF circuit 1310, or outputted to the memory 1320 for further processing.
  • the audio circuit 1360 may also include an earbud jack to provide communication of the peripheral earphones with the electronic device 1300.
  • the short-range wireless transmission module 1370 may be a WIFI (wireless fidelity) module or a Bluetooth module.
  • the electronic device 1300 can help the user to send and receive emails and browse the network through the short-range wireless transmission module 1370. Pages and access to streaming media, etc., which provide users with wireless broadband Internet access.
  • FIG. 13 shows the short-range wireless transmission module 1370, it can be understood that it does not belong to the essential configuration of the electronic device 1300, and may be omitted as needed within the scope of not changing the essence of the invention.
  • the processor 1380 is a control center for the electronic device 1300 that connects various portions of the entire electronic device using various interfaces and lines, by running or executing software programs and/or modules stored in the memory 1320, and recalling stored in the memory 1320. The data, performing various functions and processing data of the electronic device 1300, thereby performing overall monitoring of the electronic device.
  • the processor 1380 may include one or more processing cores.
  • the processor 1380 may integrate an application processor and a modem processor, where the application processor mainly processes an operating system, a user interface, an application, and the like.
  • the modem processor primarily handles wireless communications. It will be appreciated that the above described modem processor may also not be integrated into the processor 1380.
  • the electronic device 1300 further includes a power source 1390 (such as a battery) for powering various components.
  • the power source can be logically connected to the processor 1380 through a power management system to manage functions such as charging, discharging, and power management through the power management system.
  • Power supply 1390 may also include any one or more of a DC or AC power source, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and the like.
  • the electronic device 1300 may further include a camera, a Bluetooth module, and the like, and details are not described herein.
  • the display unit of the electronic device 1300 is a touch screen display.
  • the electronic device 1300 also includes a memory, and one or more programs, one or more of which are stored in the memory and configured to be executed by one or more processors.
  • the above one or more programs include instructions for executing an e-book document processing method, including:
  • a page image is generated based on the layout layout data.
  • FIG. 1 to FIG. 10 are only preferred embodiments of the present disclosure, and those skilled in the art can design more embodiments based on this, and therefore are not described herein.
  • RAM random access memory
  • ROM read only memory
  • electrically programmable ROM electrically erasable programmable ROM
  • registers hard disk, removable disk, CD-ROM, or technical field Any other form of storage medium known.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Artificial Intelligence (AREA)
  • Computational Linguistics (AREA)
  • General Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Multimedia (AREA)
  • Document Processing Apparatus (AREA)
  • User Interface Of Digital Computer (AREA)
  • Editing Of Facsimile Originals (AREA)
  • Processing Or Creating Images (AREA)
  • Character Input (AREA)

Abstract

公开了一种电子书文档的处理方法、终端及电子设备,该方法包括:获取电子书文档(S101);按照预设分段方式对电子书文档的内容分段,生成多个片段(S102);将多个片段组成一个有序的片段组(S103);在片段组中选择一个片段,将片段作为当前片段(S104);解析当前片段的内容,生成布局排版数据(S105);根据布局排版数据生成页面图像(S106)。每次对电子书操作时,仅解析电子书文档中的一个片段,每次在操作过程中需要处理的数据量仅为一个片段的数据量,所以处理数据的效率就得到提升,进而缩短了读取电子书的时间;而且,移动终端在处理电子书文档的一个片段时,无论是解析操作还是后续的生成页面图像的操作,需处理的数据量都相对较小,所以降低电子书对移动终端内存的占用。

Description

一种电子书文档的处理方法、 终端及电子设备 本申请基于申请号为 201310485775.8、 申请日为 2013/10/16的中国专利申请提出, 并 要求该中国专利申请的优先权, 该中国专利申请的全部内容在此引入本申请作为参考。 技术领域
本本公公开开涉涉及及数数据据处处理理领领域域,, 更更具具体体的的说说,, 涉涉及及电电子子书书文文档档的的处处理理方方法法、、 终终端端及及电电子子设设 备备。。
1100 背背景景技技术术
随随着着移移动动终终端端的的不不断断普普及及,, 在在移移动动终终端端上上阅阅读读和和编编辑辑电电子子书书也也越越来来越越普普及及,, 在在某某些些场场 合合,, 部部分分电电子子书书已已经经替替代代了了纸纸质质图图书书,, 成成为为日日常常生生活活中中主主要要的的阅阅读读手手段段。。 用用于于阅阅读读电电子子书书 的的移移动动终终端端有有很很多多,, 例例如如,, 智智能能手手机机、、 平平板板电电脑脑或或电电子子阅阅读读器器等等等等。。
目目前前,, 电电子子书书文文档档主主要要采采用用 HHTTMMLL((HHyyppeertrteexxtt MMaarrkkuupp LLaanngguuaaggee,,超超文文本本标标记记语语言言;;))来来编编 1155 辑辑,, 对对于于采采用用 HHTTMMLL编编辑辑的的电电子子书书文文档档可可以以简简称称为为 HHTTMMLL文文档档。。 在在用用户户通通过过移移动动终终端端打打 开开电电子子书书以以后后,, 移移动动终终端端会会读读取取电电子子书书的的 HHTTMMLL文文档档到到内内存存中中,, 对对 HHTTMMLL文文档档翻翻译译后后使使 电电子子书书转转换换成成用用户户可可通通过过移移动动终终端端观观看看的的页页面面图图像像。。 上上述述翻翻译译的的过过程程主主要要包包括括解解析析步步骤骤、、 全全分分页页步步骤骤、、 生生成成页页面面对对象象步步骤骤和和生生成成页页面面图图像像步步骤骤等等。。
在在研研究究和和实实践践过过程程中中,, 发发明明人人发发现现上上述述处处理理电电子子书书的的方方式式至至少少存存在在以以下下问问题题:: 2200 通通常常,, 一一个个移移动动终终端端的的存存储储能能力力和和计计算算能能力力均均是是有有限限制制的的。。 如如果果电电子子书书的的容容量量相相对对较较 大大,, 那那么么移移动动终终端端读读取取电电子子书书时时处处理理的的数数据据就就会会占占用用很很多多的的存存储储资资源源,, 从从而而导导致致读读取取电电子子 书书文文档档时时,, 解解析析和和全全分分页页步步骤骤的的运运行行效效率率下下降降,, 延延长长移移动动终终端端读读取取电电子子书书的的时时间间。。 而而且且,, 电电子子书书容容量量越越大大,, 越越可可能能导导致致在在解解析析和和全全分分页页的的步步骤骤消消耗耗过过多多的的内内存存资资源源,, 从从而而使使移移动动终终 端端运运行行效效率率下下降降,, 严严重重时时,, 移移动动终终端端可可能能出出现现卡卡顿顿或或死死机机等等现现象象。。
2255 因因此此,, 如如何何提提供供一一种种读读取取时时间间短短且且内内存存占占用用低低的的电电子子书书高高效效率率处处理理方方法法,, 成成为为目目前前最最 需需要要解解决决的的问问题题。。 发发明明内内容容
有有鉴鉴于于此此,, 本本公公开开的的设设计计目目的的在在于于,, 提提供供一一种种电电子子书书文文档档的的处处理理方方法法、、 终终端端及及电电子子设设 3300 备备,, 使使得得移移动动终终端端在在读读取取电电子子书书时时,, 可可缩缩短短读读取取时时间间,, 且且降降低低内内存存的的占占用用。。
根根据据本本公公开开实实施施例例的的第第一一方方面面,, 本本公公开开实实施施例例提提出出一一种种电电子子书书文文档档的的处处理理方方法法,, 该该方方 法法包包括括以以下下步步骤骤::
获获取取电电子子书书文文档档;;
按按照照预预设设分分段段方方式式对对电电子子书书文文档档的的内内容容分分段段,, 生生成成多多个个片片段段;; 在片段组中选择一个片段, 将片段作为当前片段;
解析当前片段的内容, 生成布局排版数据;
根据布局排版数据生成页面图像。
本公开实施例中, 将电子书文档分成多个片段, 每次对电子书操作时, 仅解析电子书 文档中的一个片段, 并利用该片段生成的布局排版数据生成页面图像, 所以每次在操作过 程中需要处理的数据量仅为一个片段的数据量, 所以在移动终端读取电子书文档时, 处理 数据的效率就得到提升, 进而缩短了读取电子书的时间; 而且, 分段操作实现对电子书文 档分批次处理, 移动终端在处理电子书文档的一个片段时, 无论是解析操作还是后续的生 成页面图像的操作,需处理的数据量都相对较小,所以降低电子书对移动终端内存的占用。
在一个实施例中, 在生成布局排版数据的步骤之后, 方法还包括:
记录当前片段在片段组中的位置信息;
判断布局排版数据的数据量是否低于预设值, 如果数据量低于预设值, 根据位置信息 选择当前片段的下一个片段, 并将下一个片段作为当前片段, 返回解析当前片段内的内容 的步骤。
本方案实现了保证布局排版数据的数据量足以生成页面图像。
在一个实施例中, 按照预设分段方式对电子书文档的内容进行分段的步骤, 包括下述 子步骤:
确定分段值;
将电子书文档的内容拆分成多个与分段值大小相同的片段。
本方案实现了具体分段的方式。
在一个实施例中, 方法还包括:
判断片段的起始点处的信息是否完整, 如果不完整, 在片段的起始点向上一个片段移 动确定信息的起始点, 并将信息的起始点作为片段的起始点和上一个片段的结束点。
本方案避免一个完整的信息被拆分到两个片段中去,进一步对片段的起始点和结束点 进行精确的划分。
在一个实施例中, 方法还包括:
判断片段的起始点处的信息是否完整, 如果不完整, 在片段的起始点向片段的结束点 移动确定信息的结束点, 并将信息的结束点作为片段的起始点和上一个片段的结束点。
本方案避免一个完整的信息被拆分到两个片段中去,进一步对片段的起始点和结束点 进行精确的划分。
在一个实施例中, 方法还包括:
判断片段的结束点处的信息是否完整, 如果不完整, 在片段的结束点向下一个片段移 动确定信息的结束点, 并将信息的结束点作为片段的结束点和下一个片段的起始点。
本方案避免一个完整的信息被拆分到两个片段中去,进一步对片段的起始点和结束点 进行精确的划分。 在一个实施例中, 方法还包括:
判断片段的结束点处的信息是否完整, 如果不完整, 在片段的结束点向片段的起始点 移动确定信息的起始点, 并将信息的起始点作为片段的结束点和下一个片段的起始点。
本方案避免一个完整的信息被拆分到两个片段中去,进一步对片段的起始点和结束点 进行精确的划分。
在一个实施例中, 方法还包括:
判断布局排版数据中包含的排版原子数据在预设时间内是否被使用, 如果没有被使 用, 删除排版原子数据。
本方案实现节省移动终端的内存。
在一个实施例中, 方法还包括:
判断布局排版数据中包含的排版原子数据占用的内存是否大于预设值,如果大于预设 值, 删除排版原子数据。
本方案实现节省移动终端的内存。
根据本公开实施例的第二方面, 本公开实施例还提出了一种终端, 该终端包括: 获取模块, 用于获取电子书文档;
分段模块, 用于按照预设分段方式对电子书文档的内容分段, 生成多个片段; 组成模块, 用于将多个片段组成一个有序的片段组;
选择模块, 用于在片段组中选择一个片段, 将片段作为当前片段;
解析模块, 用于解析当前片段的内容, 生成布局排版数据;
生成模块, 用于根据布局排版数据生成页面图像。
在一个实施例中, 终端还包括:
记录模块, 用于记录当前片段在片段组中的位置信息;
第一判断模块, 用于判断布局排版数据的数据量是否低于预设值;
第一执行模块, 用于在数据量低于预设值时, 根据位置信息选择当前片段的下一个片 段, 并将下一个片段作为当前片段, 返回执行解析模块。
在一个实施例中, 分段模块包括:
分段值确定单元, 用于确定分段值;
拆分单元, 用于将电子书文档的内容拆分成多个与分段值大小相同的片段。
在一个实施例中, 终端还包括:
第二判断模块, 用于判断片段的起始点处的信息是否完整;
第二执行模块, 用于在片段的起始点处的信息不完整时, 在片段的起始点向前移动确 定标签的起始点, 并将标签的起始点作为片段的起始点。
在一个实施例中, 终端还包括:
第三判断模块, 用于判断片段的起始点处的信息是否完整;
第三执行模块, 用于在判断片段的起始点处的信息不完整时, 在片段的起始点向片段 的结束点移动确定信息的结束点,并将信息的结束点作为片段的起始点和上一个片段的结 束点。
在一个实施例中, 终端还包括:
第四判断模块, 用于判断片段的结束点处的信息是否完整;
第四执行模块, 用于在片段的结束点处的信息不完整时, 在片段的结束点向下一个片 段移动确定信息的结束点, 并将信息的结束点作为片段的结束点和下一个片段的起始点。
在一个实施例中, 终端还包括:
第五判断模块, 用于判断片段的结束点处的信息是否完整;
第五执行模块, 用于在片段的结束点处的信息不完整时, 在片段的结束点向片段的起 始点移动确定信息的起始点, 并将信息的起始点作为片段的结束点和下一个片段的起始 点。
在一个实施例中, 终端还包括:
第六判断模块,用于判断布局排版数据中包含的排版原子数据在预设时间内是否被使 用;
第六执行模块,用于在排版原子数据在预设时间内没有被使用时,删除排版原子数据。 在一个实施例中, 终端还包括:
第七判断模块,用于判断布局排版数据中包含的排版原子数据占用的内存是否大于预 设值;
第七执行模块, 用于在排版原子数据占用的内存大于预设值时, 删除排版原子数据。 根据本公开实施例的第三方面, 本公开实施例还提出了一种电子设备, 该电子设备包 括有存储器,以及一个或者一个以上的程序,其中一个或者一个以上程序存储于存储器中, 且经配置以由一个或者一个以上处理器执行所述一个或者一个以上程序包含用于进行以 下操作的指令:
获取电子书文档;
按照预设分段方式对所述电子书文档的内容分段, 生成多个片段;
将所述多个片段组成一个有序的片段组;
在所述片段组中选择一个片段, 将所述片段作为当前片段;
解析所述当前片段的内容, 生成布局排版数据;
根据所述布局排版数据生成页面图像。
本公开实施例的其它特征和优点将在随后的说明书中阐述, 并且, 部分地从说明书中 变得显而易见, 或者通过实施本公开实施例而了解。本公开实施例的目的和其他优点可通 过在所写的说明书、 权利要求书、 以及附图中所特别指出的结构来实现和获得。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性的,并不能限制本公开。 下面通过附图和实施例, 对本公开实施例的技术方案做进一步的详细描述。 附图说明
此处的附图被并入说明书中并构成本说明书的一部分, 示出了符合本发明的实施例, 并与说明书一起用于解释本发明的原理。
图 1为本公开实施例提供的一种电子书文档的处理方法的示例性流程图;
图 2为本公开实施例提供的另一种电子书文档的处理方法的示例性流程图; 图 3为本公开实施例提供的又一种电子书文档的处理方法的示例性流程图; 图 4为本公开实施例提供的一种 HTML文档的处理方法的示例性流程框架图; 图 5为本公开实施例提供的一种终端的模块示意图;
图 6为本公开实施例提供的分段模块的模块示意图;
图 7为本公开实施例提供的另一种终端的模块示意图;
图 8为本公开实施例提供的又一种终端的模块示意图;
图 9为本公开实施例提供的又一种终端的模块示意图;
图 10为本公开实施例提供的电子设备的结构示意图。 具体实施方式
目前, 移动终端读取电子书文档的过程主要包括以下步骤: 解析步骤: 移动终端将电 子书文档解析成布局排版数据; 全分页步骤: 对上述布局排版数据全分页处理得到页面框 架; 生成页面对象步骤: 利用页面框架和布局排版数据生成对应的页面对象; 生成页面图 像步骤: 对页面对象进行渲染, 生成页面图像。 经过上述步骤的处理, 通过移动终端, 用 户可以看到电子书的页面图像。
为了更好的实现移动终端读取电子书文档,本公开实施例提供了一种电子书文档的处 理方法、 终端及电子设备, 以使移动终端读取电子书的时间短, 内存占用低。 由于上述电 子书文档的处理方法、 终端及电子设备的具体实现存在多种方式, 下面通过具体实施例进 行详细说明:
请参见图 1所示, 图 1所示的为一种电子书文档的处理方法, 该方法可以包括以下步 骤:
步骤 101、 获取电子书文档;
其中,在用户通过移动终端打开电子书以后,移动终端会读取电子书文档到内存储器, 此时, 移动终端已经获取到电子书文档。
在本公开实施例中, 电子书文档可以为流式电子书文档, 这里所述的流式电子书文档 是指所描述的文字, 图片等信息没有固定的版面位置, 当排版参数(如版面宽度, 字号和 行间距等)发生变化时需要对版面重新布局排版以适应新的排版参数的电子书文档。 该流 式电子书文档包括如 HTML文档在内的文档, HTML文档是由标签组成的, 所以在移动 终端获取到电子书的 HTML文档的同时, 也获取到了构成 HTML文档的标签。
步骤 102、 按照预设分段方式对电子书文档的内容分段, 生成多个片段; 其中, 预设分段方式可以有多种实现形式, 例如: 确定分段值, 将电子书文档的内容 拆分成多个与分段值大小相同的片段。具体的, 分段值代表电子书文档拆分后每个片段的 大小, 所以分段值可以是预先设定的, 也可以是根据移动终端用户自行设定的, 还可以是 根据实验计算出来的。
步骤 103、 将多个片段组成一个有序的片段组;
其中, 经过上述步骤对电子书文档分段后, 电子书文档被分成了很多个片段。 为了保 证按照原有电子书文档内容的顺序处理数据, 所以需要建立片段与片段之间的联系, 将这 些片段按照原有的顺序组成一个片段组。
步骤 104、 在片段组中选择一个片段, 将片段作为当前片段;
其中, 在片段组中可以根据用户的需求选择其中的任何一个片段, 以便于后续步骤对 选择的片段进行处理。
步骤 105、 解析当前片段内的内容, 生成布局排版数据;
其中, 某个片段的内容仅为完整的电子书文档内容的一部分, 需要解析的数据相对于 完整的电子书文档要少很多的内容, 所以解析的速度明显变快, 而且, 数据量减少, 从而 使占用的内存也大大的降低。
另外, 在布局排版数据生成以后, 为了节省移动终端的内存, 还可以包括如下步骤: 判断布局排版数据中包含的排版原子数据在预设时间内是否被使用, 如果没有被使用, 删 除排版原子数据。 其中, 如果该布局排版数据很久没有被使用过, 那么为了节省该布局排 版数据所占用的内存, 可以先删除该排版原子数据, 如果以后需要, 那么可以从新生成该 排版原子数据。
另外, 在布局排版数据生成以后, 为了节省移动终端的内存, 还可以包括如下步骤: 判断布局排版数据中包含的排版原子数据占用的内存是否大于预设值, 如果大于预设值, 删除排版原子数据。 其中, 如果布局排版数据过多, 会导致后续步骤处理速度过慢, 可以 删除全部的布局排版数据, 也可以删除预定数量的布局排版数据。
步骤 106、 根据布局排版数据生成页面图像。
其中,布局排版数据的数据量是由一个片段生成的,布局排版数据的数据量相对较少, 在生成页面图像的过程中, 无论是时间的占用还是内存的占用均减少很多。
在步骤 106 中, 具体还可以包括以下三个步骤, 1、 对布局排版数据进行分页处理, 生成页面布局框架; 2、 根据页面布局框架和布局排版数据生成页面对象; 3、 对页面对象 进行渲染生成页面图像。
在图 1所示的实施例中, 将电子书文档分成多个片段, 每次对电子书操作时, 仅解析 电子书文档中的一个片段, 并利用该片段生成的布局排版数据生成页面图像, 所以每次在 操作过程中需要处理的数据量仅为一个片段的数据量, 所以在移动终端读取电子书文档 时, 处理数据的效率就得到提升, 进而缩短了读取电子书的时间; 而且, 分段操作实现对 电子书文档分批次处理, 移动终端在处理电子书文档的一个片段时, 无论是解析操作还是 后续的生成页面图像的操作, 需处理的数据量都相对较小, 所以降低电子书对移动终端内 存的占用。 请参见图 2所示,图 2所示的为另一种电子书文档的处理方法,该方法包括以下步骤: 步骤 201、 获取电子书文档;
其中,在用户通过移动终端打开电子书以后,移动终端会读取电子书文档到内存储器, 此时, 移动终端已经获取到电子书文档。 在本公开实施例中, 电子书文档可以为 HTML 文档, HTML文档是由标签组成的, 所以在移动终端获取到电子书的 HTML文档的同时, 也获取到了构成 HTML文档的标签。
步骤 202、 按照预设分段方式对电子书文档的内容分段, 生成多个片段;
其中, 预设分段方式可以有多种实现形式, 下面介绍一种分段方法, 该方法为: 确定 分段值, 将电子书文档的内容拆分成多个与分段值大小相同的片段。 具体的, 分段值代表 电子书文档拆分后每个片段的大小, 所以分段值可以是预先设定的, 也可以是根据移动终 端用户自行设定的, 还可以是根据实验计算出来的。
步骤 203、 将多个片段组成一个有序的片段组;
其中, 经过上个步骤对电子书文档分段后, 电子书文档被分成了很多个片段。 为了保 证按照原有电子书文档内容的顺序处理数据, 所以需要建立片段与片段之间的联系, 将这 些片段按照原有的顺序组成一个片段组。
步骤 204、 在片段组中选择一个片段, 将片段作为当前片段;
其中, 在片段组中可以根据用户的需求选择其中的任何一个片段, 以便于后续步骤对 选择的片段进行处理。
步骤 205、 解析当前片段内的内容, 生成布局排版数据;
其中, 某个片段的内容仅为完整的电子书文档内容的一部分, 需要解析的数据相对于 完整的电子书文档要少很多的内容, 所以解析的速度明显变快, 而且, 数据量减少, 从而 使占用的内存也大大的降低。
步骤 206、 记录当前片段在片段组中的位置信息;
其中, 位置信息的具体实现存在多种方式, 例如, 可以采用字节来记录位置信息, 也 可以采用标识来记录位置信息, 等等。 对于电子书文档为 HTML文档而言, 可以采用字 节偏移量的方式来记录片段的位置信息, 字节偏移量的单位为字节。对于位置信息在此不 做具体的限制, 只要能够记录片段在片段组中的位置, 都在本公开方案的保护范围内。
步骤 207、 判断布局排版数据的数据量是否低于预设值, 若是, 则执行步骤 208, 否 则, 执行步骤 209;
其中, 布局排版数据是由当前片段解析生成的, 每个片段解析生成的布局排版数据的 数据量是一定的, 如果已经生成的布局排版数据的数据量足够生成页面图像, 那么就利用 该布局排版数据生成页面图像;如果已经生成的布局排版数据的数据量不足以生成页面图 像, 那么就需要再解析下一个片段以生成布局排版数据, 再将之前剩余的布局排版数据和 下一个片段生成的布局排版数据合并在一起, 生成页面图像。
步骤 208、根据位置信息选择当前片段的下一个片段,并将下一个片段作为当前片段, 返回步骤 205 ;
其中, 在步骤 206中已经记录下当前片段的位置信息, 所以可通过当前片段的位置信 息来确定下一个片段的位置信息, 从而可以选择到下一个片段中的内容。然后将下一个片 段作为当前待处理的片段, 返回步骤 205。
步骤 209、 根据布局排版数据生成页面图像。
其中,布局排版数据的数据量是由一个片段生成的,布局排版数据的数据量相对较少, 在生成页面图像的过程中, 无论是时间的占用还是内存的占用均减少很多。
在图 2所示的实施例中, 与图 1所示的实施例的不同之处在于, 判断已经生成的布局 排版数据的数据量是否足以生成页面图像, 如果是, 那么根据已有的布局排版数据生成页 面图像; 否则, 解析下一个片段, 根据已有的布局排版数据和下一个片段生成的布局排版 数据合并生成页面图像。 基于上述实施例, 在生成多个片段的步骤以后, 还可以进一步对片段的起始点和结束 点进行精确的划分。 例如, 电子书文档具体为 HTML文档, 片段的起始点是片段起始的 位置, 片段的结束点是片段结束的位置, HTML文档主要由标签构成, 标签是由两个尖括 号组成, 即标签由左尖括号 "<"和右尖括号 ">"组成, 两个尖括号之间为标签的内容。 按照 HTML语法规定, 标签必须包括左尖括号 "<"和右尖括号 ">" , 这样标签才算完 整, 因此, 分段后, 不允许存在某一个片段仅含有左尖括号或仅含有右尖括号的不对应情 况,如果存在,说明一个标签被分割到两个片段中,这样后面是无法对该标签进行解析的。 为了进一步对片段起始点和结束点进行划分, 请参见图 3所示的实施例。
请参见图 3所示,图 3所示的为又一种电子书文档的处理方法,该方法包括以下步骤: 步骤 301、 获取电子书文档;
其中,在用户通过移动终端打开电子书以后,移动终端会读取电子书文档到内存储器, 此时, 移动终端已经获取到电子书文档。 在本公开实施例中, 电子书文档可以为 HTML 文档, HTML文档是由标签组成的, 所以在移动终端获取到电子书的 HTML文档的同时, 也获取到了构成 HTML文档的标签。
步骤 302、 按照预设分段方式对电子书文档的内容分段, 生成多个片段;
其中, 预设分段方式可以有多种实现形式, 下面介绍一种分段方法, 该方法为: 确定 分段值, 将电子书文档的内容拆分成多个与分段值大小相同的片段。 具体的, 分段值代表 电子书文档拆分后每个片段的大小, 所以分段值可以是预先设定的, 也可以是根据移动终 端用户自行设定的, 还可以是根据实验计算出来的。
步骤 303、 判断片段的起始点处的信息是否完整, 若是, 则执行步骤 305; 否则, 执 行步骤 304;
其中, 在分段以后, 为了避免一个完整的信息被拆分到两个片段中去, 所以需要判断 片段的起始点处的信息是否完整, 如果不完整, 那么需要对片段的起始点和结束点进行一 定的微调, 以保证每个片段内的信息都是完整的。
步骤 304、 在片段的起始点向上一个片段移动确定信息的起始点, 并将信息的起始点 作为片段的起始点和上一个片段的结束点;
其中, 如果片段的起始点处的信息不完整, 那么说明该信息的前半部分被分到上一个 片段中, 而该信息的后半部分留在了该片段中的起始点处。 为了保证该信息的完整, 其实 存在多种实现方式, 其中一种实现方式就是步骤 304的调整方式, 为删除上一个片段中该 信息的前半部分, 并将该信息的前半部分补充到该片段的起始点处, 该片段的起始点和上 一个片段的结束点是同一个点。第二种实现方式具体为, 在片段的起始点向该片段的结束 点移动确定该信息的结束点,并将该信息的结束点作为该片段的起始点和上一个片段的结 束点。第二种方式为删除该片段中该信息的后半部分, 并将该信息的后半部分补充到上一 个片段中的结束点处。上述两种方式均可实现对各个片段的起始点和结束点处的微调, 以 实现在每个片段中所有信息均是完整的。
当然, 为了保证片段中信息的完整, 不仅可以采用步骤 304的方式, 还可以确定其他 信息的起始点作为该片段的起始点和上一个片段的结束点,虽然此种方式片段的起始点或 结束点偏移量过多, 但是同样可以解决信息完整的问题, 所以本公开提供的实施例并不局 限于步骤 304的方式, 只要能够实现信息完整的目的, 其他与步骤 304原理相同的方法均 可, 在此不作具体的限制。
步骤 305、 将多个片段组成一个有序的片段组;
其中, 经过上个步骤对电子书文档分段后, 电子书文档被分成了很多个片段。 为了保 证按照原有电子书文档内容的顺序处理数据, 所以需要建立片段与片段之间的联系, 将这 些片段按照原有的顺序组成一个片段组。
步骤 306、 在片段组中选择一个片段, 将片段作为当前片段;
其中, 在片段组中可以根据用户的需求选择其中的任何一个片段, 以便于后续步骤对 选择的片段进行处理。
步骤 307、 解析当前片段内的内容, 生成布局排版数据;
其中, 某个片段的内容仅为完整的电子书文档内容的一部分, 需要解析的数据相对于 完整的电子书文档要少很多的内容, 所以解析的速度明显变快, 而且, 数据量减少, 从而 使占用的内存也大大的降低。
步骤 308、 根据布局排版数据生成页面图像。
其中,布局排版数据的数据量是由一个片段生成的,布局排版数据的数据量相对较少, 在生成页面图像的过程中, 无论是时间的占用还是内存的占用均减少很多。
在图 3所示的实施例中, 为了避免完整的信息被拆分, 还可以判断结束点处的信息是 否完整, 具体有两种实现方法, 第一种方法为, 判断片段的结束点处的信息是否完整, 如 果不完整, 在片段的结束点向下一个片段移动确定信息的结束点, 并将信息的结束点作为 片段的结束点和下一个片段的起始点。第二种方法为, 判断片段的结束点处的信息是否完 整, 如果不完整, 在片段的结束点向片段的起始点移动确定信息的起始点, 并将信息的起 始点作为片段的结束点和下一个片段的起始点。其原理与步骤 303和步骤 304相似, 具体 解释请参见步骤 303和步骤 304, 在此不再赘述。
在图 3所示的实施例中, 与图 1所示的实施例的不同之处在于, 为了避免一个完整的 信息被拆分到两个片段中去, 进一步对片段的起始点和结束点进行精确的划分。 请参见图 4所示, 图 4所示的为本公开提供的具体应用例。本应用例介绍的是电子书 文档为 HTML文档 1时, 具体的处理方法, 该 HTML文档 1的处理方法包括如下步骤: 步骤一、 获取 HTML文档 1 ;
步骤二、 按照预设分段方式对 HTML文档 1的内容分段, 生成多个 HTML片段; 在步骤二中, HTML文档 1解析过程最大的特点是对 HTML片段解析, 解析一个小 的 HTML片段耗时即在可接受范围内; 如果每次解析过程仅处理一个小的 HTML片段, 即可缩短单次解析时间; 因此这里对 HTML进行分段解析。
首先需要确定片段的大小值 m (单位: B (字节) ) ; 再根据该 m值将大小为 n (单 位: B (字节) ) 的 HTML文档 1等分成 n/m份, 得到 (n/m-1)个断点; 由于非从头开始的 HTML文档 1解析必须保证 HTML文档 1内标签的完整性,以防止半个标签出现在 HTML 片段中, 因此需要再依次检査上述 (n/m-l)中的每个断点是否满足不在 HTML标签内部。 按照 HTML语法规定: 标签必须被左尖括号' < '和右尖括号' >'包住, 长度一般在 1024B 以 内。 根据上述限制可对上述 (n/m-1)中的每个点向前或向后査询一定的字节以满足不在 HTML标签内部的要求。
具体的, 确定片段的大小值 m 有很多种方式, 例如, 人为设定的方式或计算的方式 等。 下面重点介绍确定片段的大小值 m的计算方式:
1、 理论分析
由于 HTML文档 1解析时间由 HTML节点数直接决定, 随着 HTML文档 1的逐步增 大, HTML节点数也随之曾多, 节点的深度遍历耗时也就逐步增大; 由此可知, HTML文 档 1解析时间随 HTML文档 1大小的变化曲线应为正向递增曲线。 一般情况下当 HTML 文档 1较小时, HTML节点的数量较少, 对于现代计算机而言, HTML文档 1解析能在很 短的时间内完成深度遍历 (如面对 10K和 50K的 HTML文档 1, 其解析时间不会有太大 差别) ; 而当 HTML文档 1 的规模到达一定程度时, HTML节点数会出现非线性的急剧 增长, 这样就导致 HTML文档 1解析时间的非线性增长。 综上分析可知, HTML文档 1 解析时间随 HTML 文档 1 的变化必是先缓后急的过程, 因此曲线必存在一个拐点, 当 HTML文档 1大小到达该拐点后, HTML文档 1解析的效率急剧下降。 2、 理论验证
采用一些常用的性能分析手段对已准备好的几组 HTML物理空间进行解析性能测试, 记录每组测试结果, 通过可视化工具绘制出 HTML 文档解析时间随 HTML文档 1大小的 变化曲线, 验证上述理论分析。
3、 寻找 HTML文档 1解析性能拐点
通过理论验证后的 HTML文档 1解析时间随 HTML文档 1大小的变化曲线, 可确定 当 HTML文档 1大小在某一数值 M以下时 HTML 文档解析时间基本无较大变化, 而当 HTML物理空间大小超过某一数值 N时, HTML 文档解析时间就会急剧增加; 因此, 可 将 HTML文档 1解析性能拐点对应的 HTML文档 1的大小确定在区间 [M, N]内。 再在区 间 [M, N)内进行分析; 此时有两个需要考虑的准则: a、 HTML文档 1分段不宜过多, 否 则会增加段管理上复杂性, 这就使得片段的大小值 m不能过小; b、 单个 HTML片段的解 析时间不宜过长, 否则会出现用户等待时间过长等缺点, 这就使得段的大小值 m 不能过 大; 根据以上两个准则可先去掉区间 [M, N]内的最小值 M和最大值 N得到新的区间, 再 按照上述两个准则去掉新的区间的最大值和最小值, 依次循环下去, 若最后剩余一个值, 则此值即作为性能拐点对应的 HTML文档 1的大小, 也就是片段的大小值 m, 若最后剩 余两个数, 则取两个数的中间值作为性能拐点对应的 HTML文档 1 的大小, 也就是片段 的大小值!^
步骤三、 将多个片段组成一个有序的片段组;
解析过程最大的缺点是不可回溯, 即解析必须按照 HTML文档 1 的物理顺序执行。 如果从 HTML文档 1的非起始点开始解析, 就会存在 HTML文档 1上下文关系不全的问 题, 导致标签节点丢失等问题。 因此这里采用 HTML文档 1的解析状态区间 2来克服该 问题。
1、 新建 HTML文档 1的解析器, 此时, 该解析器的状态未经过任何 HTML文档 1解 析, 称为最初始解析状态; 最初始解析状态与具体的 HTML文档 1无关, 任意新创建的 HTML文档 1的解析器所处的状态都可为最初始状态。
2、 由于 HTML文档 1的解析状态会记录所在 HTML文档 1的字节偏移量,而 HTML 文档 1的分段结果也是采用字节偏移量来描述的,所以按照这种字节偏移量的对应关系即 可完成 HTML片段到解析区间的一一对应。 因此, 第一个 HTML片段对应第一个解析区 间, 其解析开始状态 (所处字节偏移量为第一个片段的开始)取最初始解析状态, 其解析 结束状态为完成第一个 HTML片段的解析状态 (所处的字节偏移量为第一段的结束) 。
3、第一个 HTML片段以后的每个 HTML片段对应的解析区间, 其解析开始状态从上 一个解析区间的解析结束状态继承, 其解析结束状态为完成该 HTML片段解析的状态。 这里的解析状态继承是解析状态的克隆过程, 即将一个解析状态所处的 HTML文档 1 的 字节偏移量 (用于记录解析状态在 HTML文档 1当前的位置) , 父 HTML标签节点的栈 信息 (用于记录解析状态在 HTML文档 1走过的路劲) , 解析时的上下文关系 (用于记 录解析状态落在 HTML文档 1 的文本节点内的前后文关系) 等原封不动的拷贝到另外一 个解析状态中。
按照这种方案每个 HTML片段在参与 HTML文档 1解析的过程中都从上个 HTML片 段继承解析状态, 也就继承下来了上个 HTML片段的上下文关系, 即克服了因上下关系 不全带来的问题。
步骤四、 在片段组中选择一个片段, 将片段作为当前片段;
步骤五、 解析当前片段的内容, 生成布局排版数据 3 ;
布局驱动过程:
1、 启动布局引擎, 当发现当前 HTML片段对应的排版原子数据 31为空时, 即根据 当前布局的点所对应 HTML文档 1的偏移量寻找所在的 HTML片段, 唤醒相应的 HTML 文档 1的解析区间, 对相应的 HTML片段进行翻译, 生成相应的标签节点数据 32和排版 原子数据 31。 对于标签节点数据 32而言, 由于排版原子数据 31需要依赖标签节点数据 32的样式信息, 因此随着排版原子数据 31的解析, 标签节点数据 32将随之解析。标签节 点数据 32只需保存一份, 若发现该 HTML片段所对应的标签节点数据 32已存在, 则解 析过程中不再添加相应的标签节点数据 32。
2、 进行布局排版完成分页或页面对象的生成。 同时监控当前存储资源的使用情况, 若超过规定的阈值 (如 android平台的应用内存一般被限制为 24MB以内, 而 iOS平台的 应用一般被限制在 20M以内),则删除暂时用不到的 HTML片段对应的排版原子数据 31。
步骤六、 根据布局排版数据 3生成页面图像。
HTML全分页过程:
全分页过程从 HTML物理空间的最开始进行处理, 从前到后依次对每一个 HTML片 段进行布局驱动过程的调用; 对于每一个 HTML片段而言, 当分页过程到达该片段的最 后一页时, 则会出现半页的情形, 为了保证当前片段的最后一页与下一片段连续则此时立 即解析下一片段的 HTML数据, 再从当前片段的最后一页的开始点进行分页, 若发现排 版原子不足则使用下一片段的排版原子, 以避免片段之间的断裂问题。
页面对象生成过程:
1、 通过 HTML文档 1的字节偏移量(以下称为 "请求点")确定其对应的 HTML页 空间 4中的页码 (HTML页的起始点和结束点会记录其在物理 HTML中字节偏移量) , 此时便有两种情形: 获取到页码和未获取页码。若获取到页码则说明全分页过程已到达请 求点,则使用页码对应的全分页结果;若未获取到页码则说明全分页过程尚未到达请求点, 此时启动临时解析过程获取请求点对应的页结果。
2、 根据上述获取的页结果, 调用布局驱动过程, 完成页面对象的生成。
临时解析过程:
1、 根据请求点判断所处的 HTML片段(以下称为 "请求片段") ; 新建 HTML文档 1 的解析状态区间 2, 取其解析起始状态为最初始解析状态, 取其解析结束状态为使用解 析起始状态完成请求片段的解析后的状态。
2、 调用布局驱动过程, 强制使用上述步骤中的解析状态区间 2对请求片段进行分页, 生成临时解析页空间 4。
3、 根据请求点确定其在临时解析页空间 4中的页码, 根据该页码获取临时解析页结 果。
注:当全分页过程尚未到达所需页所在的区间时,采用临时解析过程获取所需要的页, 此时虽然未使用完整的 HTML上下文关系, 但当全分页过程到达所需页后即转向全分页 结果, 以修正从非起始点解析 HTML存在的问题。 当 HTML全分页过程尚未完成时不可 通过页码进行跳转, 原因在于页码可能超过全分页过程已分页的数目, 此时可跟过 HTML 物理空间大小的百分比跳转, 当全分页完成后即可通过页码进行跳转。
本公开提供的方案在使得大尺寸 HTML文档 1在全分页过程中消耗更小的存储空间, 并降低 HTML文档 1 中靠前部分全分页结果的获取时间; 在全分页到达之前采用临时解 析使得 HTML跳转的性能提高, 并在全分页到达后修正临时解析中可能存在的问题。
如此对于用户体验而言则带来了如下有益的效果:
1、 将 HTML 文档的解析时间与全分页时间分摊到每个 HTML片段中, 明显缩短了 用户在首次打开以大尺寸 HTML文档 1为主的电子书时, 第一页的等待时间;
2、 全分页到请求点后采用全分页结果, 保证了用户在阅读以大尺寸 HTML文档 1为 主的电子书时未破坏书籍样式结构, 以最尊重原书的形式将内容呈现给用户;
3、 消耗更小的存储空间, 明显降低了用户在移动设备等存储资源较少的设备上打开 以大尺寸 HTML文档 1为主的电子书的死机概率;
4、 消耗更小的存储空间, 使得用户在阅读以大尺寸 HTML文档 1为主的电子书时阅 读软件的运行更加流畅, 提高用户操作的流畅性;
5、 采用临时解析过程, 明显缩短了用户在阅读以大尺寸 HTML文档 1为主的电子书 时同步阅读进度的等待时间;
6、 采用临时解析过程, 明显缩短了用户在阅读以大尺寸 HTML文档 1为主的电子书 时使目录跳转的等待时间;
7、 采用临时解析过程, 明显缩短了用户在阅读以大尺寸 HTML文档 1为主的电子书 时快进快退的等待时间。 请参见图 5所示, 图 5所示的为一种终端, 该终端包括获取模块 11, 用于获取电子书 文档; 分段模块 12, 用于按照预设分段方式对电子书文档的内容分段, 生成多个片段; 组 成模块 13, 用于将多个片段组成一个有序的片段组; 选择模块 14, 用于在片段组中选择 一个片段, 将片段作为当前片段; 解析模块 15, 用于解析当前片段的内容, 生成布局排版 数据; 生成模块 16, 用于根据布局排版数据生成页面图像。 其中, 请参见图 6所示, 图 6 所示的为分段模块 12的具体结构,分段模块 12包括分段值确定单元 121和拆分单元 122, 分段值确定单元 121, 用于确定分段值; 拆分单元 122, 用于将电子书文档的内容拆分成 多个与分段值大小相同的片段。
在图 5所示的实施例中, 利用分段模块 12将电子书文档分成多个片段, 每次对电子 书操作时, 解析模块 15仅解析电子书文档中的一个片段, 生成模块 16利用该片段生成的 布局排版数据生成页面图像,所以每次在操作过程中需要处理的数据量仅为一个片段的数 据量, 所以在移动终端读取电子书文档时, 处理数据的效率就得到提升, 进而缩短了读取 电子书的时间; 而且, 分段操作实现对电子书文档分批次处理, 移动终端在处理电子书文 档的一个片段时, 无论是解析操作还是后续的生成页面图像的操作, 需处理的数据量都相 对较小, 所以降低电子书对移动终端内存的占用。 在图 6所示的实施例中,利用分段值确定单元 121得到的分段值代表电子书文档拆分 后每个片段的大小, 所以分段值可以是预先设定的, 也可以是根据移动终端用户自行设定 的, 还可以是根据实验计算出来的。然后再利用拆分单元 122根据得到的分段值对电子书 文档进行拆分。 请参见图 7所示, 图 7所示的为另一种终端, 该终端包括获取模块 21, 用于获取电子 书文档; 分段模块 22, 用于按照预设分段方式对电子书文档的内容分段, 生成多个片段; 组成模块 23, 用于将多个片段组成一个有序的片段组; 选择模块 24, 用于在片段组中选 择一个片段, 将片段作为当前片段; 解析模块 25, 用于解析当前片段的内容, 生成布局排 版数据; 记录模块 26, 用于记录当前片段在片段组中的位置信息; 第一判断模块 27, 用 于判断布局排版数据的数据量是否低于预设值;第一执行模块 28,用于在数据量低于预设 值时, 根据位置信息选择当前片段的下一个片段, 并将下一个片段作为当前片段, 返回执 行解析模块 25 ; 生成模块 29, 用于根据布局排版数据生成页面图像。
在图 7所示的实施例中,与图 6所示的实施例的不同之处在于,通过第一判断模块 27 判断已经生成的布局排版数据的数据量是否足以生成页面图像, 如果是, 那么根据已有的 布局排版数据生成页面图像; 否则, 解析下一个片段, 根据已有的布局排版数据和下一个 片段生成的布局排版数据合并生成页面图像。 请参见图 8所示, 图 8所示的为又一种终端, 该终端包括获取模块 31, 用于获取电子 书文档; 分段模块 32, 用于按照预设分段方式对电子书文档的内容分段, 生成多个片段; 第二判断模块 33, 用于判断片段的起始点处的信息是否完整; 第二执行模块 34, 用于在 片段的起始点处的信息不完整时, 在片段的起始点向前移动确定标签的起始点, 并将标签 的起始点作为片段的起始点; 组成模块 35, 用于将多个片段组成一个有序的片段组; 选择 模块 36, 用于在片段组中选择一个片段, 将片段作为当前片段; 解析模块 37, 用于解析 当前片段的内容,生成布局排版数据;生成模块 38,用于根据布局排版数据生成页面图像。 在图 8所示的实施例中, 与图 5所示的实施例的不同之处在于, 为了避免一个完整的 信息被拆分到两个片段中去, 进一步对片段的起始点和结束点进行精确的划分。
在图 8所示的实施例中, 还有三种实现方式可以等同的替代上述第二判断模块 33和 第二执行模块 34的功能, 下面简要介绍这三种实现方式: 第一种替换方式, 利用第三判 断模块替换第二判断模块 33, 利用第三执行模块替换第二执行模块 34, 具体的, 第三判 断模块, 用于判断片段的起始点处的信息是否完整; 第三执行模块, 用于在判断片段的起 始点处的信息不完整时, 在片段的起始点向片段的结束点移动确定信息的结束点, 并将信 息的结束点作为片段的起始点和上一个片段的结束点。第二种替换方式, 利用第四判断模 块替换第二判断模块 33, 利用第四执行模块替换第二执行模块 34, 具体的, 第四判断模 块, 用于判断片段的结束点处的信息是否完整; 第四执行模块, 用于在片段的结束点处的 信息不完整时, 在片段的结束点向下一个片段移动确定信息的结束点, 并将信息的结束点 作为片段的结束点和下一个片段的起始点。第三种替换方式, 利用第五判断模块替换第二 判断模块 33, 利用第五执行模块替换第二执行模块 34, 具体的, 第五判断模块, 用于判 断片段的结束点处的信息是否完整; 第五执行模块, 用于在片段的结束点处的信息不完整 时, 在片段的结束点向片段的起始点移动确定信息的起始点, 并将信息的起始点作为片段 的结束点和下一个片段的起始点。
在图 8所示的实施例中,为了保证片段中信息的完整,不仅可以采用第二判断模块 33 和第二执行模块 34的功能, 还可以确定其他信息的起始点作为该片段的起始点和上一个 片段的结束点, 虽然此种方式片段的起始点或结束点偏移量过多, 但是同样可以解决信息 完整的问题, 所以本公开提供的实施例并不局限于图 8所示的方式, 只要能够实现信息完 整的目的, 其他与第二判断模块 33和第二执行模块 34的功能原理相同的模块均可, 在此 不作具体的限制。 请参见图 9所示, 图 9所示的为又一种终端, 该终端包括获取模块 41, 用于获取电子 书文档; 分段模块 42, 用于按照预设分段方式对电子书文档的内容分段, 生成多个片段; 组成模块 43, 用于将多个片段组成一个有序的片段组; 选择模块 44, 用于在片段组中选 择一个片段, 将片段作为当前片段; 解析模块 45, 用于解析当前片段的内容, 生成布局排 版数据;第六判断模块 46,用于判断布局排版数据中包含的排版原子数据在预设时间内是 否被使用; 第六执行模块 47, 用于在排版原子数据在预设时间内没有被使用时, 删除排版 原子数据。 生成模块 48, 用于根据布局排版数据生成页面图像。
在图 9所示的实施例中, 在解析模块 45生成布局排版数据以后, 为了节省移动终端 的内存, 可以利用第六判断模块 46和第六执行模块 47对其实现。 其中, 如果该布局排版 数据很久没有被使用过, 那么为了节省该布局排版数据所占用的内存, 可以先删除该排版 原子数据, 如果以后需要, 那么可以从新生成该排版原子数据。
在图 9所示的实施例中, 在解析模块 45生成布局排版数据以后, 为了节省移动终端 的内存,可以利用第七判断模块替换第六判断模块 46,利用第七执行模块替换第六执行模 块 47对其实现。 具体的, 第七判断模块, 用于判断布局排版数据中包含的排版原子数据 占用的内存是否大于预设值; 第七执行模块, 用于在排版原子数据占用的内存大于预设值 时, 删除排版原子数据。 其中, 如果布局排版数据过多, 会导致后续步骤处理速度过慢, 可以删除全部的布局排版数据, 也可以删除预定数量的布局排版数据, 从而实现了节省移 动终端内存的目的。 如图 10所示, 本公开实施例还提供了一种电子设备, 该电子设备用于实施上述实施 例中提供的电子书文档的处理方法, 具体来讲:
电子设备 1300可以包括 RF (Radio Frequency, 射频) 电路 1310、 包括有一个或一个 以上计算机可读存储介质的存储器 1320、 输入单元 1330、 显示单元 1340、 传感器 1350、 音频电路 1360、 短距离无线传输模块 1370、 包括有一个或者一个以上处理核心的处理器 1380、 以及电源 1390等部件。 本领域技术人员可以理解, 图 13中示出的电子设备结构并 不构成对电子设备的限定, 可以包括比图示更多或更少的部件, 或者组合某些部件, 或者 不同的部件布置。 其中:
RF电路 1310可用于收发信息或通话过程中, 信号的接收和发送, 特别地, 将基站的 下行信息接收后, 交由一个或者一个以上处理器 1380处理; 另外, 将涉及上行的数据发 送给基站。 通常, RF电路 1310包括但不限于天线、 至少一个放大器、 调谐器、 一个或多 个振荡器、 用户身份模块 (SIM) 卡、 收发信机、 耦合器、 LNA (Low Noise Amplifier, 低噪声放大器) 、 双工器等。 此外, RF电路 1310还可以通过无线通信与网络和其他设备 通信。 所述无线通信可以使用任一通信标准或协议, 包括但不限于 GSM(Global System of Mobile communication, 全球移动通讯系统)、 GPRS(General Packet Radio Service, 通用分 组无线服务)、 CDMA(Code Division Multiple Access,码分多址)、 WCDMA(Wideband Code Division Multiple Access, 宽带码分多址)、 LTE(Long Term Evolution,长期演进)、 电子邮件、 SMS(Short Messaging Service, 短消息服务)等。
存储器 1320可用于存储软件程序以及模块。处理器 1380通过运行存储在存储器 1320 的软件程序以及模块, 从而执行各种功能应用以及数据处理。 存储器 1320可主要包括存 储程序区和存储数据区, 其中, 存储程序区可存储操作系统、 至少一个功能所需的应用程 序 (比如声音播放功能、 图像播放功能等) 等; 存储数据区可存储根据电子设备 1300的 使用所创建的数据 (比如音频数据、 电话本等) 等。 此外, 存储器 1320可以包括高速随 机存取存储器, 还可以包括非易失性存储器, 例如至少一个磁盘存储器件、 闪存器件、 或 其他易失性固态存储器件。 相应地, 存储器 1320还可以包括存储器控制器, 以提供处理 器 1380和输入单元 1330对存储器 1320的访问。
输入单元 1330可用于接收输入的数字或字符信息, 以及产生与用户设置以及功能控 制有关的键盘、 鼠标、 操作杆、 光学或者轨迹球信号输入。 具体地, 输入单元 1330可包 括触敏表面 1331以及其他输入设备 1332。触敏表面 1331,也称为触摸显示屏或者触控板, 可收集用户在其上或附近的触摸操作(比如用户使用手指、 触笔等任何适合的物体或附件 在触敏表面 1331上或在触敏表面 1331附近的操作), 并根据预先设定的程式驱动相应的 连接装置。 可选的, 触敏表面 1331可包括触摸检测装置和触摸控制器两个部分。 其中, 触摸检测装置检测用户的触摸方位, 并检测触摸操作带来的信号, 将信号传送给触摸控制 器; 触摸控制器从触摸检测装置上接收触摸信息, 并将它转换成触点坐标, 再送给处理器 1380, 并能接收处理器 1380发来的命令并加以执行。 此外, 可以采用电阻式、 电容式、 红外线以及表面声波等多种类型实现触敏表面 1331。 除了触敏表面 1331, 输入单元 1330 还可以包括其他输入设备 1332。具体地,其他输入设备 1332可以包括但不限于物理键盘、 功能键 (比如音量控制按键、 开关按键等) 、 轨迹球、 鼠标、 操作杆等中的一种或多种。
显示单元 1340可用于显示由用户输入的信息或提供给用户的信息以及电子设备 1300 的各种图形用户接口, 这些图形用户接口可以由图形、 文本、 图标、 视频和其任意组合来 构成。显示单元 1340可包括显示面板 1341,可选的,可以采用 LCD(Liquid Crystal Display, 液晶显示器)、 OLED(Organic Light-Emitting Diode,有机发光二极管)等形式来配置显示面板 1341。 进一步的, 触敏表面 1331可覆盖在显示面板 1341之上, 当触敏表面 1331检测到 在其上或附近的触摸操作后,传送给处理器 1380以确定触摸事件的类型,随后处理器 1380 根据触摸事件的类型在显示面板 1341上提供相应的视觉输出。 虽然在图 13中, 触敏表面 1331与显示面板 1341是作为两个独立的部件来实现输入和输入功能, 但是在某些实施例 中, 可以将触敏表面 1331与显示面板 1341集成而实现输入和输出功能。
电子设备 1300还可包括至少一种传感器 1350, 比如光传感器、 运动传感器以及其他 传感器。 具体地, 光传感器可包括环境光传感器及接近传感器, 其中, 环境光传感器可根 据环境光线的明暗来调节显示面板 1341的亮度,接近传感器可在电子设备 1300移动到耳 边时, 关闭显示面板 1341和 /或背光。 作为运动传感器的一种, 重力加速度传感器可检测 各个方向上(一般为三轴)加速度的大小, 静止时可检测出重力的大小及方向, 可用于识 别手机姿态的应用 (比如横竖屏切换、 相关游戏、 磁力计姿态校准) 、 振动识别相关功能 (比如计步器、 敲击) 等; 至于电子设备 1300 还可配置的陀螺仪、 气压计、 湿度计、 温 度计、 红外线传感器等其他传感器, 在此不再赘述。
音频电路 1360、 扬声器 1361, 传声器 1362可提供用户与电子设备 1300之间的音频 接口。 音频电路 1360可将接收到的音频数据转换后的电信号, 传输到扬声器 1361, 由扬 声器 1361转换为声音信号输出;另一方面,传声器 1362将收集的声音信号转换为电信号, 由音频电路 1360接收后转换为音频数据, 再将音频数据输出处理器 1380处理后, 经 RF 电路 1310以发送给另一终端, 或者将音频数据输出至存储器 1320以便进一步处理。 音频 电路 1360还可能包括耳塞插孔, 以提供外设耳机与电子设备 1300的通信。
短距离无线传输模块 1370可以是 WIFI (wireless fidelity, 无线保真)模块或者蓝牙模 块等。 电子设备 1300通过短距离无线传输模块 1370可以帮助用户收发电子邮件、 浏览网 页和访问流式媒体等, 它为用户提供了无线的宽带互联网访问。 虽然图 13示出了短距离 无线传输模块 1370, 但是可以理解的是, 其并不属于电子设备 1300的必须构成, 完全可 以根据需要在不改变发明的本质的范围内而省略。
处理器 1380是电子设备 1300的控制中心,利用各种接口和线路连接整个电子设备的 各个部分, 通过运行或执行存储在存储器 1320内的软件程序和 /或模块, 以及调用存储在 存储器 1320内的数据, 执行电子设备 1300的各种功能和处理数据, 从而对电子设备进行 整体监控。 可选的, 处理器 1380可包括一个或多个处理核心; 优选的, 处理器 1380可集 成应用处理器和调制解调处理器, 其中, 应用处理器主要处理操作系统、 用户界面和应用 程序等, 调制解调处理器主要处理无线通信。 可以理解的是, 上述调制解调处理器也可以 不集成到处理器 1380中。
电子设备 1300还包括给各个部件供电的电源 1390 (比如电池) , 优选的, 电源可以 通过电源管理系统与处理器 1380逻辑相连, 从而通过电源管理系统实现管理充电、 放电、 以及功耗管理等功能。 电源 1390还可以包括一个或一个以上的直流或交流电源、 再充电 系统、 电源故障检测电路、 电源转换器或者逆变器、 电源状态指示器等任意组件。
尽管未示出, 电子设备 1300还可以包括摄像头、 蓝牙模块等, 在此不再赘述。 具体 在本实施例中, 电子设备 1300的显示单元是触摸屏显示器。
电子设备 1300还包括有存储器, 以及一个或者一个以上的程序, 其中一个或者一个 以上程序存储于存储器中, 且经配置以由一个或者一个以上处理器执行。 上述一个或者一 个以上程序包含的指令用于执行一个电子书文档处理方法, 包括:
获取电子书文档;
按照预设分段方式对所述电子书文档的内容分段, 生成多个片段;
将所述多个片段组成一个有序的片段组;
在所述片段组中选择一个片段, 将所述片段作为当前片段;
解析所述当前片段的内容, 生成布局排版数据;
根据所述布局排版数据生成页面图像。
而有关上述方法的具体实现方式可以参见上述方法部分, 此处就不再赘述。
需要说明的是, 图 1至图 10所示的实施例只是本公开所介绍的优选实施例, 本领域 技术人员在此基础上, 完全可以设计出更多的实施例, 因此不在此处赘述。
本说明书中各个实施例采用递进的方式描述,每个实施例重点说明的都是与其他实施 例的不同之处,各个实施例之间相同相似部分互相参见即可。对于实施例公开的装置而言, 由于其与实施例公开的方法相对应, 所以描述的比较简单, 相关之处参见方法部分说明即 可。
本领域技术人员可以理解, 可以使用许多不同的工艺和技术中的任意一种来表示信 息、 消息和信号。 例如, 上述说明中提到过的消息、 信息都可以表示为电压、 电流、 电磁 波、 磁场或磁性粒子、 光场或以上任意组合。 专业人员还可以进一步意识到,结合本文中所公开的实施例描述的各示例的单元及算 法步骤, 能够以电子硬件、 计算机软件或者二者的结合来实现, 为了清楚地说明硬件和软 件的可互换性, 在上述说明中已经按照功能一般性地描述了各示例的组成及步骤。这些功 能究竟以硬件还是软件方式来执行, 取决于技术方案的特定应用和设计约束条件。 专业技 术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认 为超出本公开的范围。
结合本文中所公开的实施例描述的方法或算法的步骤可以直接用硬件、处理器执行的 软件模块, 或者二者的结合来实施。 软件模块可以置于随机存储器 (RAM) 、 内存、 只 读存储器 (ROM) 、 电可编程 ROM、 电可擦除可编程 ROM、 寄存器、 硬盘、 可移动磁 盘、 CD-ROM、 或技术领域内所公知的任意其它形式的存储介质中。 对所公开的实施例的 上述说明, 使本领域专业技术人员能够实现或使用本公开。
对这些实施例的多种修改对本领域的专业技术人员来说将是显而易见的,本文中所定 义的一般原理可以在不脱离本公开的精神或范围的情况下, 在其它实施例中实现。 因此, 本公开将不会被限制于本文所示的这些实施例,而是要符合与本文所公开的原理和新颖特 点相一致的最宽的范围。

Claims

权利要求
1、 一种电子书文档的处理方法, 其特征在于, 所述方法包括:
获取电子书文档;
按照预设分段方式对所述电子书文档的内容分段, 生成多个片段;
将所述多个片段组成一个有序的片段组;
在所述片段组中选择一个片段, 将所述片段作为当前片段;
解析所述当前片段的内容, 生成布局排版数据;
根据所述布局排版数据生成页面图像。
2、 根据权利要求 1所述的方法, 其特征在于, 在生成布局排版数据的步骤之后, 所 述方法还包括:
记录所述当前片段在所述片段组中的位置信息;
判断所述布局排版数据的数据量是否低于预设值,如果所述数据量低于预设值,根据 所述位置信息选择所述当前片段的下一个片段,并将所述下一个片段作为当前片段,返回 解析所述当前片段内的内容的步骤。
3、 根据权利要求 1所述的方法, 其特征在于, 所述按照预设分段方式对所述电子书 文档的内容进行分段的步骤, 包括下述子步骤:
确定分段值;
将所述电子书文档的内容拆分成多个与所述分段值大小相同的片段。
4、 根据权利要求 1所述的方法, 其特征在于, 所述方法还包括:
判断所述片段的起始点处的信息是否完整,如果不完整,在所述片段的起始点向上一 个片段移动确定所述信息的起始点,并将所述信息的起始点作为所述片段的起始点和所述 上一个片段的结束点。
5、 根据权利要求 1所述的方法, 其特征在于, 所述方法还包括:
判断所述片段的起始点处的信息是否完整,如果不完整,在所述片段的起始点向所述 片段的结束点移动确定所述信息的结束点,并将所述信息的结束点作为所述片段的起始点 和上一个片段的结束点。
6、 根据权利要求 1所述的方法, 其特征在于, 所述方法还包括:
判断所述片段的结束点处的信息是否完整,如果不完整,在所述片段的结束点向下一 个片段移动确定所述信息的结束点,并将所述信息的结束点作为所述片段的结束点和所述 下一个片段的起始点。
7、 根据权利要求 1所述的方法, 其特征在于, 所述方法还包括:
判断所述片段的结束点处的信息是否完整,如果不完整,在所述片段的结束点向所述 片段的起始点移动确定所述信息的起始点,并将所述信息的起始点作为所述片段的结束点 和下一个片段的起始点。
8、 根据权利要求 1所述的方法, 其特征在于, 所述方法还包括: 判断所述布局排版数据中包含的排版原子数据在预设时间内是否被使用,如果没有被 使用, 删除所述排版原子数据。
9、 根据权利要求 1所述的方法, 其特征在于, 所述方法还包括:
判断所述布局排版数据中包含的排版原子数据占用的内存是否大于预设值,如果大于 预设值, 删除所述排版原子数据。
10、 一种终端, 其特征在于, 所述终端包括:
获取模块, 用于获取电子书文档;
分段模块, 用于按照预设分段方式对所述电子书文档的内容分段, 生成多个片段; 组成模块, 用于将所述多个片段组成一个有序的片段组;
选择模块, 用于在所述片段组中选择一个片段, 将所述片段作为当前片段; 解析模块, 用于解析所述当前片段的内容, 生成布局排版数据;
生成模块, 用于根据所述布局排版数据生成页面图像。
11、 根据权利要求 10所述的终端, 其特征在于, 所述终端还包括:
记录模块, 用于记录所述当前片段在所述片段组中的位置信息;
第一判断模块, 用于判断所述布局排版数据的数据量是否低于预设值;
第一执行模块,用于在所述数据量低于预设值时,根据所述位置信息选择所述当前片 段的下一个片段, 并将所述下一个片段作为当前片段, 返回执行解析模块。
12、 根据权利要求 10所述的终端, 其特征在于, 所述分段模块包括:
分段值确定单元, 用于确定分段值;
拆分单元, 用于将所述电子书文档的内容拆分成多个与所述分段值大小相同的片段。
13、 根据权利要求 10所述的终端, 其特征在于, 所述终端还包括:
第二判断模块, 用于判断所述片段的起始点处的信息是否完整;
第二执行模块,用于在所述片段的起始点处的信息不完整时,在所述片段的起始点向 前移动确定所述标签的起始点, 并将所述标签的起始点作为所述片段的起始点。
14、 根据权利要求 10所述的终端, 其特征在于, 所述终端还包括:
第三判断模块, 用于判断所述片段的起始点处的信息是否完整;
第三执行模块,用于在判断所述片段的起始点处的信息不完整时,在所述片段的起始 点向所述片段的结束点移动确定所述信息的结束点,并将所述信息的结束点作为所述片段 的起始点和上一个片段的结束点。
15、 根据权利要求 10所述的终端, 其特征在于, 所述终端还包括:
第四判断模块, 用于判断所述片段的结束点处的信息是否完整;
第四执行模块,用于在所述片段的结束点处的信息不完整时,在所述片段的结束点向 下一个片段移动确定所述信息的结束点,并将所述信息的结束点作为所述片段的结束点和 所述下一个片段的起始点。
16、 根据权利要求 10所述的终端, 其特征在于, 所述终端还包括: 第五判断模块, 用于判断所述片段的结束点处的信息是否完整;
第五执行模块,用于在所述片段的结束点处的信息不完整时,在所述片段的结束点向 所述片段的起始点移动确定所述信息的起始点,并将所述信息的起始点作为所述片段的结 束点和下一个片段的起始点。
17、 根据权利要求 10所述的终端, 其特征在于, 所述终端还包括:
第六判断模块,用于判断所述布局排版数据中包含的排版原子数据在预设时间内是否 被使用;
第六执行模块,用于在所述排版原子数据在预设时间内没有被使用时,删除所述排版 原子数据。
18、 根据权利要求 10所述的终端, 其特征在于, 所述终端还包括:
第七判断模块,用于判断所述布局排版数据中包含的排版原子数据占用的内存是否大 于预设值;
第七执行模块,用于在所述排版原子数据占用的内存大于预设值时,删除所述排版原 子数据。
19、 一种电子设备, 其特征在于, 包括有存储器, 以及一个或者一个以上的程序, 其 中一个或者一个以上程序存储于存储器中,且经配置以由一个或者一个以上处理器执行所 述一个或者一个以上程序包含用于进行以下操作的指令:
获取电子书文档;
按照预设分段方式对所述电子书文档的内容分段, 生成多个片段;
将所述多个片段组成一个有序的片段组;
在所述片段组中选择一个片段, 将所述片段作为当前片段;
解析所述当前片段的内容, 生成布局排版数据;
根据所述布局排版数据生成页面图像。
PCT/CN2014/077413 2013-10-16 2014-05-14 一种电子书文档的处理方法、终端及电子设备 Ceased WO2015055002A1 (zh)

Priority Applications (6)

Application Number Priority Date Filing Date Title
KR1020147021518A KR20150055600A (ko) 2013-10-16 2014-05-14 전자책 문서 처리방법, 단말기, 전자기기, 프로그램 및 기록매체
RU2015122426A RU2617388C2 (ru) 2013-10-16 2014-05-14 Способ, терминал и электронное устройство для обработки документа электронной книги
JP2015542159A JP6118418B2 (ja) 2013-10-16 2014-05-14 電子書籍ドキュメント処理方法、端末、電子機器、プログラム及び記録媒体
MX2014009068A MX2014009068A (es) 2013-10-16 2014-05-14 Metodo, terminal y dispositivo electronico para procesar documento de libro electronico.
BR112014018483A BR112014018483A8 (pt) 2013-10-16 2014-05-14 Método, terminal e dispositivo eletrônico para processar documento de livro eletrônico
US14/471,702 US20150106697A1 (en) 2013-10-16 2014-08-28 Method and electronic device for processing e-book document

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201310485775.8 2013-10-16
CN201310485775.8A CN103593333B (zh) 2013-10-16 2013-10-16 一种电子书文档的处理方法、终端及电子设备

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US14/471,702 Continuation US20150106697A1 (en) 2013-10-16 2014-08-28 Method and electronic device for processing e-book document

Publications (1)

Publication Number Publication Date
WO2015055002A1 true WO2015055002A1 (zh) 2015-04-23

Family

ID=50083483

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2014/077413 Ceased WO2015055002A1 (zh) 2013-10-16 2014-05-14 一种电子书文档的处理方法、终端及电子设备

Country Status (8)

Country Link
EP (1) EP2863320B1 (zh)
JP (1) JP6118418B2 (zh)
KR (1) KR20150055600A (zh)
CN (1) CN103593333B (zh)
BR (1) BR112014018483A8 (zh)
MX (1) MX2014009068A (zh)
RU (1) RU2617388C2 (zh)
WO (1) WO2015055002A1 (zh)

Families Citing this family (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103593333B (zh) * 2013-10-16 2017-09-22 小米科技有限责任公司 一种电子书文档的处理方法、终端及电子设备
CN105224540A (zh) * 2014-05-29 2016-01-06 广州市动景计算机科技有限公司 页面排版方法及装置
CN107464011A (zh) * 2017-07-04 2017-12-12 林聪发 一种石板排版方法、装置、终端设备及可读存储介质
CN107861935A (zh) * 2017-11-13 2018-03-30 掌阅科技股份有限公司 躲避手指按压位置的文字重排方法、终端及存储介质
CN107967243A (zh) * 2017-11-22 2018-04-27 语联网(武汉)信息技术有限公司 一种支持用户自主断句的处理方法
CN109597980A (zh) * 2018-12-07 2019-04-09 万兴科技股份有限公司 Pdf文档分割方法、装置及电子设备
CN111191418B (zh) * 2019-12-06 2024-01-12 腾讯科技(深圳)有限公司 在线文档处理方法、装置、电子设备及计算机存储介质
CN112100176B (zh) * 2020-09-03 2024-01-19 北京得到信息科技有限公司 一种电子书阅读进度计算方法及系统
CN113515928B (zh) * 2021-07-13 2023-03-28 抖音视界有限公司 电子文本生成方法、装置、设备及介质
CN113946774B (zh) * 2021-10-26 2022-04-08 掌阅科技股份有限公司 书籍网页的跳转方法、计算设备及计算机存储介质
CN115146608B (zh) * 2022-05-13 2024-07-23 抖音视界有限公司 内容排版方法、装置、设备和存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102045388A (zh) * 2010-11-25 2011-05-04 汉王科技股份有限公司 在线阅读装置及在线阅读方法
CN102096674A (zh) * 2009-12-11 2011-06-15 华为技术有限公司 电子书发布和下载的方法、设备及系统
CN102214441A (zh) * 2010-04-01 2011-10-12 上海易狄欧电子科技有限公司 在电子书阅读器上打开文本格式电子书的方法
CN102682093A (zh) * 2012-04-25 2012-09-19 广州市动景计算机科技有限公司 一种移动浏览器网页分段加载方法及系统
CN103593333A (zh) * 2013-10-16 2014-02-19 小米科技有限责任公司 一种电子书文档的处理方法、终端及电子设备

Family Cites Families (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0612411A (ja) * 1992-06-25 1994-01-21 Sanyo Electric Co Ltd 文書処理装置
US5623679A (en) * 1993-11-19 1997-04-22 Waverley Holdings, Inc. System and method for creating and manipulating notes each containing multiple sub-notes, and linking the sub-notes to portions of data objects
US7143181B2 (en) * 2000-08-31 2006-11-28 Yohoo! Inc. System and method of sending chunks of data over wireless devices
JP2004280278A (ja) * 2003-03-13 2004-10-07 Sharp Corp データ処理装置、データ処理方法、データ処理プログラム、および、記録媒体
JP3905851B2 (ja) * 2003-03-24 2007-04-18 株式会社東芝 構造化文書の分割方法及びプログラム
US8977603B2 (en) * 2005-11-22 2015-03-10 Ebay Inc. System and method for managing shared collections
US20080301545A1 (en) * 2007-06-01 2008-12-04 Jia Zhang Method and system for the intelligent adaption of web content for mobile and handheld access
WO2010063070A1 (en) * 2008-12-03 2010-06-10 Ozmiz Pty. Ltd. Method and system for displaying data on a mobile terminal
US9002701B2 (en) * 2010-09-29 2015-04-07 Rhonda Enterprises, Llc Method, system, and computer readable medium for graphically displaying related text in an electronic document
RU118085U1 (ru) * 2012-01-19 2012-07-10 Плеадес Паблишинг, Лтд. Электронная книга
CN102750086A (zh) * 2012-05-31 2012-10-24 上海必邦信息科技有限公司 电子设备间实现无线分享显示页面控制的方法

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102096674A (zh) * 2009-12-11 2011-06-15 华为技术有限公司 电子书发布和下载的方法、设备及系统
CN102214441A (zh) * 2010-04-01 2011-10-12 上海易狄欧电子科技有限公司 在电子书阅读器上打开文本格式电子书的方法
CN102045388A (zh) * 2010-11-25 2011-05-04 汉王科技股份有限公司 在线阅读装置及在线阅读方法
CN102682093A (zh) * 2012-04-25 2012-09-19 广州市动景计算机科技有限公司 一种移动浏览器网页分段加载方法及系统
CN103593333A (zh) * 2013-10-16 2014-02-19 小米科技有限责任公司 一种电子书文档的处理方法、终端及电子设备

Also Published As

Publication number Publication date
EP2863320A2 (en) 2015-04-22
CN103593333A (zh) 2014-02-19
BR112014018483A8 (pt) 2017-07-11
CN103593333B (zh) 2017-09-22
RU2617388C2 (ru) 2017-04-24
KR20150055600A (ko) 2015-05-21
EP2863320A3 (en) 2015-09-02
EP2863320B1 (en) 2019-03-27
BR112014018483A2 (zh) 2017-06-20
JP2016502716A (ja) 2016-01-28
JP6118418B2 (ja) 2017-04-19
RU2015122426A (ru) 2017-01-10
MX2014009068A (es) 2015-08-27

Similar Documents

Publication Publication Date Title
WO2015055002A1 (zh) 一种电子书文档的处理方法、终端及电子设备
EP2990930B1 (en) Scraped information providing method and apparatus
RU2616536C2 (ru) Способ, устройство и терминальное устройство для отображения сообщений
CN103699292B (zh) 一种进入文本选择模式的方法和装置
CN103823835A (zh) 一种电子书目录的处理方法、装置及终端设备
RU2602791C2 (ru) Способ, устройство и система набора
CN109003194B (zh) 评论分享方法、终端以及存储介质
CN103543913A (zh) 一种终端设备操作方法、装置和终端设备
US9921735B2 (en) Apparatuses and methods for inputting a uniform resource locator
CN107590278A (zh) 一种基于ceph的文件预读方法及相关装置
CN106716351A (zh) 显示网页的方法和电子设备
US20170339230A1 (en) Method and apparatus for file management
CN108595520A (zh) 一种生成多媒体文件的方法和装置
CN104104711A (zh) 阅读历史处理方法和装置
US9589167B2 (en) Graphical code processing method and apparatus
CN107766370A (zh) 一种文件碎片评估方法及终端
CN116795780A (zh) 文档格式转换方法、装置、存储介质及电子设备
JP2016506587A (ja) ページ適応方法、ページ適応装置、端末装置、プログラム及び記録媒体
CN103617164A (zh) 网页预取方法、装置及终端设备
CN103440295A (zh) 一种多媒体文件上传方法及电子终端
US10140265B2 (en) Apparatuses and methods for phone number processing
CN106230919B (zh) 一种文件上传的方法和装置
CN102687566A (zh) 基于功率的速率选择
CN103533177A (zh) 一种消息浏览方法、装置和终端设备
CN106484615A (zh) 记录日志的方法和装置

Legal Events

Date Code Title Description
WWE Wipo information: entry into national phase

Ref document number: MX/A/2014/009068

Country of ref document: MX

ENP Entry into the national phase

Ref document number: 20147021518

Country of ref document: KR

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 1020147021518

Country of ref document: KR

ENP Entry into the national phase

Ref document number: 2015542159

Country of ref document: JP

Kind code of ref document: A

REG Reference to national code

Ref country code: BR

Ref legal event code: B01A

Ref document number: 112014018483

Country of ref document: BR

121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 14854032

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2015122426

Country of ref document: RU

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 14854032

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 112014018483

Country of ref document: BR

Kind code of ref document: A2

Effective date: 20140728