EP2100235A2 - Method and apparatus to automatically commit files to worm status - Google Patents
Method and apparatus to automatically commit files to worm statusInfo
- Publication number
- EP2100235A2 EP2100235A2 EP06733911A EP06733911A EP2100235A2 EP 2100235 A2 EP2100235 A2 EP 2100235A2 EP 06733911 A EP06733911 A EP 06733911A EP 06733911 A EP06733911 A EP 06733911A EP 2100235 A2 EP2100235 A2 EP 2100235A2
- Authority
- EP
- European Patent Office
- Prior art keywords
- file
- period
- worm
- autocommit
- commit
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F21/00—Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
- G06F21/70—Protecting specific internal or peripheral components, in which the protection of a component leads to protection of the entire computer
- G06F21/78—Protecting specific internal or peripheral components, in which the protection of a component leads to protection of the entire computer to assure secure storage of data
- G06F21/80—Protecting specific internal or peripheral components, in which the protection of a component leads to protection of the entire computer to assure secure storage of data in storage media based on magnetic or optical technology, e.g. disks with sectors
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/10—File systems; File servers
- G06F16/11—File system administration, e.g. details of archiving or snapshots
- G06F16/122—File system administration, e.g. details of archiving or snapshots using management policies
- G06F16/125—File system administration, e.g. details of archiving or snapshots using management policies characterised by the use of retention policies
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/10—File systems; File servers
- G06F16/18—File system types
- G06F16/1805—Append-only file systems, e.g. using logs or journals to store data
- G06F16/181—Append-only file systems, e.g. using logs or journals to store data providing write once read many [WORM] semantics
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/10—File systems; File servers
- G06F16/18—File system types
- G06F16/1865—Transactional file systems
Definitions
- At least one embodiment of the present invention pertains to networked storage systems and, more particularly, to a method and apparatus to automatically commit files to WORM status.
- WORM write once, read many
- industries such as the financial services and healthcare industries
- businesses are required by strict records-retention regulations to archive important data, such as emails, transaction information, patient records, audit information, as well as other types of documents and data.
- records-retention regulations include, for example, Securities Exchange Commission (SEC) Rule 17a-4 (17 C.F.R. ⁇ 240.17a-4(f)), which regulates broker-dealers; Health Insurance Portability and Accountability Act (HIPAA), which regulates companies in the healthcare industry; Sarbanes-Oxley (SOX), which regulates publicly traded companies; 21 C.F.R.
- SEC Securities Exchange Commission
- HIPAA Health Insurance Portability and Accountability Act
- SOX Sarbanes-Oxley
- a networked storage system may include one or more storage servers, which may be storage appliances.
- a storage server may provide services related to the organization of data on mass storage devices, such as disks.
- Some of these storage servers are commonly referred to as filers or file servers.
- An example of such a storage server is any of the Filer products made by Network Appliance, Inc. in Sunnyvale, California.
- the storage appliance may be implemented with a special-purpose computer or a general-purpose computer.
- various networked storage systems may include different numbers of storage servers.
- Various applications, including compliance applications may be permitted to create and modify data on a storage appliance.
- Some compliance applications do not have a built-in capability to assign a retention time to files or to commit files to WORM status. The files therefore may need to be committed to WORM status manually (e.g., by an administrator copying the files to a WORM storage device).
- some compliance applications may not have a capability to notify a storage system of when the application has completed the modifying operations on a file so that the file may be treated as closed.
- a compliance application communicates with a storage server via open communications protocols, such as NFS or CIFS, the network traffic may not be indicative of the status of a file with respect to the status of the file as being open or closed for further modifications.
- NFS does not have a mechanism to indicate when a file is closed.
- CIFS does have a mechanism of indicating that a file has been closed, but there are many applications that will close a file and then reopen it for writing again.
- a system and method are provided to commit files to WORM status.
- the system comprises a configuration component to set an autocommit period; a scanner to detect that the autocommit period has expired for a file; and a commit component to commit the file to write once read many (WORM) status.
- WORM write once read many
- Figure 1 is a schematic block diagram of an environment including a storage system that may be used with one embodiment of the present invention
- Figure 2 is a schematic block diagram of a storage server that may be used with one embodiment of the present invention.
- Figure 3 is a schematic block diagram of a storage operating system that may be used with one embodiment of the present invention.
- Figure 4 is a schematic block diagram of a WORM component, according to one embodiment of the present invention.
- Figure 5 is a flow chart illustrating a method to automatically commit files to WORM status, according to embodiments of the invention.
- a file may be considered to have a WORM status if the file cannot be deleted or modified until a predetermined end of retention period.
- an administrator may be permitted to configure a storage system to automatically commit files to WORM status after the files have not been modified for a predetermined period of time. This predetermined period of time may be referred to as an autocommit period.
- the autocommit period may be dependent upon a particular application, as for some applications the files may need to remain opened for writing longer than for other applications.
- the autocommit period may be applicable to files storing financial records.
- a user e.g., an administrator
- the storage server may automatically commit a file to WORM status if the file has not been changed for the autocommit period.
- a file that has been committed to WORM status remains unmodifiable for a predetermined retention period, which may be a default value.
- the storage server may send out an Enhanced Messaging Service (EMS) message to a designated recipient (e.g., an administrator) every time the system automatically commits a file to WORM status.
- EMS Enhanced Messaging Service
- the present invention may be implemented in the context of a storage-oriented network, e.g., a network that includes one or more storage servers that store and retrieve data on behalf of one or more clients.
- a storage-oriented network e.g., a network that includes one or more storage servers that store and retrieve data on behalf of one or more clients.
- Such a network may be used, for example, to provide multiple users with access to shared data or to backup mission critical data.
- An example of such a network is illustrated in Figure 1.
- FIG. 1 is a schematic block diagram of an environment 100 including a storage system 110 that may be advantageously used with one embodiment of the present invention.
- the storage system 110 may be configured to access information requested by clients such as a client 140 via a network 150.
- the storage system 110 may store files created or modified by an application 142 running on the client 140.
- the application 142 may be a compliance application.
- the storage system 110 may store data on any type of attached array of writable storage device media such as video tape, optical, DVD, magnetic tape, bubble memory, electronic random access memory, micro- electro mechanical and any other similar media adapted to store information, including data and parity information.
- the data is preferably stored on disks 120, such as hard disk drives (HDD) and/or direct access storage devices (DASD), of an array 130.
- HDD hard disk drives
- DASD direct access storage devices
- storage of information on array 130 may be implemented as one or more storage "volumes,” such as a volume 132 and a volume 134, that comprise a collection of physical storage disks 120 cooperating to define an overall logical arrangement of virtual block number (vbn) space on the volumes.
- Each logical volume is generally, although not necessarily, associated with its own file system.
- the disks within a logical volume are typically organized as one or more groups, wherein each group may be operated as a redundant array of independent disks (RAID).
- RAID implementations such as a RAID-4 level implementation, enhance the reliability/integrity of data storage through the redundant writing of data "stripes" across a given number of physical disks in the RAID group, and through the appropriate storing of parity information with respect to the striped data.
- An illustrative example of a RAID implementation is a RAID-4 level implementation, although it will be understood that other types and levels of RAID implementations may be used in accordance with the inventive principles described herein.
- a volume may be designated, e.g., at the time of volume creation, as a WORM volume, such that at least some files on a WORM volume may be committed to a WORM status to remain unmodifiable for a predetermined retention period.
- the storage system 110 may service client requests over the computer network 150.
- the computer network 150 may comprise a point-to-point connection or a shared medium, such as a local area network.
- the computer network 150 may be embodied as an Ethernet network or a Fibre Channel (FC) network.
- the client 140 may communicate with the storage system over network 150 by exchanging discrete frames or packets of data according to pre-defined protocols, such as the Transmission Control Protocol/Internet Protocol (TCP/IP).
- TCP/IP Transmission Control Protocol/Internet Protocol
- the client 140 may be a general-purpose computer configured to execute applications 142.
- the client 140 may interact with the storage system 110 in accordance with a client/server model of information delivery.
- the client may request the services of the storage system, and the system may return the results of the services requested by the client by exchanging packets over the network 150.
- the clients may issue packets including file-based access protocols, such as the Common Internet File System (CIFS) protocol or Network File System (NFS) protocol, over TCP/IP when accessing information in the form of files and directories.
- file-based access protocols such as the Common Internet File System (CIFS) protocol or Network File System (NFS) protocol
- NFS Network File System
- the client may issue packets including block-based access protocols, such as the Small Computer Systems Interface (SCSI) protocol encapsulated over TCP (iSCSI) and SCSI encapsulated over Fibre Channel (FCP), when accessing information in the form of blocks.
- SCSI Small Computer Systems Interface
- iSCSI SCSI encapsulated over TCP
- FCP Fibre Channel
- a storage system 200 comprises a processor 222, a memory 224, a network adaptor 226, and a storage adaptor 228, interconnected by a system bus 250.
- the memory 224 comprises storage locations that are addressable by the processor and adaptors for storing software program code.
- a storage operating system 300 portions of which are typically resident in memory and executed by the processing elements, functionally organizes the system 200 by, inter alia, invoking storage operations executed by the storage system. It will be apparent to those skilled in the art that other processing and memory means, including various computer readable media, may be used for storing and executing program instructions pertaining to the inventive technique described herein.
- the network adaptor 226 comprises the mechanical, electrical and signaling circuitry needed to connect the storage system 200 to clients (e.g., the clients 140 of Figure 10) over a computer network.
- FIG. 3 illustrates the operating system 300 in greater details according to one embodiment of the invention.
- the term "storage operating system” generally refers to the computer-executable code operable on a computer that manages data access and may implement file system semantics, such as the Data ONTAP® storage operating system, implemented as a microkernel, and available from Network Appliance, Inc. of Sunnyvale, California, which implements a Write Anywhere File Layout (WAFLTM) file system.
- the storage operating system can also be implemented as an application program operating over a general-purpose operating system, such as UNIX® or Windows NT®, or as a general-purpose operating system with configurable functionality, which is configured for storage applications.
- the storage operating system 300 comprises a series of software layers organized to form an integrated network protocol stack or, more generally, a multi-protocol engine that provides data paths for clients to access information stored on the storage system using block and file access protocols.
- the protocol stack includes a media access layer 310 of network drivers (e.g., gigabit Ethernet drivers) that interfaces to network protocol layers, such as the IP layer 312 and its supporting transport mechanisms, the TCP layer 314 and the User Datagram Protocol (UDP) layer 316.
- a file system protocol layer provides multi-protocol file access and, to that end, includes support for the Direct Access File System (DAFS) protocol 318, the NFS protocol 320, the CIFS protocol 322 and the Hypertext Transfer Protocol (HTTP) protocol 324.
- a virtual interface (Vl) layer 326 implements the Vl architecture to provide direct access transport (DAT) capabilities, such as remote direct memory access (RDMA), as required by the DAFS protocol 318.
- DAT Direct Access File System
- RDMA remote direct memory
- An iSCSI driver layer 328 provides block protocol access over the TCP/IP network protocol layers, while a FC driver layer 330 receives and transmits block access requests and responses to and from the storage system.
- the FC and iSCSI drivers provide FC-specific and iSCSI-specific access control to the blocks and, thus, manage exports of LUNs to either iSCSI or FCP or, alternatively, to both iSCSI and FCP when accessing the blocks on the storage system.
- the storage operating system includes a storage module embodied as a RAID system 340 that manages the storage and retrieval of information to and from the volumes/disks in accordance with I/O operations, and a disk driver system 350 that implements a disk access protocol such as, e.g., the SCSI protocol.
- a disk access protocol such as, e.g., the SCSI protocol.
- Bridging the disk software layers with the integrated network protocol stack layers is a virtualization system that is implemented by a file system 380 interacting with virtualization modules illustratively embodied as, e.g., vdisk module 390 and SCSI target module 370.
- the vdisk module 390 is layered on the file system 380 to enable access by administrative interfaces, such as a user interface (Ul) 375, in response to a user (system administrator) issuing commands to the storage system.
- the SCSI target module 370 is disposed to provide a translation layer of the virtualization system between the block (LUN) space and the file system space, where LUNs are represented as blocks.
- the Ul 375 is disposed over the storage operating system in a manner that enables administrative or user access to the various layers and systems.
- the file system 380 is illustratively a message-based system that provides logical volume management capabilities for use in access to the information stored on the storage devices, such as disks. That is, in addition to providing file system semantics, the file system 380 provides functions normally associated with a volume manager. These functions include (i) aggregation of the disks, (ii) aggregation of storage bandwidth of the disks, and (iii) reliability guarantees, such as mirroring and/or parity (RAID).
- RAID mirroring and/or parity
- the file system 380 illustratively implements a write anywhere file system having an on-disk format representation that is block-based using, e.g., 4 kilobyte (kB) blocks and using index nodes ("inodes”) to identify files and file attributes (such as creation time, access permissions, size and block location).
- kB kilobyte
- inodes index nodes
- the file system 380 may include a WORM component 382.
- the WORM component 382 may be configured to identify files that, according to a preset criteria, may be considered to be closed (i.e., when an application completed modification operations on the file) and commit files that have not been modified for a predetermined period of time to WORM status.
- FIG. 4 is a schematic block diagram of a WORM component 400, according to one embodiment of the present invention.
- the WORM component 400 may include a compliance clock 410, a configuration component 420, a scanner 430, and a commit component 440.
- the compliance clock 410 in one embodiment, is different from a system clock in that that the compliance clock 410 has certain security features that restrict any user from modifying the time on it. In one embodiment, the commands that could be used to modify the clock are disabled for the compliance clock 410.
- the time the compliance clock has is periodically written to all of the volumes on the storage system, both the WORM volumes and the non-WORM volumes.
- the compliance clock time that was last written to that volume is compared to the current compliance clock time for the system and, if the compliance clock time for the volume is earlier than the compliance clock time for the system, the system's compliance clock time is moved back to match the compliance clock time that was last written for the volume.
- the configuration component 420 may be utilized to define various settings associated with a WORM volume.
- the configuration component 420 permits an administrator to specify an autocommit period (e.g., utilizing a CLI interface or a GUI interface) that indicates when a file may be committed to WORM status after the file has been closed for writing.
- Other settings may include retention period for files that have been committed to WORM status, as well as a definition of what constitutes a modification operation performed on a file.
- the configuration component 420 may define a modification operation to include a change to a file's contents, but to exclude any change to a file's attributes. Every time a file in the storage system is created or modified (e.g., by a compliance application), a modification time stamp (here, referred to as "mtime”) is updated.
- mtime modification time stamp
- the scanner 430 may be configured to scan the files on a WORM volume and to identify those files that are closed and hence are ready to be committed to WORM status. This determination may be made, in one embodiment, by comparing the modification time stamp for a file (the mtime) with the current compliance clock value (compliance time). In one embodiment, the modification time stamp for a file (the mtime) may be stored in the file's inode along with other metadata related to the file. If the difference between the mtime and the compliance time for a file exceeds the autocommit period, the scanner 430 may identify the file as ready to be committed to WORM status. The scanner 430 then communicates information regarding the files that are ready to be committed to WORM status to the commit component 440, and the commit component 440 commits such files to WORM status.
- the operations involved in the commit process include setting the file's read-only attribute and associating an end of retention time for the file.
- the commit component 440 may determine retention time for the file utilizing a default retention time value for the volume and the modification time for a file associated with the file. Specifically, retention time may be calculated by increasing mtime by the retention period. In one embodiment, the system will protect a file from deletion and modification until retention time has been reached.
- FIG. 5 is a flowchart illustrating a method 500 to automatically commit files to WORM status, according to one embodiment of the present invention.
- the method may be performed by processing logic of the storage system that may comprise hardware (e.g., dedicated logic, programmable logic, microcode, etc.), software (such as run on a general purpose computer system or a dedicated machine), or a combination of both.
- processing logic e.g. the configuration component 420
- the autocommit period setting may be, for example, a system- wide parameter. In some embodiments, the autocommit period may be set on a per volume basis.
- the processing logic sets the retention time on a file to a default or a user-defined value (operation 504) and defines criteria to determine whether a file is closed by checking if the file has been modified (operation 506).
- the scanner 430 starts scanning the files on a WORM volume (e.g., by accessing inodes associated with the files). For each file, the scanner 430 determines whether the difference between the current compliance time and the file's mtime is greater than the autocommit period (operation 510). If the difference between the current compliance time and the file's mtime is greater than or equal to the autocommit period, the file is committed to WORM status at operation 512.
- the scanner 430 continues to scan the files on the WORM volume at operation 514.
- This apparatus may be specially constructed for the required purposes, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer.
- a computer program may be stored in a computer readable storage medium, such as, but is not limited to, any type of disk including floppy disks, optical disks, CD-ROMs, and magnetic-optical disks, read-only memories (ROMs), random access memories (RAMs) 1 EPROMs, EEPROMs, magnetic or optical cards, or any type of media suitable for storing electronic instructions, and each coupled to a computer system bus.
- a machine-readable medium includes any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer).
- a machine-readable medium includes read only memory ("ROM”); random access memory (“RAM”); magnetic disk storage media; optical storage media; FLASH memory devices; electrical, optical, acoustical or other form of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.); etc.
- ROM read only memory
- RAM random access memory
- magnetic disk storage media includes magnetic disk storage media; optical storage media; FLASH memory devices; electrical, optical, acoustical or other form of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.); etc.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Computer Security & Cryptography (AREA)
- Computer Hardware Design (AREA)
- Software Systems (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Description
Claims
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/US2006/002725 WO2007086844A2 (en) | 2006-01-25 | 2006-01-25 | Method and apparatus to automatically commit files to worm status |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP2100235A2 true EP2100235A2 (en) | 2009-09-16 |
Family
ID=38309632
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP06733911A Withdrawn EP2100235A2 (en) | 2006-01-25 | 2006-01-25 | Method and apparatus to automatically commit files to worm status |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP2100235A2 (en) |
| WO (1) | WO2007086844A2 (en) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9514150B2 (en) | 2013-04-19 | 2016-12-06 | Hewlett Packard Enterprise Development Lp | Automatic WORM-retention state transitions |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7392234B2 (en) * | 1999-05-18 | 2008-06-24 | Kom, Inc. | Method and system for electronic file lifecycle management |
| US7590807B2 (en) * | 2003-11-03 | 2009-09-15 | Netapp, Inc. | System and method for record retention date in a write once read many storage system |
| JP4401863B2 (en) * | 2004-05-14 | 2010-01-20 | 株式会社日立製作所 | Storage system |
| US20060010301A1 (en) * | 2004-07-06 | 2006-01-12 | Hitachi, Ltd. | Method and apparatus for file guard and file shredding |
-
2006
- 2006-01-25 EP EP06733911A patent/EP2100235A2/en not_active Withdrawn
- 2006-01-25 WO PCT/US2006/002725 patent/WO2007086844A2/en not_active Ceased
Non-Patent Citations (1)
| Title |
|---|
| See references of WO2007086844A2 * |
Also Published As
| Publication number | Publication date |
|---|---|
| WO2007086844A3 (en) | 2009-12-23 |
| WO2007086844A2 (en) | 2007-08-02 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN101443760B (en) | System and method for record retention date in a write once read much storage system | |
| EP1875393B1 (en) | Architecture for supporting sparse volumes | |
| JP4310338B2 (en) | Single write multiple read storage system and method for implementing the same | |
| US7774610B2 (en) | Method and apparatus for verifiably migrating WORM data | |
| EP1849056B9 (en) | System and method for enabling a storage system to support multiple volume formats simultaneously | |
| US7752401B2 (en) | Method and apparatus to automatically commit files to WORM status | |
| US20070124341A1 (en) | System and method for restoring data on demand for instant volume restoration | |
| US9449007B1 (en) | Controlling access to XAM metadata | |
| US8793226B1 (en) | System and method for estimating duplicate data | |
| US20090030983A1 (en) | System and method for non-disruptive check of a mirror | |
| US8719535B1 (en) | Method and system for non-disruptive migration | |
| US7694095B2 (en) | Managing snapshots using messages | |
| US8135676B1 (en) | Method and system for managing data in storage systems | |
| US20050210028A1 (en) | Data write protection in a storage area network and network attached storage mixed environment | |
| US7882086B1 (en) | Method and system for portset data management | |
| US7926049B1 (en) | System and method for determining differences between software configurations | |
| US7640279B2 (en) | Apparatus and method for file-level replication between two or more non-symmetric storage sites | |
| JP2008539521A (en) | System and method for restoring data on demand for instant volume restoration | |
| WO2007086844A2 (en) | Method and apparatus to automatically commit files to worm status | |
| US7836020B1 (en) | Method and apparatus to improve server performance associated with takeover and giveback procedures |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20080801 |
|
| AK | Designated contracting states |
Kind code of ref document: A2 Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LI LT LU LV MC NL PL PT RO SE SI SK TR |
|
| AX | Request for extension of the european patent |
Extension state: AL BA HR MK YU |
|
| R17D | Deferred search report published (corrected) |
Effective date: 20091223 |
|
| RIC1 | Information provided on ipc code assigned before grant |
Ipc: G06F 3/06 20060101AFI20100112BHEP |
|
| R17P | Request for examination filed (corrected) |
Effective date: 20080801 |
|
| DAX | Request for extension of the european patent (deleted) | ||
| RAP1 | Party data changed (applicant data changed or rights of an application transferred) |
Owner name: NETAPP, INC. |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20170801 |