From Nathan_Rutman at xyratex.com Wed Feb 2 18:34:24 2011 From: Nathan_Rutman at xyratex.com (Nathan Rutman) Date: Wed, 2 Feb 2011 10:34:24 -0800 Subject: [Twg] Lustre Requirements and Roadmap In-Reply-To: <765768D8E686B84BACC0EEE5FEB80DCE5481051041@MAILBOXCLUSTER.datadirect.datadirectnet.com> References: <765768D8E686B84BACC0EEE5FEB80DCE5481051041@MAILBOXCLUSTER.datadirect.datadirectnet.com> Message-ID: <73AED5C780AE05478241DB067651A92101D11FAF@XYUS-EX22.xyus.xyratex.com> Is it possible to get a list of all the Lustre-centric projects / enhancements that various groups are working on already? Wide striping? Imperative recovery? Readdir+? NRS? Should these be removed, or at least noted as "in progress by _responsible party_" on the OpenSFS requirements list? On Jan 27, 2011, at 12:44 PM, Atul Vidwansa wrote: > > Summary of F2F Requirements Discussion (1/26/2011) > -------------------------------------------------- > In the last few weeks, the TWG members gathered the following requirements from within their organizations and the Lustre community: > > - Improved error reporting > - Online filesystem integrity checks > - Integrated performance monitoring > - Completing OSD restructuring/kDMU work > - Size on MDS > - BTRFS investigation > - Imperative recovery > - Improved admin/mgmt tools > - Improved metadata performance > - LNET Channel Bonding > - end-to-end data integrity > - interface to HSM > > Further discussion within the TWG organized these requirements into two high priority categories: > > * improve metadata performance > * improve scalability and reliability of backend file system > > At our 1/20 meeting, TWG members discussed several options to meet these > goals: > > * metadata performance > - clustered metadata servers (CMD) > - network request scheduler (NRS) > - RPC aggregation > - SMP scaling > - subtree lockings > - size on MDS > > * backend storage > - OSD restructuring > - online fsck > - btrfs evaluation > > Much of the design and implementation for these features was started by Sun/Oracle. At the F2F meeting in Chicago, the OpenSFS board decided that there is an opportunity to acknowledge any technical debt in these initial designs and requested that, instead of an RFI for these features, the TWG specify our long-term requirements for Lustre and then engage the Lustre community to discuss architectures to meet these requirements. > > Therefore, the TWG will spend the next two meetings refining the requirements for metadata and backend storage. We will then spend the following two weeks leading discussions with the Lustre community to define the architecture and features that will meet the requirements. > From these discussions, the TWG will create a roadmap that OpenSFS can use to direct feature development for Lustre 2.2 and beyond. > > The following is an incomplete list of requirements to motivate further discussion : > > * metadata performance > > GOAL: improve file system scalability and interactive > performance > > requirements: min max > - # files in file system 100 billion 1 trillion > - # files in directory 50 million 10 billion > - file creates / sec 100 thousand 30 thousand > (aggregate) (single client) > - directory lisings / sec > - open files per process - 100 thousand > - file system capacity 30 PB 100 PB > - # clients 30 thousand ?00 thousand > > > * backend storage > > GOAL: provide reliable, scalable backing store for > Lustre servers > > requirements: > - large LUNs (min 32 TB) > - end-to-end data integrity (T10 PI or equivalent) > - no performance impact for file system repair > - framework to enable alternatives to ldiskfs > - direct I/O mode > - ?? > > ______________________________________________________________________ This email may contain privileged or confidential information, which should only be used for the purpose for which it was sent by Xyratex. No further rights or licenses are granted to use such information. If you are not the intended recipient of this message, please notify the sender by return and delete it. You may not use, copy, disclose or rely on the information contained in it. Internet email is susceptible to data corruption, interception and unauthorised amendment for which Xyratex does not accept liability. While we have taken reasonable precautions to ensure that this email is free of viruses, Xyratex does not accept liability for the presence of any computer viruses in this email, nor for any losses caused as a result of viruses. Xyratex Technology Limited (03134912), Registered in England & Wales, Registered Office, Langstone Road, Havant, Hampshire, PO9 1SA. The Xyratex group of companies also includes, Xyratex Ltd, registered in Bermuda, Xyratex International Inc, registered in California, Xyratex (Malaysia) Sdn Bhd registered in Malaysia, Xyratex Technology (Wuxi) Co Ltd registered in The People's Republic of China and Xyratex Japan Limited registered in Japan. ______________________________________________________________________ From carrier at cray.com Thu Feb 3 09:41:31 2011 From: carrier at cray.com (John Carrier) Date: Thu, 3 Feb 2011 03:41:31 -0600 Subject: [Twg] 2011-01-27 meeting minutes and concall reminder Message-ID: Hi all, Minutes from last week's meeting are attached. Please send any corrections to the reflector. I have sent a summary of the requirements discussion to discuss at lists.opensfs.org. Please continue the discussion on that reflector, which is open for public subscription. And, finally, a reminder that we have a call later this morning to continue discussing requirements for the architecture. We need to prepare for the community conference calls that will start next week including its schedule. Thanks, --jc -------------- next part -------------- OpenSFS Technical Working Group Meeting minutes : 01/27/2011 Concall: start : 9:30a PT, end : 10:32a PT) Next meeting: Thursday, 2/3/2011 @ 9:30a PT/12:30p ET Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. attending Name Organization email ----------------- -------------- ---------------------------- John Carrier Cray carrier at cray.com Kit Westeneat DDN kwesteneat at ddn.com Justin Miller IU jupmille at indiana.edu Steve Simms IU ssimms at indiana.edu Damian Hazen LBL dhazen at lbl.gov Marc Stearman LLNL marc at llnl.gov Ken Hornstein NRL kenh at cmf.nrl.navy.mil Dave Dillow ORNL dillowda at ornl.gov Sarp Oral ORNL oralhs at ornl.gov Andreas Dilger Whamcloud adilger at whamcloud.com Eric Barton Whamcloud eeb at whamcloud.com Robert Read Whamcloud rread at whamcloud.com Nathan Rutman Xyratex nathan_rutman at xyratex.com Peter Bojanic Xyratex peter_bojanic at xyratex.com Agenda summary of F2F discussion on 1/26 Discussion John briefly reviewed his notes from the F2F meeting on 1/26 (see email titled "[Twg] Lustre Requirements and Roadmap" in the archive). The main point is that the OpenSFS Board felt that the RFP should be based on requirements, not features and asked the TWG to pursue a roadmap based on performance requirements for Lustre from the community. Based on our previous efforts to gather feature requirements, we propose that we focus on requirements for metadata performance and the backend storage without specifying the solutions. Dave lead a discussion of how to state performance requirements. His users need 'du' to be faster. Should the requirement specify stats per second or the particular command he wants to go faster? Eric discussed the problems with the existing architecture. Several groups are pursuing particular features to improve metadata performance. Andreas thought that OpenSFS should fund features that vendors don't want to fund, such as portals cleanup. Eric mentioned the OSD restructuring, which is needed to complete work with btrfs. The working group needs to identify requirements that motivate these projects before we can get them in an OpenSFS RFP. Eric lead a discussion of considering projects composed of multiple parts. Not all projects will yield a direct benefit but are pieces of implementation needed to reach a longer-term goal. The working group will need to consider creating requirements in terms of these intermediate steps. There is also the concern for projects with multiple collaborators. Whereever development contracts go, designs need to be in the open. Requirements will define what is delivered for the RFP, but there must be agreement on the underlying architecture to develop them. Sun's HPCS contract with Cray was one example of structuring RFPs to allow developers to design, architect, and implement features based on requirements. Eric mentioned that the requirements list didn't include capacities or rates. And the focus on metadata and backend storage misses requirements to make Lustre easier to manage. OpenSFS wants to keep this discussion out in the open. Therefore, we are moving the discussion of requirements, architecture, and roadmap to discuss at lists.opensfs.org. John will send out the initial list from his notes to this reflector and lustre-community at lists.lustre.org. meeting end : 10:32 PT From carrier at cray.com Thu Feb 3 09:47:20 2011 From: carrier at cray.com (John Carrier) Date: Thu, 3 Feb 2011 03:47:20 -0600 Subject: [Twg] 2011-01-27 meeting minutes and concall reminder In-Reply-To: References: Message-ID: Should have included this in the original email: Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. -----Original Message----- From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of John Carrier Sent: Thursday, February 03, 2011 1:42 AM To: twg at lists.opensfs.org Subject: [Twg] 2011-01-27 meeting minutes and concall reminder Hi all, Minutes from last week's meeting are attached. Please send any corrections to the reflector. I have sent a summary of the requirements discussion to discuss at lists.opensfs.org. Please continue the discussion on that reflector, which is open for public subscription. And, finally, a reminder that we have a call later this morning to continue discussing requirements for the architecture. We need to prepare for the community conference calls that will start next week including its schedule. Thanks, --jc From adilger at whamcloud.com Thu Feb 3 18:44:17 2011 From: adilger at whamcloud.com (Andreas Dilger) Date: Thu, 3 Feb 2011 11:44:17 -0700 Subject: [Twg] Lustre Requirements and Roadmap In-Reply-To: <73AED5C780AE05478241DB067651A92101D11FAF@XYUS-EX22.xyus.xyratex.com> References: <765768D8E686B84BACC0EEE5FEB80DCE5481051041@MAILBOXCLUSTER.datadirect.datadirectnet.com> <73AED5C780AE05478241DB067651A92101D11FAF@XYUS-EX22.xyus.xyratex.com> Message-ID: On the TWG call today, we discussed a list of requirements that will be posted in the next couple of days. I will annotate that list with a status of idea, design, code, test, and active (for those few projects that are actively being worked on). That will help in doing prioritization to know how close a project is to completion. Unfortunately, I don't think there are many active projects under development yet. Cheers, Andreas On 2011-02-02, at 11:34, "Nathan Rutman" wrote: > Is it possible to get a list of all the Lustre-centric projects / enhancements that various groups are working on already? > Wide striping? Imperative recovery? Readdir+? NRS? > Should these be removed, or at least noted as "in progress by _responsible party_" on the OpenSFS requirements list? > > > On Jan 27, 2011, at 12:44 PM, Atul Vidwansa wrote: > >> >> Summary of F2F Requirements Discussion (1/26/2011) >> -------------------------------------------------- >> In the last few weeks, the TWG members gathered the following requirements from within their organizations and the Lustre community: >> >> - Improved error reporting >> - Online filesystem integrity checks >> - Integrated performance monitoring >> - Completing OSD restructuring/kDMU work >> - Size on MDS >> - BTRFS investigation >> - Imperative recovery >> - Improved admin/mgmt tools >> - Improved metadata performance >> - LNET Channel Bonding >> - end-to-end data integrity >> - interface to HSM >> >> Further discussion within the TWG organized these requirements into two high priority categories: >> >> * improve metadata performance >> * improve scalability and reliability of backend file system >> >> At our 1/20 meeting, TWG members discussed several options to meet these >> goals: >> >> * metadata performance >> - clustered metadata servers (CMD) >> - network request scheduler (NRS) >> - RPC aggregation >> - SMP scaling >> - subtree lockings >> - size on MDS >> >> * backend storage >> - OSD restructuring >> - online fsck >> - btrfs evaluation >> >> Much of the design and implementation for these features was started by Sun/Oracle. At the F2F meeting in Chicago, the OpenSFS board decided that there is an opportunity to acknowledge any technical debt in these initial designs and requested that, instead of an RFI for these features, the TWG specify our long-term requirements for Lustre and then engage the Lustre community to discuss architectures to meet these requirements. >> >> Therefore, the TWG will spend the next two meetings refining the requirements for metadata and backend storage. We will then spend the following two weeks leading discussions with the Lustre community to define the architecture and features that will meet the requirements. >> From these discussions, the TWG will create a roadmap that OpenSFS can use to direct feature development for Lustre 2.2 and beyond. >> >> The following is an incomplete list of requirements to motivate further discussion : >> >> * metadata performance >> >> GOAL: improve file system scalability and interactive >> performance >> >> requirements: min max >> - # files in file system 100 billion 1 trillion >> - # files in directory 50 million 10 billion >> - file creates / sec 100 thousand 30 thousand >> (aggregate) (single client) >> - directory lisings / sec >> - open files per process - 100 thousand >> - file system capacity 30 PB 100 PB >> - # clients 30 thousand ?00 thousand >> >> >> * backend storage >> >> GOAL: provide reliable, scalable backing store for >> Lustre servers >> >> requirements: >> - large LUNs (min 32 TB) >> - end-to-end data integrity (T10 PI or equivalent) >> - no performance impact for file system repair >> - framework to enable alternatives to ldiskfs >> - direct I/O mode >> - ?? >> >> > ______________________________________________________________________ > This email may contain privileged or confidential information, which should only be used for the purpose for which it was sent by Xyratex. No further rights or licenses are granted to use such information. If you are not the intended recipient of this message, please notify the sender by return and delete it. You may not use, copy, disclose or rely on the information contained in it. > > Internet email is susceptible to data corruption, interception and unauthorised amendment for which Xyratex does not accept liability. While we have taken reasonable precautions to ensure that this email is free of viruses, Xyratex does not accept liability for the presence of any computer viruses in this email, nor for any losses caused as a result of viruses. > > Xyratex Technology Limited (03134912), Registered in England & Wales, Registered Office, Langstone Road, Havant, Hampshire, PO9 1SA. > > The Xyratex group of companies also includes, Xyratex Ltd, registered in Bermuda, Xyratex International Inc, registered in California, Xyratex (Malaysia) Sdn Bhd registered in Malaysia, Xyratex Technology (Wuxi) Co Ltd registered in The People's Republic of China and Xyratex Japan Limited registered in Japan. > ______________________________________________________________________ > > > _______________________________________________ > twg mailing list > twg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/twg-opensfs.org From dillowda at ornl.gov Fri Feb 4 23:58:28 2011 From: dillowda at ornl.gov (David Dillow) Date: Fri, 04 Feb 2011 18:58:28 -0500 Subject: [Twg] Meeting notes from Feb 13 con call Message-ID: <1296863908.23654.27.camel@lap75545.ornl.gov> So much for COB today. :/ Here's my rough notes, and Shay's much better capture of the discussion from yesterday's meeting. I'll send my rough draft of the requirements under separate cover. -- Dave Dillow National Center for Computational Science Oak Ridge National Laboratory (865) 241-6602 office -------------- next part -------------- OpenSFS Technical Working Group Meeting minutes for Feb 3, 2011 Agenda Requirements gathering Listing of projects in progress Schedule for architecture meeting Are we even ready for the meeting? If we put it off, can prinicpals still attend? Attending: David Dillow ORNL dillowda at ornl.gov Shea Seager OpenSFS shay at opensfs.org Kit Westeneat DDN kwesteneat at ddn.com Chris Morrone LLNL morrone2 at llnl.gov Justin Miller IU jupmille at indiana.edu Steve Simms IU ssimms at indiana.edu Andreas Dilger Whamcloud adilger at whamcloud.com Eric Barton Whamcloud eeb at whmcloud.com Cory Spitz Cray spitzcor at cray.com Nathan Hale IU Requirements Gathering Eric -- 10 billion files? what does ls do? break posix in such a large directory? what is driving the requirement? open files per process ==> open files per client? total open files per system? * scalable fault managment fast failover -- O(log n) on service interruption bound pause + fault handling replication to acheive fault tolerance sync vs async replication quality of service * per job/cluster/user * guaranteed share of IOPS/BW allows priority levels so simulation doesn't stutter during checkpoint -- how to handle fairness? keep ls from going away while jon opens checkpoint files metadata -- find has a similar set of requirements full bandwidth and IOPs for single shared files Integrity check requirements -- online vs offline Allowing dynamic configuration of networking vs static config -- implementation issue-- killing old RPC conneciton Changing OST/MDT/MGT creation -- binding targets to specific local networks explicitly Improved LNET bandwidth and reliability -- channel bonding Patchless server and better support for upstream kernels Administrative server shutdown -- clean shutdown instead of crashing -- tell clients it is going away, no new actions until it is back -- ancillry "rolling upgrades" -- ability to quiese the clients activity to a server without unmounting filesystem space balancing migration -- empty an OST -- rebalance over new OST/MDT as added -- pool management (Use quota/ACLS for access to pool storage) Mandatory pool authorization for use Improved security -- user name mapping -- propogating authorization from MDT to OST through trusted channels rather than untrusted clients ADIO/MPI-IO improvements collective open? Improved small file IO performance Better ability for user to specify expected access patterns -- fadvise() to say "I will use this range, get locks appropriately" Improved analysis and visualization tools? Linux AIO implementaion readdir+/statlite/collection open from POSIX HPC ext WG Better userspacce tools -- redo llapi -- lctl Listing of projects in progress -- kerberized connections per user to MDT, per client to OST (Lustre 2) -- ORNL metadata through 2.x stabilzation imperative recovery through LNET -- wide striping -- mostly complete -- OSD restructuring -- well under way, no current funding -- IU/NRL uid mapping -- LLNL stabilaztion and Lustre on ZFS on Linux -- DDN -- bugfixes for 1.8, possible plans for HPCFS work? Schedule for architecture meeting After brief discussion, we don't feel we're ready for the architecture discussion. We need to refine our requirements prior to the meeting on Feb 10th, and use that meeting to generate at least a rough prioritization to guide the arch discussion. Currently plan to move the architecure discussion to Feb 17th. -------------- next part -------------- A non-text attachment was scrubbed... Name: 2-3-11 TWG concall.doc Type: application/msword Size: 50176 bytes Desc: not available URL: From dillowda at ornl.gov Sat Feb 5 00:01:42 2011 From: dillowda at ornl.gov (David Dillow) Date: Fri, 04 Feb 2011 19:01:42 -0500 Subject: [Twg] Rough requirements Message-ID: <1296864102.23654.31.camel@lap75545.ornl.gov> Extremely rough draft, I'm sure I didn't capture everything for our discussion Thursday, and didn't get to incorporate discussion on the lists. My comments/unfinished text is in [[ ]] pairs. We need to divide things up by area, and try to polish enough to send to the open reflector by late Monday. Performance requirements Metadata Performance [[ mainly an import of John's list ]] * metadata performance GOAL: improve file system scalability and interactive performance requirements: Q2 2012 Q1 2014 -------------- ----------------- - # files in file system 100 billion 1 trillion - # files in directory 50 million 10 billion - aggregate file creates/s 100 thousand ? - single file creates/s ? 30 thousand - directory listings/s ? ? - open files per client - 100 thousand - file system capacity 30 PB 100 PB - # clients 30 thousand ? - ... [[ There is a question as to what's driving 10 billion files in a directory? How does one expect ls to work in that directory? Is that expected to be POSIX compliant? ]] [[ open files per client instead of open files per process? or is that total open files per system or job? We've already passed that it seems... ]] * backend storage GOAL: provide reliable, scalable backing store for Lustre servers requirements: - large LUNs (min?, max?) - end-to-end data integrity (ie provide resiliency that T10 PI gives local file systems) - low performance impact for file system repair [[ suggest 32 TB for minimum max LUN size supported near term, 64 TB for mid-term ]] [[ Eric notes that find has a similar set of requirements as ls ]] [[ full bandwidth and IOPs for a single shared files -- this is wide stripe ]] [[ Improved small file IO performance ]] Quality of service Existing deployments currently have no mechanism to balance the performance needs of interactive users against the needs of large-scale compute jobs -- it is possible and likely that directory listings will encounter absurdly long execution times when competing against a 200,000 core checkpoint operation. To maintain usability in such scenarios, Lustre must be able to allocate a {job,cluster,user} a share of IOPS and bandwidth consumate with the priority levels assigned by an administrator. Scalable fault management While it is already the case that today's supercomputers have a marked dependence on their file systems for productive use, this dependency will continue to rise as we see more and more center-wide file systems. To minimize the downtime for the entire center, reliability must increase and recovery from faults must be bounded in time. Lustre must be able to recovery in O(log n) time or better as a mid-term goal to meet this requirement. [[ Replication to achieve fault tolerance was mentioned -- requirement vs implementation detail? ]] Improved configuration of Lustre The current mechanism of using module parameters is relatively inflexible and can constrain large, complex Lustre configurations. As a short- to mid-term requirement, Lustre must have the ability to allow dynamic configuration of the LNET interfaces and routes. This ability should include the ability to bind targets (OST/MDT/etc) to specific LNET interfaces, and to hot-add or hot-remove LNET interfaces. Allowing for controlled partial-system maintenance Currently, to upgrade a Lustre installation or perform maintenance on a subset of the comprising hardware, one must unmount the filesystem from all clients or risk hanging processes until the hardware is back online (maintenance) or other odd, undefined client behavior once the upgrade completes. To allow more flexible administration, the filesystem must be able to handle these situations gracefully, and allow the clients to avoid attempting to use hardware known to be down. [[ AKA Administrative Server Shutdown ]] Balancing storage use Currently, ensuring a balanced use of the storage space available to Lustre relies on a haphazard set of setting default stripping, storage pools, and manual rebalancing of overfull OSTs. As a mid-term goal, Lustre must be able to allow automatic emptying of an OST, migrating the data to other devices in the filesystem. Similarly, Lustre must be able to rebalance the storage load over new OSTs as they are added. Additionally, Lustre must be able require authorization for use of specific storage pools. more management [[ how to handle integrity check requirements, split online/offline? Improved analysis and visualization tools? Better userspace tools -- redo llapi -- lctl ]] Improved storage semantics/interfaces [[ need verbiage, possible requirements: ADIO/MPI-IO improvements collective open? Better ability for user to specify expected access patterns -- fadvise() to say "I will use this range, get the appropriate locks" Support Linux AIO to make O_DIRECT more useful readdir+/statlite/collective open from POSIX HECEWG ]] Misc [[ Improved LNET bandwidth and reliability -- channel bonding Improved security -- user name mapping -- propagating authorization from MDT to OST through trusted channels rather than untrusted clients Patchless server and better support for upstream kernels ]] -- Dave Dillow National Center for Computational Science Oak Ridge National Laboratory (865) 241-6602 office From adilger at whamcloud.com Tue Feb 8 08:06:51 2011 From: adilger at whamcloud.com (Andreas Dilger) Date: Tue, 8 Feb 2011 00:06:51 -0800 Subject: [Twg] Rough requirements In-Reply-To: <1296864102.23654.31.camel@lap75545.ornl.gov> References: <1296864102.23654.31.camel@lap75545.ornl.gov> Message-ID: I'm just adding some notes for each item to give an idea where it stands today. idea = just something that was discussed design = a written design exists code = some code exists Also, I won't be able to attend the Thursday meeting, possibly Eric will be able to attend. On 2011-02-04, at 16:01, David Dillow wrote: > Performance requirements > > Metadata Performance > [[ mainly an import of John's list ]] > > * metadata performance > > GOAL: improve file system scalability and interactive > performance > > requirements: Q2 2012 Q1 2014 > -------------- ----------------- > - # files in file system 100 billion 1 trillion > - # files in directory 50 million 10 billion The 2012 numbers could be attainable with CMD + ldiskfs, if one considers 1-2B inodes/fs, 0.5-1M inodes/dir and 50-100 MDTs. I don't think the 2014 numbers are achievable with ldiskfs, unless we really push the number of MDTs to the ~500 range. It is definitely higher than the number we were thinking about for CMD. > - aggregate file creates/s 100 thousand ? > - single file creates/s ? 30 thousand > - directory listings/s ? ? > - open files per client - 100 thousand > - file system capacity 30 PB 100 PB > - # clients 30 thousand ? I would put most of these requirements in the "idea/code" stage. Some CMD code exists, some needs to be written and has only been discussed. > [[ There is a question as to what's driving 10 billion files in a directory? > How does one expect ls to work in that directory? Is that expected to > be POSIX compliant? ]] > [[ open files per client instead of open files per process? or is that > total open files per system or job? We've already passed that it seems... ]] > > * backend storage > > GOAL: provide reliable, scalable backing store for > Lustre servers > > requirements: > - large LUNs (min?, max?) > [[ suggest 32 TB for minimum max LUN size supported near term, 64 TB for > mid-term ]] code exists in e2fsprogs, ext4, needs testing, integrated testing with Lustre. Also needs additional fixes to support large objects (> 2TB) > - end-to-end data integrity > (ie provide resiliency that T10 PI gives local file systems) idea - T10 for ldiskfs design - end-to-end checksum design for HPCS > - low performance impact for file system repair idea/code - slow OST avoidance, improved MDS object allocation algorithm > [[ Eric notes that find has a similar set of requirements as ls ]] > [[ full bandwidth and IOPs for a single shared files -- this is wide stripe ]] > [[ Improved small file IO performance ]] > > > Quality of service > > Existing deployments currently have no mechanism to balance the > performance needs of interactive users against the needs of large-scale > compute jobs -- it is possible and likely that directory listings will > encounter absurdly long execution times when competing against a 200,000 > core checkpoint operation. To maintain usability in such scenarios, Lustre > must be able to allocate a {job,cluster,user} a share of IOPS and bandwidth > consumate with the priority levels assigned by an administrator. code - prototype NRS code exists, needs review, testing > Scalable fault management > > While it is already the case that today's supercomputers have a > marked dependence on their file systems for productive use, this dependency > will continue to rise as we see more and more center-wide file systems. To > minimize the downtime for the entire center, reliability must increase and > recovery from faults must be bounded in time. Lustre must be able to recovery > in O(log n) time or better as a mid-term goal to meet this requirement. > [[ Replication to achieve fault tolerance was mentioned -- > requirement vs implementation detail? ]] idea > Improved configuration of Lustre > > The current mechanism of using module parameters is relatively > inflexible and can constrain large, complex Lustre configurations. As a > short- to mid-term requirement, Lustre must have the ability to allow > dynamic configuration of the LNET interfaces and routes. This ability should > include the ability to bind targets (OST/MDT/etc) to specific LNET interfaces, > and to hot-add or hot-remove LNET interfaces. idea > Allowing for controlled partial-system maintenance > > Currently, to upgrade a Lustre installation or perform maintenance > on a subset of the comprising hardware, one must unmount the filesystem from > all clients or risk hanging processes until the hardware is back online > (maintenance) or other odd, undefined client behavior once the upgrade > completes. To allow more flexible administration, the filesystem must be able > to handle these situations gracefully, and allow the clients to avoid > attempting to use hardware known to be down. > [[ AKA Administrative Server Shutdown ]] idea > Balancing storage use > > Currently, ensuring a balanced use of the storage space available > to Lustre relies on a haphazard set of setting default stripping, storage > pools, and manual rebalancing of overfull OSTs. As a mid-term goal, Lustre > must be able to allow automatic emptying of an OST, migrating the data > to other devices in the filesystem. Similarly, Lustre must be able to > rebalance the storage load over new OSTs as they are added. Additionally, > Lustre must be able require authorization for use of specific storage pools. idea - object migration idea - improved MDS object allocation > more management > [[ > how to handle integrity check requirements, split online/offline? > > Improved analysis and visualization tools? code - some work has been done on this to convert Lustre RPCTRACE logs to OTF for viewing in VampirTrace. Is that an open-source tool? > Better userspace tools > -- redo llapi > -- lctl > ]] > > Improved storage semantics/interfaces > [[ need verbiage, possible requirements: > ADIO/MPI-IO improvements > collective open? idea > Better ability for user to specify expected access patterns > -- fadvise() to say "I will use this range, get the > appropriate locks" idea > Support Linux AIO to make O_DIRECT more useful idea > readdir+/statlite/collective open from POSIX HECEWG idea/code - some code proposed to kernel developers in the past, got bogged down in a morass of feature bloat. > Misc > [[ > Improved LNET bandwidth and reliability > -- channel bonding idea > Improved security > -- user name mapping code for 1.8. Needs efficient NID matching, per discussions with Eric in the distant past. > -- propagating authorization from MDT to OST through > trusted channels rather than untrusted clients idea > Patchless server and better support for upstream kernels idea/design - detailed descriptions of how to remove each patch, some work done for RHEL6/Lustre 2.1 Cheers, Andreas -- Andreas Dilger Principal Engineer Whamcloud, Inc. From dillowda at ornl.gov Tue Feb 8 21:20:24 2011 From: dillowda at ornl.gov (David Dillow) Date: Tue, 08 Feb 2011 16:20:24 -0500 Subject: [Twg] Rough requirements In-Reply-To: References: <1296864102.23654.31.camel@lap75545.ornl.gov> Message-ID: <1297200024.9572.21.camel@obelisk.thedillows.org> On Tue, 2011-02-08 at 03:06 -0500, Andreas Dilger wrote: > I'm just adding some notes for each item to give an idea where it > stands today. Thanks for the update, Andreas. I don't think I'll merge these into the requirements document at this time, but they will be useful for the prioritization and architecture discussions. -- Dave Dillow National Center for Computational Science Oak Ridge National Laboratory (865) 241-6602 office From dillowda at ornl.gov Tue Feb 8 21:40:33 2011 From: dillowda at ornl.gov (David Dillow) Date: Tue, 08 Feb 2011 16:40:33 -0500 Subject: [Twg] Updated draft requirements posted to discuss Message-ID: <1297201233.9572.51.camel@obelisk.thedillows.org> A few small updates, added some text for storage semantics and kernel patches. There are number of items marked by [[ ]] pairs that need expansion. Chris, I think the "redo llapi and lctl" items came from you -- I don't know what you had in mind there, so please provide some text, or redirect me if I misremember the origin. John, can you explain some of the driving forces behind the 10 billion files in a directory -- how does one plan to work on such a system; will it be expected to be POSIX compliant? Steve/Justin -- while I think the user name mapping was mentioned by Eric, I seem to recall that you guys are working in this area for the WAN projects. Can you elaborate some text here? All: we need to fill out the table quite a bit. -- Dave Dillow National Center for Computational Science Oak Ridge National Laboratory (865) 241-6602 office From carrier at cray.com Tue Feb 8 23:13:26 2011 From: carrier at cray.com (John Carrier) Date: Tue, 8 Feb 2011 17:13:26 -0600 Subject: [Twg] Updated draft requirements posted to discuss In-Reply-To: <1297201233.9572.51.camel@obelisk.thedillows.org> References: <1297201233.9572.51.camel@obelisk.thedillows.org> Message-ID: Dave Dillow wrote: > John, can you explain some of the driving forces behind the 10 billion > files in a directory -- how does one plan to work on such a system; will > it be expected to be POSIX compliant? This is an HPCS requirement. Note that this was a prediction of capacity needs made in 2006 for system that was to have been delivered in 2010. Though the mission partners were pretty good at guessing how I/O would have to scale, some of their requirements are still a stretch today. --jc From carrier at cray.com Thu Feb 10 07:31:28 2011 From: carrier at cray.com (John Carrier) Date: Thu, 10 Feb 2011 01:31:28 -0600 Subject: [Twg] architecture meeting reminder Message-ID: This is a reminder of our meeting tomorrow 2/10 at 9:30a PT / 12:30 ET to continue the detailed discussion of requirements for the Lustre roadmap. Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Thanks, --jc From carrier at cray.com Thu Feb 10 17:34:06 2011 From: carrier at cray.com (John Carrier) Date: Thu, 10 Feb 2011 11:34:06 -0600 Subject: [Twg] FW: Fujitsu Requirements Message-ID: -----Original Message----- From: Eric Barton [mailto:eeb at whamcloud.com] Sent: Thursday, February 10, 2011 9:33 AM To: John Carrier Subject: Fujitsu Requirements John, Here is a presentation from Fujitsu on their Lustre requirements. It's public information and can be posted on mailing lists and the website if required.... Cheers, Eric -------------- next part -------------- A non-text attachment was scrubbed... Name: Fujistu Lustre Requirements.pdf Type: application/pdf Size: 327198 bytes Desc: Fujistu Lustre Requirements.pdf URL: From carrier at cray.com Wed Feb 16 05:43:08 2011 From: carrier at cray.com (John Carrier) Date: Tue, 15 Feb 2011 23:43:08 -0600 Subject: [Twg] 2011-02-10 meeting minutes Message-ID: Attached are the minutes from last Thursday's meeting when we reviewed Fujitsu's requirements. Dave sent our updated requirements to on Saturday. We were to have sent new requirements to the list by EOB Monday. Our next meeting is Thursday, 2/17/2011 @ 9:30a PT/12:30p ET. (Dial-in numbers are 715-726-4994 or 866-304-8294, meeting ID and password are 7012090.) The agenda for the meeting is to discuss requirement priorities and prepare for our future architecture meetings. Please review the requirements list and send your priorities _before_ the meeting. (Dave would prefer EOB Wednesday :) Send requirements and priorities to and any process concerns or questions to . Thanks, --jc -------------- next part -------------- OpenSFS Technical Working Group Meeting minutes : 02/10/2011 Concall: start : 9:30a PT, end : 11:30a PT Next meeting: Thursday, 2/17/2011 @ 9:30a PT/12:30p ET Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Deadline - requirements by 5pm EST Monday - priorities by 5pm EST attending Name Organization email ----------------- -------------- ---------------------------- Cory Spitz Cray spitzcor at cray.com John Carrier Cray carrier at cray.com Justin Miller IU jupmille at indiana.edu Steve Simms IU ssimms at indiana.edu Chris Morrone LLNL morrone2 at llnl.gov Marc Stearman LLNL marc at llnl.gov Shay Seager OpenSFS shay at opensfs.org Dave Dillow ORNL dillowda at ornl.gov Galen Shipman ORNL gshipman at ornl.gov Eric Barton Whamcloud eeb at whamcloud.com Agenda requirement discussion and prioritization Discussion Eric forwarded a document he received from Fujitsu "Fujitsu Lustre Requirements.pdf". John forwarded the doc to the reflector and we spent the remainder of the meeting discussing the items they listed. Note the meeting ran for an additional, unplanned, hour in order to complete the review of Fujitsu's requirements. In the notes below, the requirement is on one line and the discussion appears, indented, on the lines below it. We went through Fujitsu's doc page by page, but topics on each page may not have been covered in order. * 512KB block size for ldiskfs This is a backend requirement. Not sure if ldiskfs is right target, which is an implementation detail. A larger blocksize gives more performance and gets by scaling requirements, but there may be other backend solutions that do what Fujitsu really needs. * 100 PB file system size We expect file system size to scale with the memory size of the mainframe. A typical rule of thumb is to track node memory, size of normal checkpoint data, and the number of checkpoints to store during a run. LLNL expects to have 50 PB in 2011, 100 PB in 2012. ORNL has 10PB today. They bought spindles to get the bandwidth needed and then sized the capacity to meet storage requirements. Since capacity of drives doubles every 18 months, agreed that 100PB is a reasonable target for 2012. * 10K OSS nodes 2K is achievable in near term, but 10K is aggressive and will require changing derived attributes. To scale the OSS beyond 2K, Lustre will need fast health detection to avoid problems caused by distributed communications between clients and servers. We need to track scaling of OSS as another feature that depends on the health network. Nonetheless, we need to understand how many OpenSFS customers will build systems of this size. * 10K OSTs Max # of OSTs with today's architecture will be a stripe-size of 1500, with minor format changes. There will be a significant amount of work to go beyond this limit. 10K OSTs and 10K OSS implies ratio of 1 OST per OSS. Can single OST use all the bandwidth from a single OSS. Not clear what the architectural requirements are. How much bandwidth is needed? How many rails of communication are there? Consensus is that the goal should be 1K OSSs. The # of OSTs will vary to get max BW from the OSS. One OST/OSS would reduce admin cost. * 100K clients ORNL is moving to fewer, fatter nodes. BlueGene (LLNL) uses i/o forwarding to funnel node requests to fewer clients. If Fujitsu needs high node counts, they could also use I/O forwarding to mitigate the client requirements. Consensus that 64K clients in 2012 and 128K in 2014 is reasonable. * 100TB OST size This is a back-end storage requirement. Consensus that 32TB should be the near-term target and 64TB is reasonable long-term (2014). Beyond this, need new backend file system. * 100K max subdirectories Group believes the limit is not the file system itself but at the VFS layer. Lustre is constrained by the Linux architecture. * big-endian support Lustre servers will be little-endian. Big-endian support is available in Lustre 2.x for clients, but we will need to test the code to be sure it works. * varying page-sizes This topic arose during the big-endian discussion. Lustre must allow clients and servers to use different page sizes. (Expect smaller on client, larger on server.) * advisory file locking This is currently supported, though no one was sure it was fully operational. Also not sure how it was used, though one person suggested the adio drive for MPI-IO might. * server cache Read cache exists. Write cache is messy. Agreed not to require it. * arbitrary OST assignment Goal is to assign specific stripes to specific OSTs rather than heuristically. This has a potential impact on WAN operations. * 200K file creates/sec This is an aggregate requirement for creates from all nodes in the system. We looked at LLNL's 1.5 million core Sequoia system to put this in context. They expect 100K file create/s. If each rank creates its own file, then the checkpoint will need 15 seconds to create all files from all cores. Someone then asked the question of how this time compared to the time each core spends writing its file. There is 1 GB per core and 512 GB/s is the bandwidth target. At 15 secs, LLNL estimates opens will take ~4% of the write time. 200K creates/s may be a reasonable short-term goal. Some thought we should expect this to quadruple by 2014. * 100K metadata ops There is a difference between 1.5 million cores creating files and a single process running "ls -l" on the files created by the 1.5 million cores. We need to separate aggregate metadata ops from requirements for small numbers of processes. This may lead to two requirements: one for aggregate stats, one for single client. * assured disk i/o The requirement should be for zero down-time while checking the file system. We discussed separting this requirement for a Lustre-level fsck and a low-level fsck. Both systems should have online checking and scrubbing. Failover doesn't work if doing off-line fsck. * 1PB file size This is a consistent with shared-file requirements from others. * dynamic OST addition/deletion Ceph provides a model for adding storage dynamically. It uses the CRUSH algorithm to determine striping. For Lustre, we need to consider the locality of the new OSTs while enabling dynamic striping. Could create metadata to create heirarchical map of storage. With equivalent choices, get random stripe assignments; with unequivalent choices, get directed i/o. other topics mentioned during this discussion: * space-determined migration, migration can be driven by policy or command line * adaptive object layouts needs research: CRUSH is not spectacularly good or bad * replicas not RAID across OSTs, but policy-driven replication - for certain files, need to get copied somewhere else - if complete backend disaster, then have recovery mechanism * Exascale and POSIX Fujitsu's exascale numbers are really extreme and show how exascale is stretching POSIX. Instead of exposing POSIX to such extreme scales, we could explore architecture changes. Applications must still see POSIX APIs and write to file objects as expected, but without overloading the namespace held by the metadata server to show all files. An example would be the use of data sets on the MDS, where instead of exposing all of the files in the namespace, there would a single file with metadata for all of the other subordinate objects in the data set. The challenge is expressing this in terms of requirements. We agreed to explore the goal of alternate access methods and namespaces. For example, look to support multiple namespaces and consider a generalized object model vs posix interface. * MDS scaling We all know that Sun had worked on SMP scaling to improve MDS performance. If Lustre can scale on SMPs, then there will be wider range of systems to run on. If can't exploit additional cores, esp on servers, then limiting ourselves. The requirement is that we expect Lustre to run on reasonable hardware and that we expect the software system to scale and perform with underlying hardware. * patchless servers Should be required for better support for new Linux, security patches etc * scalability The Fujitsu system appears to be large enough that there should be locality of reference to the OSSs to avoid contention. Having all clients talking to all servers will cause turbulence on the interconnection network. ORNL attempted to resolve this problem with LNET fine grain routing. The routing forced clients to access the router connected to specific servers in order to avoid contention on their IB I/O network. Locality is part of the story for avoiding network congestion between clients and servers. Clients spend time checking connectivity of all servers in the file system. Localization would also reduce the number of servers mapped by the clients and avoid using loads of memory to keep connection state. Caching proxy servers will also alleviate congestion problems. Long-term, procxies will be needed to reach 1 million clients. * snapshots Replicas and snapshots will be needed to do Lustre backups. The requirement is for consistent use of file system at point in time. Next Meeting (2/17) continue discussion on - send additional requirements by 5pm EST Monday - send priorities by 5pm EST Wednesday From ssseager at gmail.com Wed Feb 16 17:41:45 2011 From: ssseager at gmail.com (Shay Seager) Date: Wed, 16 Feb 2011 09:41:45 -0800 Subject: [Twg] Document Repository - Please Review Message-ID: <7DEB0CB1-C5E0-4E1B-A806-89DCD01E14E1@gmail.com> Hello OpenSFS Members, The Communication working group reviewed many document storage applications over the past week to decide what was the best venue to securely store and work on internal documents. After comparing many choices we think that Google Docs will be our best option. We reviewed the concerned voiced in the past about using Google docs and I have our responses listed: 1-You need a google account. With any repository application you will need an account to access the files. With Google docs individuals were afraid that if they didn't have a google email address they would not be able to access the documents and internally keeping track of double e-mail addresses for everyone would be a hassle. However, we found out that you can and will need to create a google account with your corporate e-mail address. Creating an account with your corporate email will allow us to verify your membership within OpenSFS and the process of creating an account will mirror that of any of the repository options. 2- Security. We found that Google docs offers the same security options as it's more expensive competitors (ie. Zoho), where documents can be shared by the creator and only seen by the individuals which the creator shared the document with. 3 - Google Docs is not blocked by any members sites as far as we know. 4- Trackable changes. Google Docs will show you who is viewing a document when you are viewing it and will track changes within a document. 5- Price. Google Docs was the only free venue that met all of our needs. Using a free account for document storage allows more of OpenSFS's funds to go back into OpenSFS Lustre Projects. The price range for other repositories started at $1000/year-$2000/year. We are looking for member feedback on the decision and storage options. Please e-mail me with any comments/questions/concerns. We will implement a document repository system by February 23rd, without feedback we will move ahead with Google Documents. Warmest Regards, Shay Seager and the Communication Working Group - - - Shay Seager Open SFS Secretary www.opensfs.org (925) 290-7641 shay at opensfs.org From hamilton5 at llnl.gov Tue Feb 22 18:58:38 2011 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Tue, 22 Feb 2011 10:58:38 -0800 Subject: [Twg] FW: Lustre 2.1 Release - DRAFT Message-ID: Hi all, Below is the draft announcement for the Lustre 2.1 Release. Please send comments to Peter or myself. We can also discuss this on today's call at 1:30pm PST: 866-914-3976 (925-424-8105)x534986# Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA 94551-9900 E-Mail: pgh at llnl.gov Phone: 925-423-1332 Fax: 925-423-8719 ___________________________________ From: Peter Jones [mailto:pjones at whamcloud.com] Sent: Tuesday, February 22, 2011 10:41 AM To: Hamilton, Pam Subject: Lustre 2.1 Release - DRAFT Hi Pam Here is the draft of the message that I plan to send to lustre-discuss first thing tomorrow. Please let me know ahead of time if there is anything below that OpenSFS would like to change. We can discuss on the call this afternoon if you like Regards Peter -------------------------------------------------------------------------------------------------------------------- Hi there There has been much discussion within the Lustre community about the future of the Lustre 2.x codeline with the following outcome. Roles -I have taken on the role of Release Manager for the Lustre 2.1 release and Oleg Drokin (green at whamcloud.com) will be the Technical Lead for this release. Issue Tracking -Issues relating to this release will be tracked in Whamcloud's JIRA system - http://jira.whamcloud.com . Signup is open and free. -To see the present list of blockers, please use the filter "Lustre 2.1 Blockers". This can be conveniently accessed by selecting Manage Filters and then Popular. Source Control -The code for the release will be made from Whamcloud's git instance - http://git.whamcloud.com/ -Patches contributed by engineers from third party organizations will be according to arrangement similar to the kernel (see http://wiki.whamcloud.com/display/PUB/Submitting+Changes for details). The outcome will be that no single organization will own the copyright to this release Testing -The latest build can be downloaded from the http://build.whamcloud.com/ -Testing results from both Whamcloud and third party organizations will be stored in Maloo, the Whamcloud test database - http://maloo.whamcloud.com. See http://wiki.whamcloud.com/display/PUB/Using+Maloo for details on how to use Maloo either to view progress or to upload your own testing results. Weekly Call -A weekly status call will take place Tuesday at 1:30pm PT. This call is open to any interested parties. 866-914-3976 534986# This was considered the most expedient plan for the Lustre 2.1 release, but a different approach may be taken for ongoing Lustre 2.x releases. This is still under consideration within the Lustre community. The Lustre community organizations - EOFS, HPCFS, and OpenSFS - have all expressed support for these plans and we look forward to collaborating with the community for this release. A Lustre 2.1 Google group has been setup as a forum to discuss this release - http://groups.google.com/group/lustre-21. Please feel free to signup for this mailing list whether you are interested in collaborating in this release or just observing the progress. Regards Peter NB\ Lustre is a trademark of the Oracle Corporation -------------- next part -------------- An HTML attachment was scrubbed... URL: From morrone2 at llnl.gov Tue Feb 22 22:25:28 2011 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Tue, 22 Feb 2011 14:25:28 -0800 Subject: [Twg] [Rpwg] FW: Lustre 2.1 Release - DRAFT In-Reply-To: References: Message-ID: <4D6437D8.9010104@llnl.gov> FYI, If you are like me and wondering how to subscribe to a google group from a non-google account (e.g. your work email), here is the answer: http://groups.google.com/support/bin/answer.py?hl=en&answer=46606 On 02/22/2011 10:58 AM, Hamilton, Pam wrote: > Hi all, > > Below is the draft announcement for the Lustre 2.1 Release. Please send comments to Peter or myself. We can also discuss this on today’s call at 1:30pm PST: > 866-914-3976 (925-424-8105)x534986# > > Regards, > Pam > ___________________________________ > Pam Hamilton > Lawrence Livermore National Lab > P.O. Box 808, L-556 > Livermore, CA 94551-9900 > E-Mail: pgh at llnl.gov > Phone: 925-423-1332 Fax: 925-423-8719 > ___________________________________ > > > From: Peter Jones [mailto:pjones at whamcloud.com] > Sent: Tuesday, February 22, 2011 10:41 AM > To: Hamilton, Pam > Subject: Lustre 2.1 Release - DRAFT > > > Hi Pam > > > > Here is the draft of the message that I plan to send to lustre-discuss first thing tomorrow. Please let me know ahead of time if there is anything below that OpenSFS would like to change. We can discuss on the call this afternoon if you like > > > > Regards > > > > Peter > > > > -------------------------------------------------------------------------------------------------------------------- > > > > Hi there > > > > There has been much discussion within the Lustre community about the future of the Lustre 2.x codeline with the following outcome. > > > > Roles > > -I have taken on the role of Release Manager for the Lustre 2.1 release and Oleg Drokin (green at whamcloud.com) will be the Technical Lead for this release. > > > > Issue Tracking > > -Issues relating to this release will be tracked in Whamcloud's JIRA system - http://jira.whamcloud.com . Signup is open and free. > > -To see the present list of blockers, please use the filter "Lustre 2.1 Blockers". This can be conveniently accessed by selecting Manage Filters and then Popular. > > > > Source Control > > -The code for the release will be made from Whamcloud's git instance - http://git.whamcloud.com/ > > -Patches contributed by engineers from third party organizations will be according to arrangement similar to the kernel (see http://wiki.whamcloud.com/display/PUB/Submitting+Changes > > for details). The outcome will be that no single organization will own the copyright to this release > > > > Testing > > -The latest build can be downloaded from the http://build.whamcloud.com/ > > -Testing results from both Whamcloud and third party organizations will be stored in Maloo, the Whamcloud test database - http://maloo.whamcloud.com. See http://wiki.whamcloud.com/display/PUB/Using+Maloo for details on how to use Maloo either to view progress or to upload your own testing results. > > > > Weekly Call > > -A weekly status call will take place Tuesday at 1:30pm PT. This call is open to any interested parties. 866-914-3976 534986# > > > > This was considered the most expedient plan for the Lustre 2.1 release, but a different approach may be taken for ongoing Lustre 2.x releases. This is still under consideration within the Lustre community. > > > > The Lustre community organizations - EOFS, HPCFS, and OpenSFS - have all expressed support for these plans and we look forward to collaborating with the community for this release. > > > > A Lustre 2.1 Google group has been setup as a forum to discuss this release - http://groups.google.com/group/lustre-21. Please feel free to signup for this mailing list whether you are interested in collaborating in this release or just observing the progress. > > > > Regards > > > > Peter > > > > > > NB\ Lustre is a trademark of the Oracle Corporation >