[Twg] meeting minutes for 2011-02-17, 2011-02-24, 2011-03-03

John Carrier carrier at cray.com
Thu Mar 10 09:07:36 UTC 2011


Attached are meeting minutes for the 2011-02-17 and 2011-02-04 meetings.  There are no minutes for 2011-03-03 since that meeting was cancelled.  Please let me know if you have any corrections to make and I'll repost.

Our next meeting is Thursday, 3/10/2011 @ 9:30a PT/12:30p ET. (Dial-in numbers are 715-726-4994 or 866-304-8294, meeting ID and password are 7012090.)  Our agenda is to prioritize the requirements.  Dave posted the last update to the requirements on 2/12 to discuss at lists.opensfs.org.  I haven't seen an update since then.  I also haven't seen any further priority lists since we started discussing them at the 2/17 meeting.

The TWG needs to submit the prioritized requirements list to the board. DaveD is slotted to present these at LUG.  Please bring your top 5 items to this meeting (better yet, send them to the list before hand!).  We need to move on the list in order to identify the requirements we want to use in the RFPs.  

My apologies for not getting the minutes distributed sooner.  It's been a very hectic few weeks. 

Thanks,

--jc



-------------- next part --------------
OpenSFS Technical Working Group
Meeting minutes : 02/24/2011

Concall: start : 9:30a PT, adjourn : 9:48a PT

Attending
   Name              Organization   email
   ----------------- -------------- ----------------------------
   Eric Monserrat    Bull           eric.monserrat at bull.net
   Diego Moreno      Bull           diego.moreno at bull.net
   Pascal Barbolosi  Bull           pascal.barbolosi at bull.net
   Cory Spitz        Cray           spitzcor at cray.com
   John Carrier      Cray           carrier at cray.com
   Justin Miller     IU             jupmille at indiana.edu                            
   Steve Simms       IU             ssimms at indiana.edu                            
   Shay Seager       OpenSFS        shay at opensfs.org
   Andreas Dilger    Whamcloud      adilger at whamcloud.com
   Eric Barton       Whamcloud      eeb at whamcloud.com

Agenda
   priorities

Meeting notes
   DaveD was off on vacation.  No one had posted their requirements to
   the lists so there was nothing to reconcile during the meeting.  
   
   We were, however, joined by three representatives from Bull. [My
   apologies if I misidentified any of these people or misspelled their
   names. Please send corrections. --jc] They noted that European OFS
   is gathering requirements from their membership.  These are still
   several weeks out.  Their focus will be performance, redundancy,
   recovery, backup/restore, management, and security.  The priority
   will set the timeframe for the requirements. They defered discussing
   further until the European OFS publishes the results.

   Without priority lists, we adjourned the meeting early.
-------------- next part --------------
OpenSFS Technical Working Group
Meeting minutes : 02/17/2011

Concall: start : 9:30a PT 

Attending

   Name              Organization   email
   ----------------- -------------- ----------------------------
   Cory Spitz        Cray           spitzcor at cray.com
   John Carrier      Cray           carrier at cray.com
   Steve Simms       IU             ssimms at indiana.edu                            
   Chris Morrone     LLNL           morrone2 at llnl.gov
   Shay Seager       OpenSFS        shay at opensfs.org
   Dave Dillow       ORNL           dillowda at ornl.gov
   Sarp Oral         ORNL           oralhs at ornl.gov
   Andreas Dilger    Whamcloud      adilger at whamcloud.com
   Eric Barton       Whamcloud      eeb at whamcloud.com

Agenda
   requirements
   priorities

Requirements
   Dave posted requirements update on Saturday 2/12.  He needs more info
   on 2012/2014 projections for:
      single client creates/sec
      single client stats/sec 
      directory listings/sec 

   Eric asked about implied requirements that have not been explicitly
   stated, such as, interoperability and compatability between lustre
   versions.  

   Maintaining interop between generations of Lustre costs money for
   development and testing.  We understand that major version changes
   will make interoperability difficult, but expect that 2nd digit
   changes will not be as radical and should be interoperable with
   previous versions.

   We discussed doing rolling upgrades to transition to Lustre 2.0. The
   proposed scenario is to unmount the servers, upgrade the servers,
   reboot the servers, then mount the clients.  This means there will be
   a service disruption to upgrade.

   Regardless of the upgrade path, the expectation is that there will be
   forward and backward compatibility of one revision. 

         >> ACTION: Eric to provide text for interoperability and
                    compatibility.

Priorities
   Sequoia is a priority for LLNL.  They will be deploying Lustre on ZFS
   and the OSD work started at Sun needs to be completed.  Chris
   proposed the following for LLNL:
      - complete OSD
      - scalable fault management
      - LNET channel bonding to provide more than one pipe from one box.
         >> ACTION: John to provide text for channel bonding
                    requirements
      - balancing storage usage and stripe allocation
         >> ACTION: Chris to provide text for rebalancing OSTs

   Steve Simms proposed the following priorites for IndianaU:
      - improved back-end
      - better fault management
      - metadata performance for WAN users
      
   IU wants a new back-end to replace ldiskfs that offers a low time for
   offline integrity checking and low impact for online integrity
   checking.  Lustre RAID might be good with replicated metadata and
   triple RAID.

   We then discussed whether RAID6 was dead and looked at integrity
   checking options.  There are three items being discussed to improve
   data integrity: end-to-end data integrity, online checking, and
   triple parity.  The requirement is that we want zero interruption in
   the face of failures.  Lustre-level replication needs to provide
   better availability than is possibe from the backend hardware. Lustre
   must handle failures of the backend storage and should replicate on
   other HW to improve availability.  This requirement could also fall
   under scalable fault management.

   We ended up with two new requirements from this discussion:
         - resiliency of backend storage
         - availability of Lustre in face failures

   Fault management means faster failovers.  Failover needs to be
   quicker and there should be better diags to show what is going on
   during fault recovery.  Currently, failovers at IU take >20mins.

   IU liked size on MDS, which made the 2nd request much faster
   than the first.  We discussed what making the MDS faster really
   meant and agreed that the requirement is to improve "ls -l"
   responsiveness and make "du" faster finding space in the
   directory.
   
Next meeting (2/24) 

   Everyone needs to put forward their top 3-5 requirments from the
   list.  We don't need to rank all requirements, just prioritize our
   top ones.

   Send these priorities to the list and we'll organize the combined
   lists at the next meeting.



More information about the Twg mailing list