From carrier at cray.com Wed Dec 1 10:03:08 2010 From: carrier at cray.com (John Carrier) Date: Wed, 1 Dec 2010 04:03:08 -0600 Subject: [Twg] 2010-11-18 meeting minutes Message-ID: My apologies for the delay submitting these notes to the group. (I returned from New Orleans and was caught up in family vacation plans.) Please send me any corrections or additions that you have. I will post updated minutes to the website. A reminder that our next meeting is this Thursday (12/2) at 9:30 PT. Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Eric will be presenting his roadmap slides: http://www.opensfs.org/wp-content/uploads/2010/11/Whamcloud-community-development-slides.pdf I will have a whitepaper outline that we can also review, if there is time. Let me know if you have any other items you would like added to the agenda. We never discussed meeting duration. If we want to restrict ourselves to biweekly meetings, we may want to use 90 minutes. Feedback is welcome. Thanks, --jc -------------- next part -------------- OpenSFS Technical Working Group Meeting minutes : 11/18/2010 Location: Broadmoor Room at Ritz Carlton (SC'10) Time: 3:30-5:00 CT Attending : Name Organization email ----------------- -------------- ---------------------------- John Carrier Cray carrier at cray.com Kit Westeneat DDN kwesteneat at ddn.com Atul Vidwansa DDN avidwansa at ddn.com Joshua Walgenbach Indiana Univ. jjw at indiana.edu Marc Stearman LLNL marc at llnl.gov Chris Moronne LLNL morrone2 at llnl.gov Frank Indiviglio NOAA frank.indiviglio at noaa.gov Bryon Neitzel Oracle bryon.neitzel at oracle.com Sarp Oral ORNL oralhs at ornl.gov Galen Shipman ORNL ghsipman at ornl.gov Feyi Wang ORNL fwang2 at ornl.gov Dave Dillow ORNL dillowda at ornl.gov Ercan Kamber RAID, Inc. ercan_kamber at raidinc.com Robert Read Whamcloud rread at whamcloud.com Eric Barton Whamcloud eeb at whamcloud.com (Kit and Atul were on the phone) Agenda : - introductions - wg charter and organization - whitepaper (due 11/30) - schedule for weekly concall Notes : - charter & roadmap Galen introduced the workgroup charter as proposed by the board: - gather requirements - develop a feature roadmap - generate RFPs - work with contractors to manage deliverables The board expects each workgroup to refine its charter through a whitepaper to be delivered by the end of the month. The TWG must describe its rules for creating the roadmap and define the development model for issuing RFPs, reviewing proposals, and evaluating deliverables. The first step is developing the roadmap since this will dictate how OpenSFS will begin spending its development dollars. We discussed the need for defining a strategy based on short-, medium-, and long-term goals. The group needs to define the target architecture, not how the engineering is done, so that long-term requirements can lead to coherent development toward those goals. Marc suggested quarterly review meetings to ensure that we are keeping on track and allow us to evaluate our progress and reset priorities if necessary. We discussed the importance of defining a mission statement for project. Although exascale is on the horizon, it is important to focus on how to make Lustre scale or, in other words, keep up with Moore's law. This implies improving Lustre resiliency as failures will be more common at scale and ease administration of the file system and hardware as the system grows. This was summarized as a statement to scale, deploy, and maintain Lustre. To start the roadmap discussion, Eric has prepared a set of slides, which we will review at the next meeting: http://www.opensfs.org/wp-content/uploads/2010/11/Whamcloud-community-development-slides.pdf We need to look outside OpenSFS and include the community when scoping our requirements. We should look at all the good ideas and sort between them. Then we need to identify the near-term goals necessary to move us toward the long-term requirements. - feature development Assuming we have a roadmap and have identified distinct work projects to be completed, we then discussed how we should proceed with completing feature development. Efforts of this working group need to be kept open. The community should know our roadmap and what features we are pursuing. We need to communicate this work to Oracle so that we can be sure the completed features have a place to land without conflict in Oracle's canonical Lustre tree. Similarly, we need to canvass the members and community (including Oracle) to be sure our development efforts are not duplicating development already planned or in progress. We then discussed strategies for issuing RFPs. Galen suggested that we could initially issue an RFI to determine the cost, which the TWG would use to evaluate resource loading. We would then issue RFPs for features we could afford. We agreed that it would be easier to issue separate smaller RFPs rather than one large one. The RFPs would require that patches be made as easily inspectable changes and that sets of patches could be easily testable. The RFP should require the contractor to provide the release collateral necessary to have the feature accepted into the Oracle Lustre tree. This includes design documents, test plans, test programs, test inspections, and test results. We discussed the importance of keeping the master stable. The process is to patch, test, then land. With more groups doing development, it will be more critical to do more testing. OpenSFS needs to be responsible for gatekeeping of its development branch. We then discussed with Bryon how OpenSFS should interface with Oracle to land patches. We agreed that requiring Oracle to do all inspections and testing would be onerous. Instead, we discussed the kernel.org model where Linus has trusted lieutenants who can accept patches. Oracle would need to find outside resources who they trusted to make these recommendations and to identify an Oracle person (eg, Andreas) to accept our landing collateral. - next meetings We are reserving Thursdays at 9:30am PT on our calendars for future TWG meetings. The team prefers biweekly meetings. The group wants, however, to avoid standing meetings that fill calendars without specific goals. The WG chairs will issue a call for agenda items at the beginning of the week. If there are no items to discuss, the meeting will be cancelled. Our next meeting is Thursday 12/2. Call one of these numbers: 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Agenda for this meeting is - discuss Eric's roadmap slides - discuss the whitepaper (John will send an outline Weds PM) Action Items: * Post Eric's roadmap slides to the TWG website: complete http://www.opensfs.org/wp-content/uploads/2010/11/Whamcloud-community-development-slides.pdf * TWG email reflector should be globally accessible: complete All OpenSFS workgroup email archives are available through the OpenSFS list interface (http://lists.opensfs.org). Only TWG members, however, can submit email to the reflector. The TWG archive is here: http://lists.opensfs.org/pipermail/twg-opensfs.org/ * Post TWG information to the lustre-community reflector : TBD John will post a meeting notice later on 12/1. From seager1 at llnl.gov Wed Dec 1 21:16:38 2010 From: seager1 at llnl.gov (Seager, Mark K.) Date: Wed, 1 Dec 2010 13:16:38 -0800 Subject: [Twg] X-Pollination Telecon Notes Message-ID: <6DAF24F9BB747E47A4BAE7A6F5B80C7DF58135D3BD@NSPEXMBX-A.the-lab.llnl.gov> Here are the notes from the subject telecon yesterday. Thanks to Shay for the great notes. By the way, many comments in these notes are not attributed to a specific person because Shay can tell who is speaking. Please don't take it personally! It would be helpful if people started their comments with their name until we all become more familiar with each other and can correlate voice with name. Next call is 12/14/2010 at 9AM PST with 517-308-1709x5282795 Call in number. Regards, ++Mark ________________________________ Mark K. Seager Advanced Technology| Google Voice LLNL | The more original a discovery, the more 925-456-4214 POBOX 808, L-554 | obvious it seems afterwards. seager [at] 7000 East Ave. | -- Arthur Koestler llnl.gov Livermore, CA 94551| ________________________________ -------------- next part -------------- An HTML attachment was scrubbed... URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: 11-30-10 Working X-Pollination Group Telecon.doc Type: application/msword Size: 50176 bytes Desc: 11-30-10 Working X-Pollination Group Telecon.doc URL: From carrier at cray.com Thu Dec 2 07:48:16 2010 From: carrier at cray.com (John Carrier) Date: Thu, 2 Dec 2010 01:48:16 -0600 Subject: [Twg] draft whitepaper Message-ID: Attached is a draft whitepaper describing the function of the TWG. This is based on our one meeting at SC two weeks ago, so I fully expect it to change as more people get involved. Please feel free to add or edit the text. If there is time, I hope we can discuss the doc during the 12/2 call. Thanks, --jc -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS TWG white paper_ver1.docx Type: application/vnd.openxmlformats-officedocument.wordprocessingml.document Size: 17430 bytes Desc: OpenSFS TWG white paper_ver1.docx URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS TWG white paper_ver1.pdf Type: application/pdf Size: 185892 bytes Desc: OpenSFS TWG white paper_ver1.pdf URL: From carrier at cray.com Thu Dec 2 17:15:49 2010 From: carrier at cray.com (John Carrier) Date: Thu, 2 Dec 2010 11:15:49 -0600 Subject: [Twg] email to [lustre-community] Message-ID: Team, I had an action from the last meeting to make public the activities of our group. Toward that end, OpenSFS has created a TWG webpage and the archive of our email reflectors is open to the public for viewing. To complete the task, I propose sending the following email to the [lustre-community]. I would like to get your approval before doing so. Thanks, --jc 8<-------------------------------- The OpenSFS technical working group (TWG) is creating a roadmap to drive development of scale-out features for Lustre for HPC. The TWG website (http://www.opensfs.org/?page_id=88) contains meeting announcements and links to group documents. Its email archive is available at http://lists.opensfs.org/pipermail/twg-opensfs.org/ Please let us know if you have any questions. John Carrier (Cray) and Dave Dillow (ORNL) -------------------------------->8 From ekamber at raidinc.com Thu Dec 2 17:29:23 2010 From: ekamber at raidinc.com (Ercan Kamber) Date: Thu, 2 Dec 2010 09:29:23 -0800 Subject: [Twg] 2010-11-18 meeting minutes In-Reply-To: References: Message-ID: Hi guys, I just dialed 866-304-8294 and punched the meeting ID 7012090 and the system told me it could not recognize the ID. Is this ID and password correct? Best, Ercan Kamber, Ph.D. Vice President of Engineering RAID, Incorporated www.raidinc.com ekamber at raidinc.com office : (978) 683 6444 ext 161 lab : (978) 683 6444 ext 163 Mobile: (774) 285 2575 -----Original Message----- From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of John Carrier Sent: Wednesday, December 01, 2010 5:03 AM To: twg at lists.opensfs.org Subject: [Twg] 2010-11-18 meeting minutes My apologies for the delay submitting these notes to the group. (I returned from New Orleans and was caught up in family vacation plans.) Please send me any corrections or additions that you have. I will post updated minutes to the website. A reminder that our next meeting is this Thursday (12/2) at 9:30 PT. Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Eric will be presenting his roadmap slides: http://www.opensfs.org/wp-content/uploads/2010/11/Whamcloud-community-development-slides.pdf I will have a whitepaper outline that we can also review, if there is time. Let me know if you have any other items you would like added to the agenda. We never discussed meeting duration. If we want to restrict ourselves to biweekly meetings, we may want to use 90 minutes. Feedback is welcome. Thanks, --jc From gshipman at ornl.gov Thu Dec 2 17:30:26 2010 From: gshipman at ornl.gov (Shipman, Galen M.) Date: Thu, 02 Dec 2010 12:30:26 -0500 Subject: [Twg] email to [lustre-community] In-Reply-To: References: Message-ID: <0EE78390-02CE-4E47-BB89-FF7CAE54E99A@ornl.gov> See small change below indicating what we are targeting for scale-out (performance, resiliency, and reliability). Thanks, Galen On Dec 2, 2010, at 12:15 PM, John Carrier wrote: > Team, > > I had an action from the last meeting to make public the activities of our group. Toward that end, OpenSFS has created a TWG webpage and the archive of our email reflectors is open to the public for viewing. > > To complete the task, I propose sending the following email to the [lustre-community]. I would like to get your approval before doing so. > > Thanks, > > --jc > > 8<-------------------------------- > > The OpenSFS technical working group (TWG) is creating a roadmap to drive development of scale-out features (performance, resiliency, and reliability) for Lustre for HPC. The TWG website (http://www.opensfs.org/?page_id=88) contains meeting announcements and links to group documents. Its email archive is available at http://lists.opensfs.org/pipermail/twg-opensfs.org/ > > Please let us know if you have any questions. > > John Carrier (Cray) and Dave Dillow (ORNL) > > -------------------------------->8 > _______________________________________________ > twg mailing list > twg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/twg-opensfs.org From carrier at cray.com Thu Dec 2 17:32:22 2010 From: carrier at cray.com (John Carrier) Date: Thu, 2 Dec 2010 11:32:22 -0600 Subject: [Twg] 2010-11-18 meeting minutes In-Reply-To: References: Message-ID: I'll reset and send new email -- give me a couple minutes -----Original Message----- From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of Ercan Kamber Sent: Thursday, December 02, 2010 9:29 AM To: twg at lists.opensfs.org Subject: Re: [Twg] 2010-11-18 meeting minutes Hi guys, I just dialed 866-304-8294 and punched the meeting ID 7012090 and the system told me it could not recognize the ID. Is this ID and password correct? Best, Ercan Kamber, Ph.D. Vice President of Engineering RAID, Incorporated www.raidinc.com ekamber at raidinc.com office : (978) 683 6444 ext 161 lab : (978) 683 6444 ext 163 Mobile: (774) 285 2575 -----Original Message----- From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of John Carrier Sent: Wednesday, December 01, 2010 5:03 AM To: twg at lists.opensfs.org Subject: [Twg] 2010-11-18 meeting minutes My apologies for the delay submitting these notes to the group. (I returned from New Orleans and was caught up in family vacation plans.) Please send me any corrections or additions that you have. I will post updated minutes to the website. A reminder that our next meeting is this Thursday (12/2) at 9:30 PT. Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Eric will be presenting his roadmap slides: http://www.opensfs.org/wp-content/uploads/2010/11/Whamcloud-community-development-slides.pdf I will have a whitepaper outline that we can also review, if there is time. Let me know if you have any other items you would like added to the agenda. We never discussed meeting duration. If we want to restrict ourselves to biweekly meetings, we may want to use 90 minutes. Feedback is welcome. Thanks, --jc _______________________________________________ twg mailing list twg at lists.opensfs.org http://lists.opensfs.org/listinfo.cgi/twg-opensfs.org From andreas.dilger at oracle.com Thu Dec 2 17:33:43 2010 From: andreas.dilger at oracle.com (Andreas Dilger) Date: Thu, 2 Dec 2010 10:33:43 -0700 Subject: [Twg] 2010-11-18 meeting minutes In-Reply-To: References: Message-ID: <30784D47-1A27-4EC2-9D30-B2A1EC545B81@oracle.com> On 2010-12-02, at 10:29, Ercan Kamber wrote: > I just dialed 866-304-8294 and punched the meeting ID 7012090 > and the system told me it could not recognize the ID. > > Is this ID and password correct? I tried the same, and the meeting ID was not working... > Ercan Kamber, Ph.D. > Vice President of Engineering > RAID, Incorporated > www.raidinc.com > > ekamber at raidinc.com > office : (978) 683 6444 ext 161 > lab : (978) 683 6444 ext 163 > Mobile: (774) 285 2575 > > -----Original Message----- > From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of John Carrier > Sent: Wednesday, December 01, 2010 5:03 AM > To: twg at lists.opensfs.org > Subject: [Twg] 2010-11-18 meeting minutes > > My apologies for the delay submitting these notes to the group. (I returned from New Orleans and was caught up in family vacation plans.) Please send me any corrections or additions that you have. I will post updated minutes to the website. > > A reminder that our next meeting is this Thursday (12/2) at 9:30 PT. Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Eric will be presenting his roadmap slides: > http://www.opensfs.org/wp-content/uploads/2010/11/Whamcloud-community-development-slides.pdf > > I will have a whitepaper outline that we can also review, if there is time. Let me know if you have any other items you would like added to the agenda. > > We never discussed meeting duration. If we want to restrict ourselves to biweekly meetings, we may want to use 90 minutes. > > Feedback is welcome. Thanks, > > --jc > _______________________________________________ > twg mailing list > twg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/twg-opensfs.org Cheers, Andreas -- Andreas Dilger Lustre Technical Lead Oracle Corporation Canada Inc. From carrier at cray.com Thu Dec 2 17:35:13 2010 From: carrier at cray.com (John Carrier) Date: Thu, 2 Dec 2010 11:35:13 -0600 Subject: [Twg] 2010-11-18 meeting minutes In-Reply-To: <30784D47-1A27-4EC2-9D30-B2A1EC545B81@oracle.com> References: <30784D47-1A27-4EC2-9D30-B2A1EC545B81@oracle.com> Message-ID: Please retry. I just reset the meeting. It is working now. --jc -----Original Message----- From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of Andreas Dilger Sent: Thursday, December 02, 2010 9:34 AM To: Ercan Kamber Cc: twg at lists.opensfs.org Subject: Re: [Twg] 2010-11-18 meeting minutes On 2010-12-02, at 10:29, Ercan Kamber wrote: > I just dialed 866-304-8294 and punched the meeting ID 7012090 > and the system told me it could not recognize the ID. > > Is this ID and password correct? I tried the same, and the meeting ID was not working... > Ercan Kamber, Ph.D. > Vice President of Engineering > RAID, Incorporated > www.raidinc.com > > ekamber at raidinc.com > office : (978) 683 6444 ext 161 > lab : (978) 683 6444 ext 163 > Mobile: (774) 285 2575 > > -----Original Message----- > From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of John Carrier > Sent: Wednesday, December 01, 2010 5:03 AM > To: twg at lists.opensfs.org > Subject: [Twg] 2010-11-18 meeting minutes > > My apologies for the delay submitting these notes to the group. (I returned from New Orleans and was caught up in family vacation plans.) Please send me any corrections or additions that you have. I will post updated minutes to the website. > > A reminder that our next meeting is this Thursday (12/2) at 9:30 PT. Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Eric will be presenting his roadmap slides: > http://www.opensfs.org/wp-content/uploads/2010/11/Whamcloud-community-development-slides.pdf > > I will have a whitepaper outline that we can also review, if there is time. Let me know if you have any other items you would like added to the agenda. > > We never discussed meeting duration. If we want to restrict ourselves to biweekly meetings, we may want to use 90 minutes. > > Feedback is welcome. Thanks, > > --jc > _______________________________________________ > twg mailing list > twg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/twg-opensfs.org Cheers, Andreas -- Andreas Dilger Lustre Technical Lead Oracle Corporation Canada Inc. _______________________________________________ twg mailing list twg at lists.opensfs.org http://lists.opensfs.org/listinfo.cgi/twg-opensfs.org From carrier at cray.com Fri Dec 3 06:42:05 2010 From: carrier at cray.com (John Carrier) Date: Fri, 3 Dec 2010 00:42:05 -0600 Subject: [Twg] reminder: meeting rescheduled for 9:30PT 12/3 Message-ID: We decided during our brief meeting today to reschedule our review of Eric's roadmap slides for Friday 12/3 at 9:30PT. Concall information for this meeting is the same: 715-726-4994 or 866-304-8294, with meeting ID and password are 7012090. (I just checked the schedule and I got the start time right this time.) Eric's slides are still linked from the twg webpage: http://www.opensfs.org/?page_id=88 Thanks, --jc From carrier at cray.com Fri Dec 3 17:37:28 2010 From: carrier at cray.com (John Carrier) Date: Fri, 3 Dec 2010 11:37:28 -0600 Subject: [Twg] Whamcloud's slides (try #2) Message-ID: -------------- next part -------------- A non-text attachment was scrubbed... Name: Lustre development roadmap.pdf Type: application/pdf Size: 181785 bytes Desc: Lustre development roadmap.pdf URL: From eeb at whamcloud.com Fri Dec 3 17:36:07 2010 From: eeb at whamcloud.com (Eric Barton) Date: Fri, 3 Dec 2010 17:36:07 -0000 Subject: [Twg] Lustre development slides Message-ID: <01b901cb9310$91f89ba0$b5e9d2e0$@com> Slides! -------------- next part -------------- A non-text attachment was scrubbed... Name: Lustre development roadmap.pdf Type: application/pdf Size: 181988 bytes Desc: not available URL: From hamilton5 at llnl.gov Tue Dec 7 04:30:33 2010 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Mon, 6 Dec 2010 20:30:33 -0800 Subject: [Twg] OpenSFS Support Working Group draft white paper - Round 2 Message-ID: Hi all, Attached is another draft of the OpenSFS Support Working Group white paper. During support model discussions, the SWG felt we were lacking a good understanding of the high level requirements for OpenSFS and any release distributed by the organization. Without this understanding, we are struggling with developing the support model. Therefore, this draft of includes a section with what we see as high level requirements for OpenSFS. I'm sending to the execs and other working groups to get a discussion started and hopefully we can reach a consensus regarding the requirements. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA 94551-9900 E-Mail: pgh at llnl.gov Phone: 925-423-1332 Fax: 925-423-8719 ___________________________________ -------------- next part -------------- An HTML attachment was scrubbed... URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS SWG white paper_ver3.docx Type: application/vnd.openxmlformats-officedocument.wordprocessingml.document Size: 23051 bytes Desc: OpenSFS SWG white paper_ver3.docx URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS SWG white paper_ver3.pdf Type: application/pdf Size: 307232 bytes Desc: OpenSFS SWG white paper_ver3.pdf URL: From gshipman at ornl.gov Tue Dec 7 16:54:42 2010 From: gshipman at ornl.gov (Shipman, Galen M.) Date: Tue, 07 Dec 2010 11:54:42 -0500 Subject: [Twg] OpenSFS community meeting today Message-ID: <61720761-FE45-4134-A792-9235E69C691D@ornl.gov> The OpenSFS community meeting is scheduled today at 12:00 PM ET. 517-308-1709 Participant code: 5282795 please do join, we would like to get updates from the working group leads. Thanks, Galen From carrier at cray.com Tue Dec 7 17:46:06 2010 From: carrier at cray.com (John Carrier) Date: Tue, 7 Dec 2010 11:46:06 -0600 Subject: [Twg] OpenSFS Support Working Group draft white paper - Round 2 In-Reply-To: References: Message-ID: Hi Pam, Your description of release mentions Redhat only. I think this is appropriate for the servers, but many members are using other distributions on their cluster nodes. What are the support plans for the Lustre clients for the OpenSFS distribution? Oracle has stated that their Linux support is for OEL and other RHEL derivatives. Clients will continue to be supported on all current distributions (SLES, RHEL, etc). Thanks, --jc From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of Hamilton, Pam Sent: Monday, December 06, 2010 8:31 PM To: Execs at lists.opensfs.org; twg at lists.opensfs.org; rpwg at lists.opensfs.org; swg at lists.opensfs.org; comswg at lists.opensfs.org Subject: [Twg] OpenSFS Support Working Group draft white paper - Round 2 Hi all, Attached is another draft of the OpenSFS Support Working Group white paper. During support model discussions, the SWG felt we were lacking a good understanding of the high level requirements for OpenSFS and any release distributed by the organization. Without this understanding, we are struggling with developing the support model. Therefore, this draft of includes a section with what we see as high level requirements for OpenSFS. I'm sending to the execs and other working groups to get a discussion started and hopefully we can reach a consensus regarding the requirements. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA 94551-9900 E-Mail: pgh at llnl.gov Phone: 925-423-1332 Fax: 925-423-8719 ___________________________________ -------------- next part -------------- An HTML attachment was scrubbed... URL: From carrier at cray.com Tue Dec 7 17:50:19 2010 From: carrier at cray.com (John Carrier) Date: Tue, 7 Dec 2010 11:50:19 -0600 Subject: [Twg] draft whitepaper In-Reply-To: References: Message-ID: Team, Please take a look at the one-page whitepaper I sent last week. I have received just one comment so far. What did I miss? What did I get wrong? The board expects the TWG to complete this task ASAP. Thanks, --jc -----Original Message----- From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of John Carrier Sent: Wednesday, December 01, 2010 11:48 PM To: twg at lists.opensfs.org Subject: [Twg] draft whitepaper Attached is a draft whitepaper describing the function of the TWG. This is based on our one meeting at SC two weeks ago, so I fully expect it to change as more people get involved. Please feel free to add or edit the text. If there is time, I hope we can discuss the doc during the 12/2 call. Thanks, --jc From kwestneat at ddn.com Tue Dec 7 18:31:21 2010 From: kwestneat at ddn.com (Kit Westneat) Date: Tue, 7 Dec 2010 13:31:21 -0500 Subject: [Twg] draft whitepaper In-Reply-To: References: Message-ID: <4CFE7D79.80605@ddn.com> Hi John, The support wg white paper mentions providing the twg with support metrics in order to help plan the evolution of OpenSFS, but the twg white paper currently doesn't discuss how support issues might affect the twg planning. Is this understood as part of surveying the OpenSFS membership? If the OpenSFS members decide to invest in a project to improve the reliability of Lustre, or even a subsystem of Lustre like quotas, would this be done through the twg or swg? Thanks, Kit On 12/07/2010 12:50 PM, John Carrier wrote: > Team, > > Please take a look at the one-page whitepaper I sent last week. I have received just one comment so far. What did I miss? What did I get wrong? > > The board expects the TWG to complete this task ASAP. > > Thanks, > > --jc > > > -----Original Message----- > From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of John Carrier > Sent: Wednesday, December 01, 2010 11:48 PM > To: twg at lists.opensfs.org > Subject: [Twg] draft whitepaper > > Attached is a draft whitepaper describing the function of the TWG. This is based on our one meeting at SC two weeks ago, so I fully expect it to change as more people get involved. Please feel free to add or edit the text. If there is time, I hope we can discuss the doc during the 12/2 call. > > Thanks, > > --jc > _______________________________________________ > twg mailing list > twg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/twg-opensfs.org -- --- Kit Westneat 812-484-8485 From carrier at cray.com Tue Dec 7 18:44:32 2010 From: carrier at cray.com (John Carrier) Date: Tue, 7 Dec 2010 12:44:32 -0600 Subject: [Twg] draft whitepaper In-Reply-To: <4CFE7D79.80605@ddn.com> References: <4CFE7D79.80605@ddn.com> Message-ID: Hi Kit, Good question. All development projects are to be run through the twg. This would include features to improve reliability or administration. I think the inputs into the roadmap and project stack from the support and release working groups would come about through our survey of the OpenSFS membership. Anyone else with an opinion? I think we also need to understand the hand-off of development features with these other groups. We discussed at SC the need for preparing release collateral and passing this all to Oracle. But should this be done through the release team? Thanks, --jc -----Original Message----- From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of Kit Westneat Sent: Tuesday, December 07, 2010 10:31 AM To: twg at lists.opensfs.org Subject: Re: [Twg] draft whitepaper Hi John, The support wg white paper mentions providing the twg with support metrics in order to help plan the evolution of OpenSFS, but the twg white paper currently doesn't discuss how support issues might affect the twg planning. Is this understood as part of surveying the OpenSFS membership? If the OpenSFS members decide to invest in a project to improve the reliability of Lustre, or even a subsystem of Lustre like quotas, would this be done through the twg or swg? Thanks, Kit On 12/07/2010 12:50 PM, John Carrier wrote: > Team, > > Please take a look at the one-page whitepaper I sent last week. I have received just one comment so far. What did I miss? What did I get wrong? > > The board expects the TWG to complete this task ASAP. > > Thanks, > > --jc > > > -----Original Message----- > From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of John Carrier > Sent: Wednesday, December 01, 2010 11:48 PM > To: twg at lists.opensfs.org > Subject: [Twg] draft whitepaper > > Attached is a draft whitepaper describing the function of the TWG. This is based on our one meeting at SC two weeks ago, so I fully expect it to change as more people get involved. Please feel free to add or edit the text. If there is time, I hope we can discuss the doc during the 12/2 call. > > Thanks, > > --jc > _______________________________________________ > twg mailing list > twg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/twg-opensfs.org -- --- Kit Westneat 812-484-8485 _______________________________________________ twg mailing list twg at lists.opensfs.org http://lists.opensfs.org/listinfo.cgi/twg-opensfs.org From carrier at cray.com Tue Dec 7 18:49:25 2010 From: carrier at cray.com (John Carrier) Date: Tue, 7 Dec 2010 12:49:25 -0600 Subject: [Twg] [Swg] OpenSFS Support Working Group draft white paper - Round 2 In-Reply-To: References: Message-ID: Pam, Dan raises a good point. If we're talking Lustre 2.x, then servers will only run on RHEL and its derivatives. But many of us are running Lustre 1.8.x on SLES and do not have the option of moving to RHEL without changing the hardware platform on which the servers are running. Is the support working group considering support for Lustre 1.8.x? Thanks, --jc From: Dan Ferber [mailto:dferber at whamcloud.com] Sent: Tuesday, December 07, 2010 9:56 AM To: John Carrier; Hamilton, Pam Subject: Re: [Swg] OpenSFS Support Working Group draft white paper - Round 2 I agree SLES (and RH) need to be added for clients. I know Oracle's support policy (for their current but not new customers and partners - and for 1.8 and not 2.0) is online, and here is Whamcloud's support plans, just as another datapoint: http://wiki.whamcloud.com/display/PUB/Support+Matrix [cid:image001.png at 01CB95FC.69DA4300] Dan [cid:image002.png at 01CB95FC.69DA4300] Dan Ferber Customer Service and Business Development Manager Office: +1 651-344-1846 Cell: +1 651-344-1846 Email: dferber at whamcloud.com From: John Carrier > Date: Tue, 7 Dec 2010 11:46:06 -0600 To: "Hamilton, Pam" > Cc: "comswg at lists.opensfs.org" >, "Execs at lists.opensfs.org" >, "swg at lists.opensfs.org" >, "twg at lists.opensfs.org" >, "rpwg at lists.opensfs.org" > Subject: Re: [Swg] OpenSFS Support Working Group draft white paper - Round 2 Hi Pam, Your description of release mentions Redhat only. I think this is appropriate for the servers, but many members are using other distributions on their cluster nodes. What are the support plans for the Lustre clients for the OpenSFS distribution? Oracle has stated that their Linux support is for OEL and other RHEL derivatives. Clients will continue to be supported on all current distributions (SLES, RHEL, etc). Thanks, --jc From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of Hamilton, Pam Sent: Monday, December 06, 2010 8:31 PM To: Execs at lists.opensfs.org; twg at lists.opensfs.org; rpwg at lists.opensfs.org; swg at lists.opensfs.org; comswg at lists.opensfs.org Subject: [Twg] OpenSFS Support Working Group draft white paper - Round 2 Hi all, Attached is another draft of the OpenSFS Support Working Group white paper. During support model discussions, the SWG felt we were lacking a good understanding of the high level requirements for OpenSFS and any release distributed by the organization. Without this understanding, we are struggling with developing the support model. Therefore, this draft of includes a section with what we see as high level requirements for OpenSFS. I'm sending to the execs and other working groups to get a discussion started and hopefully we can reach a consensus regarding the requirements. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA 94551-9900 E-Mail: pgh at llnl.gov Phone: 925-423-1332 Fax: 925-423-8719 ___________________________________ _______________________________________________ swg mailing list swg at lists.opensfs.org http://lists.opensfs.org/listinfo.cgi/swg-opensfs.org -------------- next part -------------- An HTML attachment was scrubbed... URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: image001.png Type: image/png Size: 81515 bytes Desc: image001.png URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: image002.png Type: image/png Size: 8719 bytes Desc: image002.png URL: From carrier at cray.com Tue Dec 7 19:25:13 2010 From: carrier at cray.com (John Carrier) Date: Tue, 7 Dec 2010 13:25:13 -0600 Subject: [Twg] next meeting: 12/10/2010 Message-ID: A reminder that because of schedule conflicts, our meeting this week will be Friday 12/10 at 9:30a PT/ 12:30a ET. Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Proposed agenda : · whitepaper o are we ready to pass the doc to the rest of the OpenSFS working groups for review? · roadmap o Eric's roadmap slides introduce a number of development items. http://www.opensfs.org/wp-content/uploads/2010/11/Lustre-development-roadmap1.pdf o Are there additional items to consider? o We need to vote on the proposed direction and begin prioritizing work items. o Our short-term goal should be to select an item that we can use to begin working on with the developer community. For example, based on the discussion last Friday, phase1 of sync CMD seems pretty well defined and most of the code already exists. § Could this be something OpenSFS could drive to completion? § Can a short-term project proceed in advance of a complete survey of the HPC Lustre community? Let me know if you have any additions. --jc -------------- next part -------------- An HTML attachment was scrubbed... URL: From hamilton5 at llnl.gov Tue Dec 7 20:00:00 2010 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Tue, 7 Dec 2010 12:00:00 -0800 Subject: [Twg] [Swg] OpenSFS Support Working Group draft white paper - Round 2 In-Reply-To: References: Message-ID: Hi John, The issues you raise regarding support for the different Lustre versions are valid. Our working group has not had specific discussions related to which versions will be supported or not. We are still trying to get our heads around what are the high level requirements we are trying to meet. I will make a note that we must address which Lustre versions are supported. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA 94551-9900 E-Mail: pgh at llnl.gov Phone: 925-423-1332 Fax: 925-423-8719 ___________________________________ From: John Carrier [mailto:carrier at cray.com] Sent: Tuesday, December 07, 2010 10:49 AM To: Hamilton, Pam Cc: Dan Ferber; twg at lists.opensfs.org; swg at lists.opensfs.org; rpwg at lists.opensfs.org; comswg at lists.opensfs.org; Execs at lists.opensfs.org Subject: RE: [Swg] OpenSFS Support Working Group draft white paper - Round 2 Pam, Dan raises a good point. If we're talking Lustre 2.x, then servers will only run on RHEL and its derivatives. But many of us are running Lustre 1.8.x on SLES and do not have the option of moving to RHEL without changing the hardware platform on which the servers are running. Is the support working group considering support for Lustre 1.8.x? Thanks, --jc From: Dan Ferber [mailto:dferber at whamcloud.com] Sent: Tuesday, December 07, 2010 9:56 AM To: John Carrier; Hamilton, Pam Subject: Re: [Swg] OpenSFS Support Working Group draft white paper - Round 2 I agree SLES (and RH) need to be added for clients. I know Oracle's support policy (for their current but not new customers and partners - and for 1.8 and not 2.0) is online, and here is Whamcloud's support plans, just as another datapoint: http://wiki.whamcloud.com/display/PUB/Support+Matrix [cid:image001.png at 01CB9605.BC3EDB70] Dan [cid:image002.png at 01CB9605.BC3EDB70] Dan Ferber Customer Service and Business Development Manager Office: +1 651-344-1846 Cell: +1 651-344-1846 Email: dferber at whamcloud.com From: John Carrier > Date: Tue, 7 Dec 2010 11:46:06 -0600 To: "Hamilton, Pam" > Cc: "comswg at lists.opensfs.org" >, "Execs at lists.opensfs.org" >, "swg at lists.opensfs.org" >, "twg at lists.opensfs.org" >, "rpwg at lists.opensfs.org" > Subject: Re: [Swg] OpenSFS Support Working Group draft white paper - Round 2 Hi Pam, Your description of release mentions Redhat only. I think this is appropriate for the servers, but many members are using other distributions on their cluster nodes. What are the support plans for the Lustre clients for the OpenSFS distribution? Oracle has stated that their Linux support is for OEL and other RHEL derivatives. Clients will continue to be supported on all current distributions (SLES, RHEL, etc). Thanks, --jc From: twg-bounces at lists.opensfs.org [mailto:twg-bounces at lists.opensfs.org] On Behalf Of Hamilton, Pam Sent: Monday, December 06, 2010 8:31 PM To: Execs at lists.opensfs.org; twg at lists.opensfs.org; rpwg at lists.opensfs.org; swg at lists.opensfs.org; comswg at lists.opensfs.org Subject: [Twg] OpenSFS Support Working Group draft white paper - Round 2 Hi all, Attached is another draft of the OpenSFS Support Working Group white paper. During support model discussions, the SWG felt we were lacking a good understanding of the high level requirements for OpenSFS and any release distributed by the organization. Without this understanding, we are struggling with developing the support model. Therefore, this draft of includes a section with what we see as high level requirements for OpenSFS. I'm sending to the execs and other working groups to get a discussion started and hopefully we can reach a consensus regarding the requirements. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA 94551-9900 E-Mail: pgh at llnl.gov Phone: 925-423-1332 Fax: 925-423-8719 ___________________________________ _______________________________________________ swg mailing list swg at lists.opensfs.org http://lists.opensfs.org/listinfo.cgi/swg-opensfs.org -------------- next part -------------- An HTML attachment was scrubbed... URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: image001.png Type: image/png Size: 81515 bytes Desc: image001.png URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: image002.png Type: image/png Size: 8719 bytes Desc: image002.png URL: From carrier at cray.com Thu Dec 9 23:47:26 2010 From: carrier at cray.com (John Carrier) Date: Thu, 9 Dec 2010 17:47:26 -0600 Subject: [Twg] 2010-12-03 meeting minutes Message-ID: Hi all, Attached are the minutes from last week's meeting. Please let me know if you have any corrections or additions (eg, I may not have caught everyone who was on the call). This is also a reminder that we are continuing the discussion tomorrow at 9:30a PT / 12:30 ET. Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. --jc -------------- next part -------------- OpenSFS Technical Working Group Meeting minutes : 12/03/2010 Concall: start : 9:30a PT, end 10:42a PT Next meeting: Friday, 12/10/2010 @ 9:30a PT/12:30p ET Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. [Note: the original meeting was to have been on 12/2, but due to a schedule mix-up, we were forced to delay the meeting 24hrs.] Attending: Name Organization email ----------------- -------------- ---------------------------- John Carrier Cray carrier at cray.com Justin Miller Indiana Univ. ? Steve Simms Indiana Univ. ssimms at indiana.edu Josh Walgenbach Indiana Univ. jjw at indiana.edu Chris Moronne LLNL morrone2 at llnl.gov Andreas Dilger Oracle andreas.dilger at oracle.com Dave Dillow ORNL dillowda at ornl.gov Sarp Oral ORNL oralhs at ornl.gov Galen Shipman ORNL ghsipman at ornl.gov Ercan Kamber RAID, Inc. ercan_kamber at raidinc.com Eric Barton Whamcloud eeb at whamcloud.com Agenda: Review Eric's slides "Lustre Development Roadmap" [http://www.opensfs.org/wp-content/uploads/2010/11/Lustre-development-roadmap1.pdf] [Editor's note: I have broken Eric's presentation into three sections -- intro, building blocks, development tasks -- that flow from the general principles to more specific items. Also, just a reminder, these are notes and not a stand-alone description of Eric's slides.] 03 Engineering principles - Lustre has lots of code and want to keep it scaling. - not just customers experiencing instability will try something new but cannot continue development without stability = therefore, stability must be at heart of the strategy - to do this, must develop small steps and must have a clear direction where the development is going = we must think ahead to where we want to be in 5yrs so development today is progression to this goal - ALSO, fundamental to a project like lustre, must modify subsystems - always fighting schedules, then try to take shortcuts - interruption in middle of development of new structures, leads to keeping old structures in place - challenge for developer to work on new things while keeping old things in mind - technical debt: anytime you keep old code in place, then are incurring debt and there is interest accrued over time for keeping the old code around example: Sun's purchase of CFS resulted in a 3yr hiatus to remove the technical debut accrued by CFS 04 agenda far attractors = long term directions future technologies = game-changing building blocks development choices = what needs to be done now while anticipating the evolution of Lustre section 1 : long-term goals --------------------------- 05 * major requirements = extreme system size - more clients (nodes, process/node) - fault management is an issue at scale of these devices - failing of lustre today is that advertised fault tolerance works at tiny scale, but breaks down at scales that Lustre claims to attract - fan out ratio is huge (all clients to all servers) - need fault management in all-to-all - well bound time to detect failure and respond to it => if not, then spend time responding to faults = throughput of servers - 1000's of servers (OST, MDT) - HPC loads : barriers lead to "denial of service" - exascale = barriers will be evil - will find that apps remove barriers and have constant I/O loads - consistent with enterprise (cloud) goals - MDS performance - start addressing issue of multiple MDS - make lustre more general purpose - in absence of contention, it should perform like a local file system = usability - central/global file system - share with other clusters - need QoS, SLAs - security is an issue for cross campus i/o - data management - HSM is just a start - better ways to exploit different types of storage - cheap for capacity - expensive for speed - structure files and data sets - what other layouts are possible? - administration - existing tools not developed with mind of admin in mind - how to cope with scale of devices, - how to present the status of the devices section 2: Building Blocks (future technologies) ------------------------------------------------------ 06 * health network - separate fault detection from the inband protocol - currently, all client/server failures by inband communications - but server load is so high that one client among 100Ks has very high latency - hence, using latency as indication of component failure isn't very useful - thus finding out status of peer is a facility needed to cope with systems at enormous scale how? - logically separate network implementing spanning tree over all nodes - tree rooted where all data is aggregated - make global decisions - root must be fault tolerant (eg paxos) but other nodes must be redundant - result is that communication scales latency as O(log n) - how to embed this inside LNET? - LNET has diff message priorities - embed the spanning tree in LNET and give guaranteed latency across system - use for - prompt failure notifications - collectives (census) - distributed transactions 07 * distributed transactions (epochs) - have had specific adhoc solutions to distributed transactions - hard to implement CMD because don't have distr.trans, - has adhoc solution: synchronous sequential operations - ensures that always safe - okay for a start, but not general - need cleanup after failure - need to sync transactions to disk - eeb is not keen on two-phase commits - epochs implement global transactions - system wide transaction number - servers keep roll-back logs - clients keep roll-forward logs - in case of any failure of a client - server can roll-back - in case of server failure - system can roll-back to stable point - roll-forward from clients to capture new data this tool lets you do CMD, but also - aggregate RPCs - RAID files - replication 08 * caching - example = MDS caching - need both health network and distributed transactions to enable scalable caching - allows Lustre to aggregate clients on single node - with SDDs, offer home directories as fast as local storage - need to aggregate i/o and lock access - what are policies for exploiting the features efficiently 09 * layouts - currently only have RAID0 (ie striping across OSTs) - want more complex, dynamic layouts layout locks from HSM are the key currently, Lustre has objects and layout stays with object - allow files to be on MDT when small - move to objects (raid0) when larger change layout incrementally - then do seamless replication and migration - for example, if know OSTs will fail regularly, then could mirror OSTs or files within the file also enable non-posix access Discussion : Q: order of slides A: health network (6) enables distributed transactions (7) distributed transactions (7) enables a lot of caching (8) layouts (9) is independent of the other three technologies for example, distributed transactions are important, but need scalable collectives section 3 : Development tasks ------------------------------ 10 * current issues lustre 1.8 - MD performance - failover latency - ldiskfs scalability discussion: Q: AD: lustre 1.8 fixes? A: EB: not taking plunge to 2.x until see someone else to it AD: but adding features EB: strict separation of feature and bugfix releases AD: 1.10? ldiskfs 16TB EB: 1.9 would be a feature release depends on AD: new 1.x could fragment the development effort EB: any development for 1.x has to port forward for 2.x DD: problems for 1.8 also exist in 2.0 lustre 2.0 - performance regressions - smp scaling (numio) - zfs / e2e integrity - hsm measurement & test 11 * ZFS alternatives BTRFS - consider BTRFS - does it fit with OSD model ldisfkfs - overcome 16TB limits - 64bit capacity - t10diff - hardening for future proofing - look at data structure to see fi they scale 12 * technical debt = distributed resources - existing quota and grant are fragile quota for resource allocation grant = space reservation on OST so that when cache cleaned there is space for data - more robust generic implementation of distributed resource management could do both as well as other features, eg total RPCs in flight - if only one client, could increase the RPCs in flight, but when all clients are active, then only want one RPC in flight = security - secure RPC is implemented, capabilities are not - ensure operations on one server granted by another, will continue - for example, connect on client across country - all operations will be authenticated and authorized - but OSTs and which particular user is using it, not authenticated - can subvert the protection of the namespace by accessing objects directly = capabilities complete the implementation - returned with file handle to allow user to access file objects 13 * performance = multirail networking - aggregate IB for greater per node performance - failover = server side request scheduling - service thread pools in Lustre created to avoid deadlocks - request handling has been elaborated so that requests inspected and scheduled to threads for handling - have the beginnings: know some are high priority and moved to front of queue - but to do it properly need general purpose request scheduling (QoS, priority,) would allow better use of threads, gang-scheduling (avoid starvation) eg, allow viz cluster to work while doing crash dump from compute cluster 14 * CMD - full BW for all types of servers requires lots of development - requires health network and epochs - working on version of CMD based on write through and sequenced operations - if have ops on three diff servers, then order the ops so that, if there is a failure, the worst that can happen is that the link counts will not - increment link counts before creating reference - destroy references before decrementing link counts - worst case is an orphan reference (objects exist but are unused by namespace) - can be scanned with background utility - development in two phases phase1: use admin function to create directories with distrubed functions on different MDTs - access directories with 1-1 access as it is today - allow home directories on different MDT than the directories for the scratch i/o - created independence of applications using - similare to mounting differerent File system that happen to use the same OSTs phase2 : stripe directories - have files in directories on diff targets - hash filename, tells where to get MDTs for the - also start out with admin function, but thereafter, all i/o in directory is local to MDT - stripe directories give order and speed-up for file create in single directory and speed-up for large directories phase1 - reasonable to handle multiple links with xdev phase2 - less reasonable to say that rename and link will give xdev because moving files between MDTs discussion: AD: code exists already - but have implementation to get into production state EB: rough estimate - non-trivial to get into production [meeting adjourned at 10:42a PT -- two slides, fault management and administration yet to be discussed] From gshipman at ornl.gov Fri Dec 10 03:29:51 2010 From: gshipman at ornl.gov (Shipman, Galen M.) Date: Thu, 09 Dec 2010 22:29:51 -0500 Subject: [Twg] OpenSFS Conflict of Interest Policies & Procedures Message-ID: All, The board of directors approved the following conflict of interest policy/procedure regarding the RFP process that the TWG will need to comply with: The Corporation may solicit input from Participants regarding development and support activities from time to time, including recommended requirements to be included in requests for proposals (“RFP’s). The Board may authorize Working Groups to use such input and recommended requirements to generate RFP’s, review and select offeror proposals, and negotiate contracts relating to such RFP’s. As requested or permitted by the Board or the Working Group, a Participant may provide input and recommendations on the RFP process or particular RFP’s, provided, however, that a Participant may not bid on a particular RFP if the Participant or its Representatives were involved in preparing or approving the material terms of such RFP. To comply with this I suggest that we have a clear delineation of when the RFP process begins and ensure that all participants are aware of when the process will begin and provide them the opportunity to elect to step down from this function of the TWG. Thanks, Galen From carrier at cray.com Fri Dec 10 17:38:19 2010 From: carrier at cray.com (John Carrier) Date: Fri, 10 Dec 2010 11:38:19 -0600 Subject: [Twg] whitepaper link Message-ID: http://lists.opensfs.org/pipermail/twg-opensfs.org/attachments/20101202/46e7fccb/attachment-0001.pdf From jjw at indiana.edu Wed Dec 15 16:03:09 2010 From: jjw at indiana.edu (Joshua Walgenbach) Date: Wed, 15 Dec 2010 11:03:09 -0500 Subject: [Twg] IU's priorities from the Whamcloud roadmap Message-ID: <4D08E6BD.1040708@indiana.edu> Hi All, Sorry this is a bit late. After some discussion, these are the priorities in decreasing priority we'd like to put forward from Eric's roadmap for short term projects that could be supported by OpenSFS. We know not all of these may happen given resources, but we'd be happy with any of the three. 1) Lustre Error Reporting. I heard that John Hammond from TACC has some patches to make Lustre logging a bit more sane, but have not had a chance to look at them. It might be a place to start. 2) Administration and Management tools. 3) Survey of BTRFS as a backend store for OSTs. I hope there is some easy consensus on what folks want. -Josh -- Joshua Walgenbach, Data Capacitor Team, Indiana University From dillowda at ornl.gov Wed Dec 15 22:16:56 2010 From: dillowda at ornl.gov (David Dillow) Date: Wed, 15 Dec 2010 17:16:56 -0500 Subject: [Twg] ORNL priorities for roadmap Message-ID: <1292451416.30718.38.camel@lap75545.ornl.gov> My thoughts on our priorities. In many cases, the priorities in each category shift a bit depending on expected time to deployment. 1) Metadata scaling! There are many aspects to this, but this is a big pain point for us. While we expect to need to move to some form of CMD, using NRS to avoid blocking all of the MDS threads on one lock will give us much needed breathing room. Other items of interest include SMP scaling work, RPC aggregation, and subtree locking. The priorities in this category shift quite a bit depending on how long it will take to develop them, and how they lay foundation for future improvements. ORNL has some work going in this area, but that is a small part of the puzzle. 2) Failover and recovery time. It takes far too long for clients to notice failures and connect to the HA partners. It seems the health network ideas could help this, and perhaps allow us to run lower timeouts as well. ORNL has some work going in this area as well. 3) Future expansion/data protection ldiskfs has done well, but it is expected that a migration to ZFS or BTRFS is inevitable. Ensuring data integrity from client to platter and back is going to become a significant issue going forward. ORNL may be able to contribute to benchmarking and stress testing of the different backends, but I think OpenSFS will want to contract out the code review unless a member can contribute the necessary expertise. 4) Scalable management, analytics, etc. We know we have orphans and missing objects from various bugs encountered over the life of our production filesystems, so online filesystem checking is seen as a win to us -- we cannot afford the downtime to do the offline fscks required. I like the idea of the dynamic LNET configuration, but it is only truly useful if we stop requiring a server reboot to change the LNET a client comes in on -- currently the RPC layer will always talk to a client via the same LNET with which it first connected to the server. Being able to quickly purge files and run du without needing elaborate out-of-band systems will make our admin's life easier. -- Dave Dillow National Center for Computational Science Oak Ridge National Laboratory (865) 241-6602 office From carrier at cray.com Thu Dec 16 00:36:37 2010 From: carrier at cray.com (John Carrier) Date: Wed, 15 Dec 2010 18:36:37 -0600 Subject: [Twg] 2010-12-10 meeting minutes Message-ID: Hi all, Attached are my notes from last week's meeting. Please let me know if you have any corrections or additions. This email is also a reminder of our meeting tomorrow 12/16 at 9:30a PT / 12:30 ET. Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. We had two homework items from our last meeting: * review the whitepaper * get roadmap priorities from our home institutions I was told by a member of the board that the whitepaper needs to be our priority. This is our proposal for the goals of the working group and the process we will follow in pursuing those goals. We need to submit a draft of this document before we can pursue the roadmap and initiate any development discussions. Please be prepared to discuss the draft whitepaper at tomorrow's meeting. You can find a copy in the email archive: Thanks, --jc -------------- next part -------------- OpenSFS Technical Working Group Meeting minutes : 12/10/2010 Concall: start : 9:30a PT, end 10:32a PT Next meeting: Thursday, 12/16/2010 @ 9:30a PT/12:30p ET Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Attending: Name Organization email ----------------- -------------- ---------------------------- John Carrier Cray carrier at cray.com Kit Westneat DDN kwestneat at ddn.com Justin Miller Indiana Univ. jupmille at indiana.edu Josh Walgenbach Indiana Univ. jjw at indiana.edu Damian Hazen LBL dhazen at lbl.gov Chris Morrone LLNL morrone2 at llnl.gov Ken Hornstein NRL kenh at cmf.nrl.navy.mil Shay Seager OpenSFS ssseager at gmail.com Andreas Dilger Oracle andreas.dilger at oracle.com Dave Dillow ORNL dillowda at ornl.gov Ercan Kamber RAID, Inc. ercan_kamber at raidinc.com Eric Barton Whamcloud eeb at whamcloud.com Robert Reed Whamcloud rread at whamcloud.com Agenda: * whitepaper reminder * continue review Eric's slides "Lustre Development Roadmap" Whitepaper - John sent a draft of the whitepaper to the group on 12/1: Desc: OpenSFS TWG white paper_ver1.pdf URL: - John wants to send a draft of the document to the other working groups but needs feedback from the other members of the working group. Please read it next week. Send comments to the list. Roadmap slides section 3 : Development tasks (continued from 12/3/2010 meeting) ------------------------------ 15 * fault management = fault reporting - improve lustre's reporting of errors and faults - some work already done to classify errors - originally any error used C macro to report to syslog (c debug/event tracker) - problem is that errors are in system logs and no way to plug errors out of lustre into wider admin infrastructure - need to define report architecture - vendor neutral - flag fault class, severity, etc - may initially go to system event logs, but if trying to integrate lustre into vendor management infratstructure, the vendor could take feed of errors from lustre (collected on servers and clients), then channel them into someplace for fault logging or automatic administrative action based on faults reported - should anticipate health networks - recognize that health networks will be major stepping stone for scalablity of fault management of Lustre - need to anticipate how these networks should be checked discussion: Q: AD: newer mechanisms in kernel to avoid using syslogs - can use more than printk - work done at TACC to add different messages out of kernels - other channels: ex netlink -- should do something like it EB: major issue is to classify the errors consistently thinking ahead to what classes and severity of errors will be necessary in the future. require serious thought to get right classes / severities - or is there a different architecutre - need to feed into automatic tool, not parsing text messages essentially, the issue is how to put stethescope on side of lustre AD: pass messages over the network? EB: don't want to invent something that will be useless once these networks are available AD: classes of errors will report nodes as dead to the network? EB: or just how to use network for reporting errors. health networks respond to faults, this is how to report the faults Q: CM: specific admin tools? EB: no particular message in mind. don't want to second guess what vendors will use CM: work with vendors for what are the inputs they need EB: satisfy requirements, but not mandate the tool used AD: use standard tools (eg ganglia) = Online FSCK / scrubber - want to avoid lfsck, it doesn't scale - need mechanism to validate the file system and find inconsistencies in the FS introduced by bugs and crashes - sync CMD relying on orphan clean-up for failures in the middle of distributed metadata operation - need something to crawl over the entire file system based on policy to ensure anything damaged is detected and fixed - also want to do failover without doing full FS check - as long as we know underlying storage is consistent to crashes - guarantee that crawl of FS is in finite time - allow failover without lengthy check before hand - design work already done by the lustre group for Cray HPCS - but need to address OSD specific scanning - OSD-specific resilvering 16 * administration = improved configuration - move away from config files, use scripts instead - ex can configure LNET from userspace with lctrl - work is to allow network drivers to be brought up and down on individual interfaces - result puts all name handling into user space - currently have complex config with multiple networks - clumsy to configure the networking in one way, then bringup servers just to get nids in right places - need a simpler, text-driven config system. - configure file system as single text file on MGS discussion: Q: KH: config of clients via mod parameters too? EB: anything to do with file system itself should be part of textual file system configuration (eg global parameters that affect all clients). networking interfaces for clients would be LNET configs - do it with scripts like IP is configured - bring up interfaces etc, (more admin friendly) but client-specific parameters : need those in tunables set on clients, not part of file system Q: KH: keep those as mod params or put in user space? EB: want to overcome non-obviousness of creating file system and binding the server addresses. eg, currently some config information disappears after file system starts = OST/MDT pool management - OST pools are ways to name collections of OSTs (type of shorthand) - To make them more useful, have to control which users are using which pools - reserve pools for high priority projects - permissions would increase usability - rebalance pools based on OST usage , adding new OST - allow migration between peers - HSM allows migration between OST and tape. Want to extend to allow migration between different qualities of storage = global snapshots - use underlying OSD model - simple case: block, then copy - with epochs, could make it non-blocking - make snapshots as easy as it is with ZFS - consider standards to warn storage that snapshot is desired - need consistency of storage - but want to warn apps using the file system so they leave the correct state - any architecture doing global snapshots need to inform applications to save state discussion: AD: an ioctl allows application to initiate snapshot. could export to lustre to let it receive the ioctl EB: second guessing admin, requiring all apps to be crash recoverable is a burden on the apps AD: what are the APIs? EB: standards for doing it AD: unsynchronized snapshots is good EB: two cases: - app initiated snapshot - admin initiated snapshot for (2), if app is crash recoverable, go to previous state and roll forward or - have IOCTLs to tell application about freeze/thaw events - for freeze message, get in good state on file system for application consistency. - means recovery is easier later on DD: 3rd option to have apps tell FS that the app is busy and not to snapshot EB: yes, but need to consider range of apps running on FS simultaneously RR: need snapshot done first [end of slides] Discussion - next steps The working group needs to digest the information in Eric's slides. - take the roadmap back to home institutions for review - prioritize the featurs (and find features not yet discussed) - identify short-term projects for OpenSFS to pursue Andreas mentioned the process Sun followed for Cray's HPCS program. There were two RFPs: one for design followed by one for implementation. - the HPCS designs are linked here: Eric suggested one short-term project for OpenSFS would be a btrfs evaluation. BtrFS has similar properties to ZFS and is a contender for the Lustre backing store. This project would be a survey of the code and the stability of its current implementation. The result would be an understanding of the suitability of btrfs for Lustre and the scope of the development needed to create a btrfs OSD. From carrier at cray.com Thu Dec 16 19:41:40 2010 From: carrier at cray.com (John Carrier) Date: Thu, 16 Dec 2010 13:41:40 -0600 Subject: [Twg] 2010-12-16 meeting minutes Message-ID: Attached are the minutes from today's meeting. I will post the new whitepaper draft shortly. --jc -------------- next part -------------- OpenSFS Technical Working Group Meeting minutes : 12/16/2010 Concall: start : 9:30a PT, end : 10:32a PT) Next meeting: Thursday, 1/6/2011 @ 9:30a PT/12:30p ET Dial-in numbers are 715-726-4994 or 866-304-8294. The meeting ID and password are 7012090. Attending: Name Organization email ----------------- -------------- ---------------------------- John Carrier Cray carrier at cray.com Justin Miller Indiana Univ. jupmille at indiana.edu Steve Simms Indiana Univ. ssimms at indiana.edu Josh Walgenbach Indiana Univ. jjw at indiana.edu Damian Hazen LBL dhazen at lbl.gov Shay Seager OpenSFS ssseager at gmail.com Andreas Dilger Oracle andreas.dilger at oracle.com Dave Dillow ORNL dillowda at ornl.gov Ercan Kamber RAID, Inc. ercan_kamber at raidinc.com Robert Reed Whamcloud rread at whamcloud.com Agenda: * review the draft whitepaper Discussion The team thought the first draft captured the discussion at SC. We surveyed individuals and the biggest hole was not addressing our interaction with the other working groups in OpenSFS. We then went through the document section by section. Shay mentioned the need to address expected resource commitments from the member companies. We addressed this idea in a new section. Specific changes are described below: Mission statement - needed to be broader - replace "I/O performance requirements" with "stability, performance, and management requirements" - remove reference to "Moore's law", although that is at the heart of the imbalance between processors, memory, and storage performance. Roadmap - remove references to Oracle. OpenSFS is focused on the needs of the HPC community. period. - add text that requirements will come working with other working groups in addition to the OpenSFS membership and Lustre community. Feature Development - add text to describing our intent to work with the other working groups during the RFP process and require contractors to work with them during integration. Member Requirements -- new section - to help scope time commitments, we defined the expected roles for the following functions - working group co-chairs - RFP co-leads - members Meeting schedule - no change From carrier at cray.com Thu Dec 16 19:44:48 2010 From: carrier at cray.com (John Carrier) Date: Thu, 16 Dec 2010 13:44:48 -0600 Subject: [Twg] draft2 of the whitepaper Message-ID: Thanks to the discussion at today's meeting, we have a new version of the whitepaper. Attached are two PDF versions. The one labeled markup shows exactly what changed from the first version. PLEASE take a look and send any corrections (edits, additions, deletions) to me ASAP. I will send the corrected doc to all of the working groups tomorrow. Thanks for your help. --jc Ps - let me know if you want the original MS word docx version and I'll send it along. -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS TWG white paper_ver2.pdf Type: application/pdf Size: 184563 bytes Desc: OpenSFS TWG white paper_ver2.pdf URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS TWG white paper_ver2_markup.pdf Type: application/pdf Size: 209914 bytes Desc: OpenSFS TWG white paper_ver2_markup.pdf URL: From carrier at cray.com Fri Dec 17 21:44:33 2010 From: carrier at cray.com (John Carrier) Date: Fri, 17 Dec 2010 15:44:33 -0600 Subject: [Twg] Draft of TWG whitepaper Message-ID: The team spent yesterday's meeting revising our earlier internal draft of the document. The attached is ready for review by the rest of opensfs. We look forward to your comments and feedback. --jc -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS TWG white paper_ver2.pdf Type: application/pdf Size: 184563 bytes Desc: OpenSFS TWG white paper_ver2.pdf URL: From seager1 at llnl.gov Mon Dec 20 19:56:02 2010 From: seager1 at llnl.gov (Seager, Mark K.) Date: Mon, 20 Dec 2010 11:56:02 -0800 Subject: [Twg] OpenSFS Working Group Cross Fertalization Telecon Message-ID: <6DAF24F9BB747E47A4BAE7A6F5B80C7DFB74E56337@NSPEXMBX-A.the-lab.llnl.gov> When: Tuesday, December 28, 2010 9:00 AM-10:00 AM (UTC-08:00) Pacific Time (US & Canada). Where: 517-308-1709x5282795 Note: The GMT offset above does not reflect daylight saving time adjustments. *~*~*~*~*~*~*~*~*~* The working group cross fertilization telecon between Christmas and New Years holiday is hereby canceled. Have a great Holliday! Regards, ++Mark -------------- next part -------------- An HTML attachment was scrubbed... URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: not available Type: text/calendar Size: 3446 bytes Desc: not available URL: From carrier at cray.com Tue Dec 21 14:49:41 2010 From: carrier at cray.com (John Carrier) Date: Tue, 21 Dec 2010 08:49:41 -0600 Subject: [Twg] Shay's notes from 12/16 meeting Message-ID: I received these notes from Shay after I had posted mine to the reflector. I am forwarding them along since she captured more of the flow of conversation. --jc -------------- next part -------------- A non-text attachment was scrubbed... Name: 12-16-10 TWG concall.pdf Type: application/pdf Size: 100832 bytes Desc: 12-16-10 TWG concall.pdf URL: