From hamilton5 at llnl.gov Mon Jul 2 17:20:58 2012 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Mon, 2 Jul 2012 10:20:58 -0700 Subject: [cdwg] Canceled: OpenSFS Community Development WG call Message-ID: When: Wednesday, July 04, 2012 9:00 AM-10:00 AM (UTC-08:00) Pacific Time (US & Canada). Where: Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# Note: The GMT offset above does not reflect daylight saving time adjustments. *~*~*~*~*~*~*~*~*~* Cancelling this week’s meeting… -------------- next part -------------- An HTML attachment was scrubbed... URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: not available Type: text/calendar Size: 2820 bytes Desc: not available URL: From hamilton5 at llnl.gov Tue Jul 3 23:24:59 2012 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Tue, 3 Jul 2012 16:24:59 -0700 Subject: [cdwg] OpenSFS Community Development WG call Message-ID: When: Thursday, July 12, 2012 9:00 AM-10:00 AM (UTC-08:00) Pacific Time (US & Canada). Where: Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# Note: The GMT offset above does not reflect daylight saving time adjustments. *~*~*~*~*~*~*~*~*~* Hi all, We normally have the CDWG calls on Wednesday mornings but I have a conflict on Wednesday, 7/11, so I’m proposing we have our call on Thursday, 7/12, at 9am Pacific. I am also available on Tuesday, 7/10, at 9am Pacific, if enough people would prefer the call on that day instead. Just let me know. Agenda: • Discuss a process for coming to consensus regarding the next maintenance release Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# Regards, Pam -------------- next part -------------- An HTML attachment was scrubbed... URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: not available Type: text/calendar Size: 2279 bytes Desc: not available URL: From jlevi at whamcloud.com Fri Jul 6 17:14:32 2012 From: jlevi at whamcloud.com (Jodi Levi) Date: Fri, 06 Jul 2012 11:14:32 -0600 Subject: [cdwg] Lustre 2.3 Update - July 6th 2012 Message-ID: Hi there Here is an update on the Lustre 2.3 Release. Landings ========= * A number of landings made ­ see http://git.whamcloud.com/?p=fs/lustre-release.git;a=shortlog;h=refs/heads/ma ster * All Production Features have landed (SMP, 2.6.38/FC15 Client, Job Stats, OI Scrub, CLIO RPC Engine Improvements, CRC, FLOCK). Some Preview Features have also landed (OSD-OFD, OSD-ZFS, Quota Accounting) There may be some patches that land over the weekend for the Preview Features; however by Monday we will be in full Release Testing mode. Testing ========== * Testing on the 2.2.59 tag is in progress * Full Release Testing mode will begin on Monday Blockers ========== * Full list is available at http://jira.whamcloud.com/secure/IssueNavigator.jspa?mode=hide&requestId=102 05 Other ========== - We will continue to provide periodic updates on our progress on this release. In the meantime, you can always see the landings as they happen at http://git.whamcloud.com/?p=fs/lustre-release.git;a=shortlog;h=refs/heads/ma ster and follow the patch reviews and testing at http://review.whamcloud.com/#q,status:open+project:fs/lustre-release+branch: master,n,z Thank you! -Jodi -------------- next part -------------- An HTML attachment was scrubbed... URL: From hamilton5 at llnl.gov Sat Jul 7 04:23:57 2012 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Fri, 6 Jul 2012 21:23:57 -0700 Subject: [cdwg] OpenSFS Community Development WG call Message-ID: When: Thursday, July 12, 2012 8:30 AM-9:30 AM (UTC-08:00) Pacific Time (US & Canada). Where: Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# Note: The GMT offset above does not reflect daylight saving time adjustments. *~*~*~*~*~*~*~*~*~* Hi all, I’m moving this call up a half-hour to accommodate anyone who also participates in the TWG call. Agenda: • Discuss a process for coming to consensus regarding the next maintenance release Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# Regards, Pam -------------- next part -------------- An HTML attachment was scrubbed... URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: not available Type: text/calendar Size: 2495 bytes Desc: not available URL: From hamilton5 at llnl.gov Wed Jul 11 20:49:06 2012 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Wed, 11 Jul 2012 13:49:06 -0700 Subject: [cdwg] Reminder - OpenSFS CDWG call tomorrow, 7/12, at 8:30am Pacific Message-ID: Hi all, This is a reminder that we will have a call tomorrow, July 12th, at 8:30am Pacific (Call-in: 866-914-3976 (925-424-8105) Passcode: 534986#). Agenda: . Discuss a process for coming to consensus regarding the next maintenance release Background: There has been a lot of discussion lately on the mail lists about Lustre maintenance releases and how it is determined which "feature" release gets additional maintenance releases. There is also confusion with regard to maintenance releases and the Community Tree maintenance contract. The contract specifically states: "Ongoing Lustre development requires periodic Lustre feature releases which are fully tested and qualified Lustre releases featuring the latest code (both features and bugfixes) from the master codeline (git://git.opensfs.org / fs/lustre-release.git). It is expected that feature releases should occur roughly every six months. These are distinct from Lustre maintenance releases, which are funded by support contracts." Basically, "maintenance" releases for each feature release are not covered by this contract and would not be feasible at the approximately four FTEs this contract funds. That being said, the CDWG is in a position to draft a recommendation regarding a long term support strategy for releases. Personally, I agree with Andreas Dilger who recommends that an Ubuntu-style "Long Term Stability" regular release strategy should be our goal. It is simply impossible in terms of resources to consider that each six-month feature release will have its own series of maintenance releases. Furthermore, to aid sites in planning for Lustre upgrade paths, we should decide on what will be the next LTS release well in advance of its actual release date. Therefore the discussion for the call today is to discuss a process for coming to consensus regarding the next maintenance release. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA  94551-9900 E-Mail:  pgh at llnl.gov Phone:  925-423-1332          Fax:  925-423-8719 From spitzcor at cray.com Wed Jul 11 22:51:47 2012 From: spitzcor at cray.com (Cory Spitz) Date: Wed, 11 Jul 2012 17:51:47 -0500 Subject: [cdwg] Reminder - OpenSFS CDWG call tomorrow, 7/12, at 8:30am Pacific In-Reply-To: References: Message-ID: <4FFE0383.1060304@cray.com> Hi, > Therefore the discussion for the call [tomorrow, 7/12] is to discuss a process for coming to consensus regarding the next maintenance release. As Pam noted, the current contract doesn't maintenance branches, so in terms of the (current) contract, I don't think that it is right for the CDWG to expect that Whamcloud would follow any given recommendation. Having said that, I think that we should have the discussion anyway with the thought that Whamcloud is looking for that input :) That's what the previous surveys were for too. It's also clear to me that we should think about changing the next contract to cover maintenance branches in some manner. Clearly we, as a community, care deeply about it. Maybe we could entertain a discussion about a so-called "tick-tock" model too. That is a big topic, so I propose that we tackle it in an upcoming meeting. That is, not tomorrow. Another topic I'd like to raise for a future discussion is quality metrics for the feature releases. We haven't formally defined any and I think that we should take a stab at it. I realize that it doesn't do any good without more test commitments and also that we haven't all chipped in on testing as much as we wanted in the past. Maybe this is where the discussion on funding test enhancements come in, but I think that there are other options to explore in the meantime. We'll need to figure out how to leverage the OpenSFS test cluster too. Thanks, -Cory On 07/11/2012 03:49 PM, Hamilton, Pam wrote: > Hi all, > > This is a reminder that we will have a call tomorrow, July 12th, at 8:30am Pacific (Call-in: 866-914-3976 (925-424-8105) Passcode: 534986#). > > Agenda: > . Discuss a process for coming to consensus regarding the next maintenance release > > Background: > There has been a lot of discussion lately on the mail lists about Lustre maintenance releases and how it is determined which "feature" release gets additional maintenance releases. There is also confusion with regard to maintenance releases and the Community Tree maintenance contract. The contract specifically states: > > "Ongoing Lustre development requires periodic Lustre feature releases which are fully tested and qualified Lustre releases featuring the latest code (both features and bugfixes) from the master codeline (git://git.opensfs.org / fs/lustre-release.git). It is expected that feature releases should occur roughly every six months. These are distinct from Lustre maintenance releases, which are funded by support contracts." > > Basically, "maintenance" releases for each feature release are not covered by this contract and would not be feasible at the approximately four FTEs this contract funds. That being said, the CDWG is in a position to draft a recommendation regarding a long term support strategy for releases. Personally, I agree with Andreas Dilger who recommends that an Ubuntu-style "Long Term Stability" regular release strategy should be our goal. It is simply impossible in terms of resources to consider that each six-month feature release will have its own series of maintenance releases. > > Furthermore, to aid sites in planning for Lustre upgrade paths, we should decide on what will be the next LTS release well in advance of its actual release date. > > Therefore the discussion for the call today is to discuss a process for coming to consensus regarding the next maintenance release. > > Regards, > Pam > > ___________________________________ > Pam Hamilton > Lawrence Livermore National Lab > P.O. Box 808, L-556 > Livermore, CA 94551-9900 > E-Mail: pgh at llnl.gov > Phone: 925-423-1332 Fax: 925-423-8719 > > > _______________________________________________ > cdwg mailing list > cdwg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org From carrier at cray.com Thu Jul 12 14:30:40 2012 From: carrier at cray.com (John Carrier) Date: Thu, 12 Jul 2012 14:30:40 +0000 Subject: [cdwg] Lustre release notes Message-ID: During the course of several discussions with OpenSFS members since LUG, I have come to the following understanding of how Whamcloud manages Lustre releases. I prepared the notes below for my team in Cray. Peter and Pam thought these notes might also help provide a baseline for further discussion within the cdwg. Errors or omissions are mine -- please send corrections and comments to this reflector. thanks, --jc Whamcloud produces two Lustre release streams: a feature release every 6 months and a maintenance release every 3 months. Feature releases are new branches off of the master codeline that include all new code and the bugfixes accumulated during the previous 6 months. Maintenance releases apply bugfixes to the existing maintenance branch, which is currently Lustre 2.1. Whamcloud documents the release roadmap for both streams through its "Community Lustre Roadmap" wiki [1]. Feature landings are projected based on code Whamcloud produces for NRE contracts as well as code received from 3rd parties. These features and expected landings are detailed on a "Development in Progress" wiki [2]. OpenSFS currently has two contracts with Whamcloud. The first is a development contract to produce features intended to improve Lustre metadata performance. This contract stipulates that all features must land in the master codeline (aka the canonical tree). The second contract shares support with Whamcloud for the release manager and gatekeeper positions for the feature releases. It is this latter contract which has caused confusion within the Community. The contract states that Ongoing Lustre development requires periodic Lustre feature releases which are fully tested and qualified Lustre releases featuring the latest code (both features and bugfixes) from the master codeline (git://git.opensfs.org/fs/lustre-release.git). It is expected that feature releases should occur roughly every six months. These are distinct from Lustre maintenance releases, which are funded by support contracts. Nonetheless, perhaps because the title of contract is "Lustre Development Community Tree Maintenance", the expectation among OpenSFS members has been that bugfixes found after a feature release would be applied back to the feature release branch where the bugs were found. This level of maintenance is not, however, what Whamcloud agreed to provide with this contract. Landing bugfixes is a two-step process. Whamcloud lands all bugfixes against the master codeline. They then backport these patches to the maintenance branch. Because it takes considerable resources (both people and hardware time) to backport the patches to a feature branch, test the new code and release a maintenance update, Whamcloud only produces these updates when it is viable to do so. Subsequent feature releases will incorporate all bugfixes and feature landings to the master codeline at the time the branch is made. Because the Lustre 2 codeline has had substantial re-engineering work and is only recently getting wider testing exposure to a greater variety of hardware and production workloads, the initial 2.x releases have uncovered a number of bugs, especially in the new client package (CLIO) as well as interoperability with Lustre 1.8. There have been very few bugs found in the new features themselves. Over time, as the Community gains more experience and fixes more bugs in the Lustre2 codeline, there should be fewer bugs found following each feature release. For now, it will be easier for community members to key their production releases of Lustre from a Whamcloud maintenance release since this will include all of the bugfixes since the release branch was created. Community members can pull the bugfixes and apply the patches to a feature release themselves (as Cray and ORNL have done for Lustre 2.2), but this requires tracking master for relevant updates that should be applied to support their production use of the feature branch. A more strategic solution is to do more testing of a feature release candidate _before_ it is released. Even if a Community member has no interest in using a feature release in production, early testing with pre-release versions of feature releases will help identify instabilities created by the new feature with their workloads and hardware before the release is official. OpenSFS is working with Whamcloud to identify which future release branch will become the new maintenance branch. [1] http://wiki.whamcloud.com/display/PUB/Community+Lustre+Roadmap [2] http://wiki.whamcloud.com/display/PUB/Lustre+Community+Development+in+Progress -------------- next part -------------- An HTML attachment was scrubbed... URL: From nathan_rutman at xyratex.com Thu Jul 12 19:37:14 2012 From: nathan_rutman at xyratex.com (Nathan Rutman) Date: Thu, 12 Jul 2012 12:37:14 -0700 Subject: [cdwg] broader Lustre testing In-Reply-To: References: Message-ID: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> On Jul 12, 2012, at 7:30 AM, John Carrier wrote: > A more strategic solution is to do more testing of a feature release > candidate _before_ it is released. Even if a Community member has no > interest in using a feature release in production, early testing with > pre-release versions of feature releases will help identify > instabilities created by the new feature with their workloads and > hardware before the release is official. Taking a few threads that have discussed recently, regarding the stability of certain releases vs others, what maintenance branches are, what testing was done, and "which branch should I use": These questions, I think, should not need to be asked. Which version of MacOS should I use? The latest one, period. Why can't Lustre do the same thing? The answer I think lies in testing, which becomes a chicken and egg problem. I'm only going to use a "stable" release, which is the release which was tested with my applications. I know acceptance-small was run, and passed, on Master, otherwise it wouldn't be released. Hopefully it even ran on a big system like Hyperion. (Do we learn anything more about running acc-sm on other big systems? Probably not much.) But it certainly wasn't tested with my application, because I didn't test it. Because it wasn't released yet. Chicken and egg. Only after enough others make the leap am I willing to. So, it seems, we need to test pre-release versions of Lustre, aka Master, with my applications. To that end, how willing are people to set aside a day, say once every two months, to be "filesystem beta day". Scientists, run your codes, users, do your normal work, but bear in mind there may be filesystem instabilities on that day. Make sure your data is backed up. Make sure it's not in the middle of a critical week-long run. Accept that you might have to re-run it tomorrow in the worst case. Report any problems you have. What you get out of it is a much more stable Master, and an end to the question of "which version should I run". When released, you have confidence that you can move up, get the great new features and performance, and it runs your applications. More people are on the same release, so it sees even more testing. The maintenance branch is always the latest branch, you can pull in point releases with more bug fixes with ease. No more rolling your own Lustre with Frankenstein sets of patches. Latest and greatest and most stable. Pipe dream? -------------- next part -------------- An HTML attachment was scrubbed... URL: From nathan_rutman at xyratex.com Thu Jul 12 20:07:13 2012 From: nathan_rutman at xyratex.com (Nathan Rutman) Date: Thu, 12 Jul 2012 13:07:13 -0700 Subject: [cdwg] [Lustre-devel] broader Lustre testing In-Reply-To: References: Message-ID: On Jul 12, 2012, at 12:48 PM, Bruce Korb wrote: > Hi Nathan, > On 2012-07-12, at 20:37, Nathan Rutman wrote: > >> >> On Jul 12, 2012, at 7:30 AM, John Carrier wrote: >> >>> A more strategic solution is to do more testing of a feature release >>> candidate _before_ it is released. Even if a Community member has no >>> interest in using a feature release in production, early testing with >>> pre-release versions of feature releases will help identify >>> instabilities created by the new feature with their workloads and >>> hardware before the release is official. >> >> >> Taking a few threads that have discussed recently, regarding the stability of certain releases vs others, what maintenance branches are, what testing was done, and "which branch should I use": >> These questions, I think, should not need to be asked. Which version of MacOS should I use? The latest one, period. Why can't Lustre do the same thing? The answer I think lies in testing, which becomes a chicken and egg problem. I'm only going to use a "stable" release, which is the release which was tested with my applications. I know acceptance-small was run, and passed, on Master, otherwise it wouldn't be released. Hopefully it even ran on a big system like Hyperion. (Do we learn anything more about running acc-sm on other big systems? Probably not much.) But it certainly wasn't tested with my application, because I didn't test it. Because it wasn't released yet. Chicken and egg. Only after enough others make the leap am I willing to. >> So, it seems, we need to test pre-release versions of Lustre, aka Master, with my applications. To that end, how willing are people to set aside a day, say once every two months, to be "filesystem beta day". Scientists, run your codes, users, do your normal work, but bear in mind there may be filesystem instabilities on that day. Make sure your data is backed up. Make sure it's not in the middle of a critical week-long run. Accept that you might have to re-run it tomorrow in the worst case. Report any problems you have. >> What you get out of it is a much more stable Master, and an end to the question of "which version should I run". When released, you have confidence that you can move up, get the great new features and performance, and it runs your applications. More people are on the same release, so it sees even more testing. The maintenance branch is always the latest branch, you can pull in point releases with more bug fixes with ease. No more rolling your own Lustre with Frankenstein sets of patches. Latest and greatest and most stable. >> >> Pipe dream? On Jul 12, 2012, at 12:48 PM, Bruce Korb wrote: > > _I_ think so. You might get a few customers to say, "yes" but > never be able to find the appropriate round tuit. A more fruitful > approach might be to solicit customer acceptance tests. Presumably, > they've written them to hit the wrinkles that they tend to stub > their toes on. And there may be exceptions, too. (e.g. Cray might > well actually do some pre-testing -- they, too, have paying customers.) > I have no aversion to customers writing and supplying their own acceptance tests, but I think that approach doesn't work for many of the cases: - acceptance tests may not exist; acceptance may simply be testing with large production codes - tests that run in a particular environment need to be significantly generalized - tests may not be sharable for various legal reasons This also doesn't have to be an all-or-nothing proposition -- interested parties will be able to use the latest features, and will help contribute to the stability of Master, and will help reduce the "spread" of deployed systems, in a positive feedback loop. Yes, absolutely, this is effort on the part of Lustre users. But it can be balanced by the savings of efforts in roll-your-own, and risk reduction. -------------- next part -------------- An HTML attachment was scrubbed... URL: From adilger at whamcloud.com Thu Jul 12 20:25:15 2012 From: adilger at whamcloud.com (Andreas Dilger) Date: Thu, 12 Jul 2012 14:25:15 -0600 Subject: [cdwg] broader Lustre testing In-Reply-To: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> References: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> Message-ID: <9D94CDBC-4100-4087-8F9B-903E1D62D65C@whamcloud.com> On 2012-07-12, at 1:37 PM, Nathan Rutman wrote: > On Jul 12, 2012, at 7:30 AM, John Carrier wrote: >> A more strategic solution is to do more testing of a feature release >> candidate _before_ it is released. Even if a Community member has no >> interest in using a feature release in production, early testing with >> pre-release versions of feature releases will help identify >> instabilities created by the new feature with their workloads and >> hardware before the release is official. > > > Taking a few threads that have discussed recently, regarding the stability of certain releases vs others, what maintenance branches are, what testing was done, and "which branch should I use": > > These questions, I think, should not need to be asked. Which version of MacOS should I use? The latest one, period. Interesting... I _don't_ run the latest version of MacOS, and I distinctly recall people having a variety of issues with 10.7.0 when it was released. Does that mean the MacOS testing was insufficient? Partly, but it is unrealistic to test every possible usage pattern, so testing has to be "optimized" to cover the most common use cases in order to be finished within both time and cost constraints. > Why can't Lustre do the same thing? The answer I think lies in testing, which becomes a chicken and egg problem. I'm only going to use a "stable" release, which is the release which was tested with my applications. I know acceptance-small was run, and passed, on Master, otherwise it wouldn't be released. Hopefully it even ran on a big system like Hyperion. (Do we learn anything more about running acc-sm on other big systems? Probably not much.) Right. I don't think that acc-sm is the end-all in testing frameworks, and I freely admit that there is a lot more testing that could be done, both in scale and in the types of loads that are used. The acceptance-small.sh script is intended to be an "optimized" test set that can run in a few hours to give some reasonable confidence in a particular change. > But it certainly wasn't tested with my application, because I didn't test it. Because it wasn't released yet. Chicken and egg. Only after enough others make the leap am I willing to. There are all kinds of other load/stress tests (including applications) that can/should be run after the "basic" tests have been run to find new defects. When those defects are found they should be distilled down to a simple and specific test that gets added to the regular regression suite. I think it is this kind of testing that is needed moving forward. > So, it seems, we need to test pre-release versions of Lustre, aka Master, with my applications. I would caveat this to say - only test on tags which we know to be at least reasonably stable, since a lot of testing time will be wasted otherwise. > To that end, how willing are people to set aside a day, say once every two months, to be "filesystem beta day". Scientists, run your codes, users, do your normal work, but bear in mind there may be filesystem instabilities on that day. Make sure your data is backed up. Make sure it's not in the middle of a critical week-long run. Accept that you might have to re-run it tomorrow in the worst case. Report any problems you have. I'm not sure that users will be willing to do this, though some "friendly" users are known to make the leap onto new systems in order to get early/free CPU cycles on new clusters. There are also "feature tests" that need to be run at scale to validate new features, to ensure they are functional at scale, don't impact performance, and experiencing the kids of race conditions that scale testing provides. > What you get out of it is a much more stable Master, and an end to the question of "which version should I run". When released, you have confidence that you can move up, get the great new features and performance, and it runs your applications. More people are on the same release, so it sees even more testing. The maintenance branch is always the latest branch, you can pull in point releases with more bug fixes with ease. No more rolling your own Lustre with Frankenstein sets of patches. Latest and greatest and most stable. > > Pipe dream? I hope not. When I see users taking a specific release of Lustre, testing it, and then applying a patch series to their branch, the unfortunate result is more effort for the user (vendor/site, not end users) to maintain their patches, and more effort for support to determine if some _other_ bug is already fixed, or to debug a problem that appears only with a specific combination of patches applied, and then craft a different fix for that branch than the mainline. A better use case would be for users to start testing _before_ a major release is made, find/fix bugs, and merge the fixes into mainline, so when it is released in a maintenance release it will already be quite stable. This keeps the user patchset much smaller, and everyone will benefit from fixes from other testing before the release, and hopefully find fewer bugs in the field. It also avoids the issue of each user testing some cross-product of patches, and not really leveraging each others testing. Then, any bugs found in the field go into the maintenance branch and master, but there is much less of a need to "test" the maintenance branch, since the changes there should be relatively small. I think this is a reasonable approach, given that we no longer land features on maintenance branches. That means the risk of following maintenance releases is much smaller than it was in the 1.6 and 1.8 days (1.8.x only really entered "maintenance" mode with 1.8.6 or so). We've been trying to follow this model with LLNL. One issue is that 2.1.0 didn't really receive as much up-front testing as it could have, so it is getting more fixes than it should. We are working hard to land all of the LLNL (and other) bugfix patches into master and the next 2.1.x release. There is a parallel effort to test orion (2.4 development branch) so that by the time 2.4 rolls around (including features that are not in master or orion yet) it will be relatively stable and does not need its own "test effort". Are we at this nirvana yet? Not quite, but I think we are closer than ever before, and we have the chance to get there with a coordinated effort of the community. Cheers, Andreas -- Andreas Dilger Whamcloud, Inc. Principal Lustre Engineer http://www.whamcloud.com/ From morrone2 at llnl.gov Thu Jul 12 20:57:40 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Thu, 12 Jul 2012 13:57:40 -0700 Subject: [cdwg] broader Lustre testing In-Reply-To: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> References: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> Message-ID: <4FFF3A44.30106@llnl.gov> On 07/12/2012 12:37 PM, Nathan Rutman wrote: > > On Jul 12, 2012, at 7:30 AM, John Carrier wrote: > > A more strategic solution is to do more testing of a feature release > candidate _before_ it is released. Even if a Community member has no > interest in using a feature release in production, early testing with > pre-release versions of feature releases will help identify > instabilities created by the new feature with their workloads and > hardware before the release is official. > > > Taking a few threads that have discussed recently, regarding the stability of certain releases vs others, what maintenance branches are, what testing was done, and "which branch should I use": > These questions, I think, should not need to be asked. Which version of MacOS should I use? The latest one, period. Why can't Lustre do the same thing? Because we're an open source project where all of our dirty laundry is in the public. I'm sure that Apple has all kinds of internal deadlines and testing tags and things that we don't see on the outside world because it is a close-source proprietary product with vast resources to develop and test internally. The every-six month cadence is a good thing in my opinion. It forces us developers to regularly address the stability of the changes we are introducing. It provides a clear, explicit time in the schedule for developers to stop writing new bugs, and focus their effort on fixing bugs. I believe that the maintenance branch _is_ the place that you go when the question is "which version should I use"? We just need to have a decent web page that says "Want Lustre? Here's the latest stable release!" We need to increase exposure of the maintence releases, and hid the "feature" releases off on a developers page. > The answer I think lies in testing, which becomes a chicken and egg problem. I'm only going to use a "stable" release, which is the release which was tested with my applications. I know acceptance-small was run, and passed, on Master, otherwise it wouldn't be released. Hopefully it even ran on a big system like Hyperion. (Do we learn anything more about running acc-sm on other big systems? Probably not much.) But it certainly wasn't tested with my application, because I didn't test it. Because it wasn't released yet. Chicken and egg. Only after enough others make the leap am I willing to. > So, it seems, we need to test pre-release versions of Lustre, aka Master, with my applications. To that end, how willing are people to set aside a day, say once every two months, to be "filesystem beta day". Scientists, run your codes, users, do your normal work, but bear in mind there may be filesystem instabilities on that day. Make sure your data is backed up. Make sure it's not in the middle of a critical week-long run. Accept that you might have to re-run it tomorrow in the worst case. Report any problems you have. > What you get out of it is a much more stable Master, and an end to the question of "which version should I run". When released, you have confidence that you can move up, get the great new features and performance, and it runs your applications. More people are on the same release, so it sees even more testing. The maintenance branch is always the latest branch, you can pull in point releases with more bug fixes with ease. No more rolling your own Lustre with Frankenstein sets of patches. Latest and greatest and most stable. We can do a great deal more testing, and find a seriously large amount of bugs that we have been missing by getting more testing personnel allocated to Lustre. I think that's the major gap in Lustre right now. One day every two months is, I think, insufficient validating any software product, let alone something as complex as Lustre. Not that I am opposed to the idea. If you can arrange that, go for it! But that isn't good enough by itself by a long shot. We need full time personnel working on testing lustre. I would think that all of the vendors out there selling products to customers would already have alot of experience testing hardware, and other software bits. Lets apply some of that know-how to Lustre! And I think these testing personnel need to be made known to the community, so they can talk to each other, so that developers can guide their efforts, so we know what our testing converage looks like, etc. Testing needs to be a CONTINUAL process, not just something we do at the end for a specific release number. By the time we tag 2.4, it should already have been tested so frequently all along the master development cycle that the final testing will start to look like a formality to us. We should still do it, of course, but we should have confidence long before that happens. LLNL is trying to do that with the master branch as it moves to 2.4. Our coverage is mainly on zfs backends for now, but as the rest of orion lands on master, and Sequoia goes into limited production use we'll have both zfs and ldiskfs filesystems in our testbed, and test regularly all the way up to, and beyond, 2.4. The gaps in testing are NOT all an issue of insufficent scale testing, although there is admittedly a constant issue there. We need much better testing at small scale as well. And let me be really clear: when I say testing, I mean a real human being thinking up new tests all of the time. Looking at logs all of the time (so even when the test app succeeded, we'll catch the timeouts and reconnections and things that should not be happening, and are symptoms of bugs). Powering things off randomly. Literally pulling cables out while an evil, pathologically bad IO workload is running. We need real people to test all of the things that it is really easy for a human to do, and would take years for developers to automate with any reliability. The automated regression suite that we use is great. We should continue to improve that over time. But I would content that it is not, and never will be, sufficient to tells us if Lustre is stable. I would argue that the regressions tests are, in fact, a very low bar. And Lustre is just too complicated, networks are too complicated, we have too few developers, to ever come up with an automated suite with any thing but a relatively low confidence level in the stability of the software. And human testers are given a very different set of goals then developers. A developer's job is to make things work. A tester's is to do whatever they can to break it. And then create a good report of how they broke it so the developers can fix it. I also agree that I don't want to continue in this mode of "we'll only run it when LLNL/ORNL runs it and says its good". So we need more human testers. And to get back to the topic of making every single release a "stable" release: That ignores the fact that we have roughly a decade of seriously buggy, undocumented code that we're dealing with. It just will not happen. Period. We have to accept that and move forward. We can strive from this point on to make every release better than the last. But developers are human. Every time we add new features, we're going to add new bugs. We'll also fix bugs. But we're going to add new ones as well. So we deal with that by having "maintenance" releases. The maintenance release is maintained for a "long" period of time, but add NO new features. No new support for new kernels. No fantastic new performance improvements. Just bug fixes. The maintenance release is what vendors should build products upon, because that is where we'll land only bug fixes. So it is far more likely to only improve with time, whereas "master" (and therefore the "feature" releases which are just tags on master every 6 months), will also introduce destabilizing new features. We'll endevour to make the new features as stable as we are capable of doing, and we can do better if we have more testers, but we have to be pragmatic. "Every tag should be completely stable" is impossible. "Every tag on the maintenance branch should be more stable than the last" is an achievable goal. Chris From hamilton5 at llnl.gov Fri Jul 13 05:11:11 2012 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Thu, 12 Jul 2012 22:11:11 -0700 Subject: [cdwg] Lustre Development Community Tree Maintenance Q2 2012 report Message-ID: Hi all, Attached is the Q2 2012 quarterly report from Whamcloud for the Lustre Development Community Tree Maintenance contract. Please send any questions/comments to cdwg at lists.opensfs.org. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA  94551-9900 E-Mail:  pgh at llnl.gov Phone:  925-423-1332          Fax:  925-423-8719 -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS-Whamcloud Tree Report - Q2 2012 FINAL.pdf Type: application/pdf Size: 393998 bytes Desc: OpenSFS-Whamcloud Tree Report - Q2 2012 FINAL.pdf URL: From morrone2 at llnl.gov Fri Jul 13 17:59:28 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Fri, 13 Jul 2012 10:59:28 -0700 Subject: [cdwg] Lustre maintenance releases Message-ID: <50006200.1040803@llnl.gov> Hi folks, Attached is a graphic (lustre_releases.png) that I've made to try to illustrate the difference between "feature" and "maintenance" releases. Note that this doesn't fully represent every tag and branch in the repo, but this should be all that most end users and non-developer folks should likely need to concern themselves with. The horizontal arrow represents the "master" branch of lustre, on which we land both bug fixes and all new features. We follow a development cadence of 6 months between major tags on master. We've been calling those tags "feature" releases. Perhaps we need to change that terminology in the future to reduce confusion. The vertical arrows represent the "maintenance" branches. Code changes along the vertical branches should ONLY be bug fixes. This is the place to go if you are an institution or vendor that plans to have a system that needs to be stable for 18 months, 3 years, whatever, and you are not planning to do and significant OS upgrades. I know we have systems like that here at LLNL, and there are other systems out there that operate that way. The "maintenance" branch is where we point the user that just wants a stable lustre, and isn't interested in the latest bleeding-edge features. We are doing this today. The only part of my diagram that has not been decided is if we should officially declare that a new maintenance branch will begin on every third feature release. So the "18 months" is not official yet, and we don't yet know if "b2_4" will be the next maintenance branch. I would like us to decide on that rather soon, so we can all begin shifting our internal schedules to align with the 2.4 release in March 2013 (or whatever release we decide to designate). After that, an obvious question might be: "how long lived are these branches?". I am going to propose, to start the conversation, that we make the stated lifetime of the maintenance branches 3 years. That would make the timeline for stable maintenance branches look like the attached timeline graphic (lustre_maintenance_timeline.png). A new maintenance branch would start every 18 months, and live for 36 months. If we maintain our rule that a maintenance branch receives only bug fixes, after the first 18 months I suspect that a maintenance branch's required effort will have tapered off quite a bit, so hopefully overlapping the supported stable maintenance branches will not prove an unacceptable burden on the development community. This proposal would mean that at any point in time there are basically three major branches that the general public will be aware of: Current stable maintenance branch Previous stable maintenance branch Development branch (master) And frankly, most end users will only look at one of the maintenance branches. Anyone not already using lustre should always pick the current stable maintenance branch. Anyone running an older machine that has no plans to upgrade their OS will stick with the maintenance branch that they are already on. And of course, if you absolutely need some new feature in lustre, you can always pick up the latest feature release from master. But you do that knowing that your upgrade path is always horizontally on the master branch until you reach the next maintenance release that forks off vertically (as represented in the lustre_releases.png). Chris -------------- next part -------------- A non-text attachment was scrubbed... Name: lustre_releases.png Type: image/png Size: 30403 bytes Desc: not available URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: lustre_maintenance_timeline.png Type: image/png Size: 21416 bytes Desc: not available URL: From adilger at whamcloud.com Fri Jul 13 18:16:58 2012 From: adilger at whamcloud.com (Andreas Dilger) Date: Fri, 13 Jul 2012 12:16:58 -0600 Subject: [cdwg] Lustre maintenance releases In-Reply-To: <50006200.1040803@llnl.gov> References: <50006200.1040803@llnl.gov> Message-ID: On 2012-07-13, at 11:59 AM, Christopher J. Morrone wrote: > Attached is a graphic (lustre_releases.png) that I've made to try to illustrate the difference between "feature" and "maintenance" releases. Note that this doesn't fully represent every tag and branch in the repo, but this should be all that most end users and non-developer folks should likely need to concern themselves with. > > The horizontal arrow represents the "master" branch of lustre, on which we land both bug fixes and all new features. We follow a development cadence of 6 months between major tags on master. We've been calling those tags "feature" releases. Perhaps we need to change that terminology in the future to reduce confusion. > > The vertical arrows represent the "maintenance" branches. Code changes along the vertical branches should ONLY be bug fixes. This is the place to go if you are an institution or vendor that plans to have a system that needs to be stable for 18 months, 3 years, whatever, and you are not planning to do and significant OS upgrades. I know we have systems like that here at LLNL, and there are other systems out there that operate that way. > > The "maintenance" branch is where we point the user that just wants a stable lustre, and isn't interested in the latest bleeding-edge features. > > We are doing this today. The only part of my diagram that has not been decided is if we should officially declare that a new maintenance branch will begin on every third feature release. So the "18 months" is not official yet, and we don't yet know if "b2_4" will be the next maintenance branch. > > I would like us to decide on that rather soon, so we can all begin shifting our internal schedules to align with the 2.4 release in March 2013 (or whatever release we decide to designate). > > After that, an obvious question might be: "how long lived are these branches?". I am going to propose, to start the conversation, that we make the stated lifetime of the maintenance branches 3 years. That would make the timeline for stable maintenance branches look like the attached timeline graphic (lustre_maintenance_timeline.png). A new maintenance branch would start every 18 months, and live for 36 months. One minor note is that the dates on this second diagram are incorrect. It shows 2.1 starting in March 2013, when that is actually the start of 2.4. Everything needs to be back-dated 18 months so that 2.1 starts in Sept 2011, but is otherwise fine. > If we maintain our rule that a maintenance branch receives only bug fixes, after the first 18 months I suspect that a maintenance branch's required effort will have tapered off quite a bit, so hopefully overlapping the supported stable maintenance branches will not prove an unacceptable burden on the development community. > > This proposal would mean that at any point in time there are basically three major branches that the general public will be aware of: > > Current stable maintenance branch > Previous stable maintenance branch > Development branch (master) > > And frankly, most end users will only look at one of the maintenance branches. Anyone not already using lustre should always pick the current stable maintenance branch. Anyone running an older machine that has no plans to upgrade their OS will stick with the maintenance branch that they are already on. > > And of course, if you absolutely need some new feature in lustre, you can always pick up the latest feature release from master. But you do that knowing that your upgrade path is always horizontally on the master branch until you reach the next maintenance release that forks off vertically (as represented in the lustre_releases.png). > > Chris > > _______________________________________________ > cdwg mailing list > cdwg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org Cheers, Andreas -- Andreas Dilger Whamcloud, Inc. Principal Lustre Engineer http://www.whamcloud.com/ From morrone2 at llnl.gov Fri Jul 13 18:21:05 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Fri, 13 Jul 2012 11:21:05 -0700 Subject: [cdwg] Lustre maintenance releases In-Reply-To: <50006200.1040803@llnl.gov> References: <50006200.1040803@llnl.gov> Message-ID: <50006711.20800@llnl.gov> Sorry, I had my second timeline dates wrong. See attached lustre_maintenance_timeline2.png for the correction. On 07/13/2012 10:59 AM, Christopher J. Morrone wrote: > Hi folks, > > Attached is a graphic (lustre_releases.png) that I've made to try to > illustrate the difference between "feature" and "maintenance" releases. > Note that this doesn't fully represent every tag and branch in the > repo, but this should be all that most end users and non-developer folks > should likely need to concern themselves with. > > The horizontal arrow represents the "master" branch of lustre, on which > we land both bug fixes and all new features. We follow a development > cadence of 6 months between major tags on master. We've been calling > those tags "feature" releases. Perhaps we need to change that > terminology in the future to reduce confusion. > > The vertical arrows represent the "maintenance" branches. Code changes > along the vertical branches should ONLY be bug fixes. This is the place > to go if you are an institution or vendor that plans to have a system > that needs to be stable for 18 months, 3 years, whatever, and you are > not planning to do and significant OS upgrades. I know we have systems > like that here at LLNL, and there are other systems out there that > operate that way. > > The "maintenance" branch is where we point the user that just wants a > stable lustre, and isn't interested in the latest bleeding-edge features. > > We are doing this today. The only part of my diagram that has not been > decided is if we should officially declare that a new maintenance branch > will begin on every third feature release. So the "18 months" is not > official yet, and we don't yet know if "b2_4" will be the next > maintenance branch. > > I would like us to decide on that rather soon, so we can all begin > shifting our internal schedules to align with the 2.4 release in March > 2013 (or whatever release we decide to designate). > > After that, an obvious question might be: "how long lived are these > branches?". I am going to propose, to start the conversation, that we > make the stated lifetime of the maintenance branches 3 years. That > would make the timeline for stable maintenance branches look like the > attached timeline graphic (lustre_maintenance_timeline.png). A new > maintenance branch would start every 18 months, and live for 36 months. > > If we maintain our rule that a maintenance branch receives only bug > fixes, after the first 18 months I suspect that a maintenance branch's > required effort will have tapered off quite a bit, so hopefully > overlapping the supported stable maintenance branches will not prove an > unacceptable burden on the development community. > > This proposal would mean that at any point in time there are basically > three major branches that the general public will be aware of: > > Current stable maintenance branch > Previous stable maintenance branch > Development branch (master) > > And frankly, most end users will only look at one of the maintenance > branches. Anyone not already using lustre should always pick the > current stable maintenance branch. Anyone running an older machine that > has no plans to upgrade their OS will stick with the maintenance branch > that they are already on. > > And of course, if you absolutely need some new feature in lustre, you > can always pick up the latest feature release from master. But you do > that knowing that your upgrade path is always horizontally on the master > branch until you reach the next maintenance release that forks off > vertically (as represented in the lustre_releases.png). > > Chris > -------------- next part -------------- A non-text attachment was scrubbed... Name: lustre_maintenance_timeline2.png Type: image/png Size: 22785 bytes Desc: not available URL: From cliffw at whamcloud.com Thu Jul 12 20:25:00 2012 From: cliffw at whamcloud.com (Cliff White) Date: Thu, 12 Jul 2012 13:25:00 -0700 Subject: [cdwg] [Lustre-devel] broader Lustre testing In-Reply-To: References: Message-ID: There are certainly examples of this working for other products, for example (it's been a good number of years) at one time the main QA benchmark for the Oracle database was a customer-furnished test (the 'Churchill' test) which exercised the database throughly. It would also be useful to have data from those using standard IO tests (IOR, iozone, etc) as we could easily expand the existing tests with different parameter sets. However, in the HPC space, i suspect obtaining/generating the data set needed to replicate some customer situations would be a challenge. cliffw On Thu, Jul 12, 2012 at 1:07 PM, Nathan Rutman wrote: > > On Jul 12, 2012, at 12:48 PM, Bruce Korb wrote: > > Hi Nathan, > > > On 2012-07-12, at 20:37, Nathan Rutman wrote: > > > On Jul 12, 2012, at 7:30 AM, John Carrier wrote: > > A more strategic solution is to do more testing of a feature release > candidate _before_ it is released. Even if a Community member has no > interest in using a feature release in production, early testing with > pre-release versions of feature releases will help identify > instabilities created by the new feature with their workloads and > hardware before the release is official. > > > > Taking a few threads that have discussed recently, regarding the stability > of certain releases vs others, what maintenance branches are, what testing > was done, and "which branch should I use": > These questions, I think, should not need to be asked. Which version of > MacOS should I use? The latest one, period. Why can't Lustre do the same > thing? The answer I think lies in testing, which becomes a chicken and egg > problem. I'm only going to use a "stable" release, which is the release > which was tested *with my applications*. I know acceptance-small was > run, and passed, on Master, otherwise it wouldn't be released. Hopefully > it even ran on a big system like Hyperion. (Do we learn anything more > about running acc-sm on other big systems? Probably not much.) But it > certainly wasn't tested with my application, because I didn't test it. > Because it wasn't released yet. Chicken and egg. Only after enough > others make the leap am I willing to. > So, it seems, we need to test pre-release versions of Lustre, aka Master, > *with my applications*. To that end, how willing are people to set aside > a day, say once every two months, to be "filesystem beta day". Scientists, > run your codes, users, do your normal work, but bear in mind there may be > filesystem instabilities on that day. Make sure your data is backed up. > Make sure it's not in the middle of a critical week-long run. Accept that > you might have to re-run it tomorrow in the worst case. Report any > problems you have. > What you get out of it is a much more stable Master, and an end to the > question of "which version should I run". When released, you have > confidence that you can move up, get the great new features and > performance, and it runs your applications. More people are on the same > release, so it sees even more testing. The maintenance branch is always the > latest branch, you can pull in point releases with more bug fixes with > ease. No more rolling your own Lustre with Frankenstein sets of patches. > Latest and greatest and most stable. > > Pipe dream? > > > On Jul 12, 2012, at 12:48 PM, Bruce Korb wrote: > > > _I_ think so. You might get a few customers to say, "yes" but > never be able to find the appropriate round tuit. A more fruitful > approach might be to solicit customer acceptance tests. Presumably, > they've written them to hit the wrinkles that they tend to stub > their toes on. And there may be exceptions, too. (e.g. Cray might > well actually do some pre-testing -- they, too, have paying customers.) > > > I have no aversion to customers writing and supplying their own acceptance > tests, but I think that approach doesn't work for many of the cases: > - acceptance tests may not exist; acceptance may simply be testing with > large production codes > - tests that run in a particular environment need to be significantly > generalized > - tests may not be sharable for various legal reasons > > This also doesn't have to be an all-or-nothing proposition -- interested > parties will be able to use the latest features, and will help contribute > to the stability of Master, and will help reduce the "spread" of deployed > systems, in a positive feedback loop. > > Yes, absolutely, this is effort on the part of Lustre users. But it can > be balanced by the savings of efforts in roll-your-own, and risk reduction. > > > > _______________________________________________ > Lustre-devel mailing list > Lustre-devel at lists.lustre.org > http://lists.lustre.org/mailman/listinfo/lustre-devel > > -- cliffw Support Guy WhamCloud, Inc. www.whamcloud.com -------------- next part -------------- An HTML attachment was scrubbed... URL: From hamilton5 at llnl.gov Mon Jul 16 19:27:29 2012 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Mon, 16 Jul 2012 12:27:29 -0700 Subject: [cdwg] Lustre Support Survey Results Message-ID: Hi all, Attached are the results from the Lustre Support Survey conducted by the Community Development Working Group. Please send any comments/questions to cdwg at lists.opensfs.org. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA  94551-9900 E-Mail:  pgh at llnl.gov Phone:  925-423-1332          Fax:  925-423-8719 -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS Survey Results June 2012.pdf Type: application/pdf Size: 523924 bytes Desc: OpenSFS Survey Results June 2012.pdf URL: From carrier at cray.com Tue Jul 17 06:55:21 2012 From: carrier at cray.com (John Carrier) Date: Tue, 17 Jul 2012 06:55:21 +0000 Subject: [cdwg] Lustre maintenance releases In-Reply-To: <50006711.20800@llnl.gov> References: <50006200.1040803@llnl.gov> <50006711.20800@llnl.gov> Message-ID: Chris, A colleague asked if OpenSFS had considered the Linux kernel model of alternating even numbered feature releases with odd numbered stable maintenance releases. He said this made it easier for users to know when a version would be usable for production. If we apply this model to your proposal (18m maintenance releases, 3y total support for each maintenance branch), then we would need to reduce the number of feature releases to just one, 9months after the maintenance release. Is that too long between releases? Does this model prevent additions of new features in the odd numbered releases? Thanks, --jc -----Original Message----- From: cdwg-bounces at lists.opensfs.org [mailto:cdwg-bounces at lists.opensfs.org] On Behalf Of Christopher J. Morrone Sent: Friday, July 13, 2012 11:21 AM To: cdwg at lists.opensfs.org Subject: Re: [cdwg] Lustre maintenance releases Sorry, I had my second timeline dates wrong. See attached lustre_maintenance_timeline2.png for the correction. On 07/13/2012 10:59 AM, Christopher J. Morrone wrote: > Hi folks, > > Attached is a graphic (lustre_releases.png) that I've made to try to > illustrate the difference between "feature" and "maintenance" releases. > Note that this doesn't fully represent every tag and branch in the > repo, but this should be all that most end users and non-developer > folks should likely need to concern themselves with. > > The horizontal arrow represents the "master" branch of lustre, on > which we land both bug fixes and all new features. We follow a > development cadence of 6 months between major tags on master. We've > been calling those tags "feature" releases. Perhaps we need to change > that terminology in the future to reduce confusion. > > The vertical arrows represent the "maintenance" branches. Code > changes along the vertical branches should ONLY be bug fixes. This is > the place to go if you are an institution or vendor that plans to have > a system that needs to be stable for 18 months, 3 years, whatever, and > you are not planning to do and significant OS upgrades. I know we > have systems like that here at LLNL, and there are other systems out > there that operate that way. > > The "maintenance" branch is where we point the user that just wants a > stable lustre, and isn't interested in the latest bleeding-edge features. > > We are doing this today. The only part of my diagram that has not > been decided is if we should officially declare that a new maintenance > branch will begin on every third feature release. So the "18 months" > is not official yet, and we don't yet know if "b2_4" will be the next > maintenance branch. > > I would like us to decide on that rather soon, so we can all begin > shifting our internal schedules to align with the 2.4 release in March > 2013 (or whatever release we decide to designate). > > After that, an obvious question might be: "how long lived are these > branches?". I am going to propose, to start the conversation, that we > make the stated lifetime of the maintenance branches 3 years. That > would make the timeline for stable maintenance branches look like the > attached timeline graphic (lustre_maintenance_timeline.png). A new > maintenance branch would start every 18 months, and live for 36 months. > > If we maintain our rule that a maintenance branch receives only bug > fixes, after the first 18 months I suspect that a maintenance branch's > required effort will have tapered off quite a bit, so hopefully > overlapping the supported stable maintenance branches will not prove > an unacceptable burden on the development community. > > This proposal would mean that at any point in time there are basically > three major branches that the general public will be aware of: > > Current stable maintenance branch > Previous stable maintenance branch > Development branch (master) > > And frankly, most end users will only look at one of the maintenance > branches. Anyone not already using lustre should always pick the > current stable maintenance branch. Anyone running an older machine > that has no plans to upgrade their OS will stick with the maintenance > branch that they are already on. > > And of course, if you absolutely need some new feature in lustre, you > can always pick up the latest feature release from master. But you do > that knowing that your upgrade path is always horizontally on the > master branch until you reach the next maintenance release that forks > off vertically (as represented in the lustre_releases.png). > > Chris > From peter_bojanic at xyratex.com Tue Jul 17 11:42:33 2012 From: peter_bojanic at xyratex.com (Peter Bojanic) Date: Tue, 17 Jul 2012 08:42:33 -0300 Subject: [cdwg] Lustre maintenance releases In-Reply-To: References: <50006200.1040803@llnl.gov> <50006711.20800@llnl.gov> Message-ID: Back to the future! Lustre 1.2, 1.4, 1.6, and 1.8 were the stable production releases, reserving the odd numbers for development releases. The practice was decidedly dropped as of Lustre 2.0. Bojanic On 2012-07-17, at 03:55 , John Carrier wrote: > Chris, > > A colleague asked if OpenSFS had considered the Linux kernel model of alternating even numbered feature releases with odd numbered stable maintenance releases. He said this made it easier for users to know when a version would be usable for production. > > If we apply this model to your proposal (18m maintenance releases, 3y total support for each maintenance branch), then we would need to reduce the number of feature releases to just one, 9months after the maintenance release. Is that too long between releases? Does this model prevent additions of new features in the odd numbered releases? > > Thanks, > > --jc > > > -----Original Message----- > From: cdwg-bounces at lists.opensfs.org [mailto:cdwg-bounces at lists.opensfs.org] On Behalf Of Christopher J. Morrone > Sent: Friday, July 13, 2012 11:21 AM > To: cdwg at lists.opensfs.org > Subject: Re: [cdwg] Lustre maintenance releases > > Sorry, I had my second timeline dates wrong. See attached lustre_maintenance_timeline2.png for the correction. > > On 07/13/2012 10:59 AM, Christopher J. Morrone wrote: >> Hi folks, >> >> Attached is a graphic (lustre_releases.png) that I've made to try to >> illustrate the difference between "feature" and "maintenance" releases. >> Note that this doesn't fully represent every tag and branch in the >> repo, but this should be all that most end users and non-developer >> folks should likely need to concern themselves with. >> >> The horizontal arrow represents the "master" branch of lustre, on >> which we land both bug fixes and all new features. We follow a >> development cadence of 6 months between major tags on master. We've >> been calling those tags "feature" releases. Perhaps we need to change >> that terminology in the future to reduce confusion. >> >> The vertical arrows represent the "maintenance" branches. Code >> changes along the vertical branches should ONLY be bug fixes. This is >> the place to go if you are an institution or vendor that plans to have >> a system that needs to be stable for 18 months, 3 years, whatever, and >> you are not planning to do and significant OS upgrades. I know we >> have systems like that here at LLNL, and there are other systems out >> there that operate that way. >> >> The "maintenance" branch is where we point the user that just wants a >> stable lustre, and isn't interested in the latest bleeding-edge features. >> >> We are doing this today. The only part of my diagram that has not >> been decided is if we should officially declare that a new maintenance >> branch will begin on every third feature release. So the "18 months" >> is not official yet, and we don't yet know if "b2_4" will be the next >> maintenance branch. >> >> I would like us to decide on that rather soon, so we can all begin >> shifting our internal schedules to align with the 2.4 release in March >> 2013 (or whatever release we decide to designate). >> >> After that, an obvious question might be: "how long lived are these >> branches?". I am going to propose, to start the conversation, that we >> make the stated lifetime of the maintenance branches 3 years. That >> would make the timeline for stable maintenance branches look like the >> attached timeline graphic (lustre_maintenance_timeline.png). A new >> maintenance branch would start every 18 months, and live for 36 months. >> >> If we maintain our rule that a maintenance branch receives only bug >> fixes, after the first 18 months I suspect that a maintenance branch's >> required effort will have tapered off quite a bit, so hopefully >> overlapping the supported stable maintenance branches will not prove >> an unacceptable burden on the development community. >> >> This proposal would mean that at any point in time there are basically >> three major branches that the general public will be aware of: >> >> Current stable maintenance branch >> Previous stable maintenance branch >> Development branch (master) >> >> And frankly, most end users will only look at one of the maintenance >> branches. Anyone not already using lustre should always pick the >> current stable maintenance branch. Anyone running an older machine >> that has no plans to upgrade their OS will stick with the maintenance >> branch that they are already on. >> >> And of course, if you absolutely need some new feature in lustre, you >> can always pick up the latest feature release from master. But you do >> that knowing that your upgrade path is always horizontally on the >> master branch until you reach the next maintenance release that forks >> off vertically (as represented in the lustre_releases.png). >> >> Chris >> > > > _______________________________________________ > cdwg mailing list > cdwg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org From adilger at whamcloud.com Tue Jul 17 14:46:11 2012 From: adilger at whamcloud.com (Andreas Dilger) Date: Tue, 17 Jul 2012 07:46:11 -0700 Subject: [cdwg] Lustre maintenance releases In-Reply-To: References: <50006200.1040803@llnl.gov> <50006711.20800@llnl.gov> Message-ID: <45931C56-9AF2-4E6B-848F-CF5E1373CAFA@whamcloud.com> On 2012-07-17, at 4:42, Peter Bojanic wrote: > On 2012-07-17, at 03:55 , John Carrier wrote: >> A colleague asked if OpenSFS had considered the Linux kernel model of alternating even numbered feature releases with odd numbered stable maintenance releases. > > Back to the future! Lustre 1.2, 1.4, 1.6, and 1.8 were the stable production releases, reserving the odd numbers for development releases. The practice was decidedly dropped as of Lustre 2.0. I wasn't necessarily fond of that decision, but at the same time Linux has dropped the "even means stable" convention as well and is just using 3.1, 3.2, 3.3, ... release numbers, and some of them are picked up for longer term stable releases (usually in line with vendor kernels like 2.6.32), others not. >> If we apply this model to your proposal (18m maintenance releases, 3y total support for each maintenance branch), then we would need to reduce the number of feature releases to just one, 9months after the maintenance release. Is that too long between releases? Does this model prevent additions of new features in the odd numbered releases? There will still be new features included in each release (whether slated for long term maintenance or not). The difference is that after the designated maintenance release is made, it will continue to get bug fixes over time. I wouldn't object to going back to the even/odd convention if this helped people to distinguish between the releases. However, I don't think this should be a driver for when and how the releases are made. If we wanted 2.3 as a feature release, and 2.4 as a maintenance release it would work out fine, though we can't retroactively change 2.1. For later releases we could skip the even numbers until it was time for a maintenance release, like 2.5, 2.7, 2.8, then 2.9, 2.11, 2.12 or 3.0 (depending on how major the features are, e.g. exascale stuff). Cheers, Andreas >> -----Original Message----- >> From: cdwg-bounces at lists.opensfs.org [mailto:cdwg-bounces at lists.opensfs.org] On Behalf Of Christopher J. Morrone >> Sent: Friday, July 13, 2012 11:21 AM >> To: cdwg at lists.opensfs.org >> Subject: Re: [cdwg] Lustre maintenance releases >> >> Sorry, I had my second timeline dates wrong. See attached lustre_maintenance_timeline2.png for the correction. >> >> On 07/13/2012 10:59 AM, Christopher J. Morrone wrote: >>> Hi folks, >>> >>> Attached is a graphic (lustre_releases.png) that I've made to try to >>> illustrate the difference between "feature" and "maintenance" releases. >>> Note that this doesn't fully represent every tag and branch in the >>> repo, but this should be all that most end users and non-developer >>> folks should likely need to concern themselves with. >>> >>> The horizontal arrow represents the "master" branch of lustre, on >>> which we land both bug fixes and all new features. We follow a >>> development cadence of 6 months between major tags on master. We've >>> been calling those tags "feature" releases. Perhaps we need to change >>> that terminology in the future to reduce confusion. >>> >>> The vertical arrows represent the "maintenance" branches. Code >>> changes along the vertical branches should ONLY be bug fixes. This is >>> the place to go if you are an institution or vendor that plans to have >>> a system that needs to be stable for 18 months, 3 years, whatever, and >>> you are not planning to do and significant OS upgrades. I know we >>> have systems like that here at LLNL, and there are other systems out >>> there that operate that way. >>> >>> The "maintenance" branch is where we point the user that just wants a >>> stable lustre, and isn't interested in the latest bleeding-edge features. >>> >>> We are doing this today. The only part of my diagram that has not >>> been decided is if we should officially declare that a new maintenance >>> branch will begin on every third feature release. So the "18 months" >>> is not official yet, and we don't yet know if "b2_4" will be the next >>> maintenance branch. >>> >>> I would like us to decide on that rather soon, so we can all begin >>> shifting our internal schedules to align with the 2.4 release in March >>> 2013 (or whatever release we decide to designate). >>> >>> After that, an obvious question might be: "how long lived are these >>> branches?". I am going to propose, to start the conversation, that we >>> make the stated lifetime of the maintenance branches 3 years. That >>> would make the timeline for stable maintenance branches look like the >>> attached timeline graphic (lustre_maintenance_timeline.png). A new >>> maintenance branch would start every 18 months, and live for 36 months. >>> >>> If we maintain our rule that a maintenance branch receives only bug >>> fixes, after the first 18 months I suspect that a maintenance branch's >>> required effort will have tapered off quite a bit, so hopefully >>> overlapping the supported stable maintenance branches will not prove >>> an unacceptable burden on the development community. >>> >>> This proposal would mean that at any point in time there are basically >>> three major branches that the general public will be aware of: >>> >>> Current stable maintenance branch >>> Previous stable maintenance branch >>> Development branch (master) >>> >>> And frankly, most end users will only look at one of the maintenance >>> branches. Anyone not already using lustre should always pick the >>> current stable maintenance branch. Anyone running an older machine >>> that has no plans to upgrade their OS will stick with the maintenance >>> branch that they are already on. >>> >>> And of course, if you absolutely need some new feature in lustre, you >>> can always pick up the latest feature release from master. But you do >>> that knowing that your upgrade path is always horizontally on the >>> master branch until you reach the next maintenance release that forks >>> off vertically (as represented in the lustre_releases.png). >>> >>> Chris >>> >> >> >> _______________________________________________ >> cdwg mailing list >> cdwg at lists.opensfs.org >> http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org > > _______________________________________________ > cdwg mailing list > cdwg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org From morrone2 at llnl.gov Tue Jul 17 17:47:23 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Tue, 17 Jul 2012 10:47:23 -0700 Subject: [cdwg] Lustre maintenance releases In-Reply-To: <45931C56-9AF2-4E6B-848F-CF5E1373CAFA@whamcloud.com> References: <50006200.1040803@llnl.gov> <50006711.20800@llnl.gov> <45931C56-9AF2-4E6B-848F-CF5E1373CAFA@whamcloud.com> Message-ID: <5005A52B.7020105@llnl.gov> On 07/17/2012 07:46 AM, Andreas Dilger wrote: > On 2012-07-17, at 4:42, Peter Bojanic wrote: >> On 2012-07-17, at 03:55 , John Carrier wrote: >>> A colleague asked if OpenSFS had considered the Linux kernel model of alternating even numbered feature releases with odd numbered stable maintenance releases. >> > >> Back to the future! Lustre 1.2, 1.4, 1.6, and 1.8 were the stable production releases, reserving the odd numbers for development releases. The practice was decidedly dropped as of Lustre 2.0. > > I wasn't necessarily fond of that decision, but at the same time Linux has dropped the "even means stable" convention as well and is just using 3.1, 3.2, 3.3, ... release numbers, and some of them are picked up for longer term stable releases (usually in line with vendor kernels like 2.6.32), others not. > >>> If we apply this model to your proposal (18m maintenance releases, 3y total support for each maintenance branch), then we would need to reduce the number of feature releases to just one, 9months after the maintenance release. Is that too long between releases? Does this model prevent additions of new features in the odd numbered releases? > > There will still be new features included in each release (whether slated for long term maintenance or not). The difference is that after the designated maintenance release is made, it will continue to get bug fixes over time. > > I wouldn't object to going back to the even/odd convention if this helped people to distinguish between the releases. However, I don't think this should be a driver for when and how the releases are made. > > If we wanted 2.3 as a feature release, and 2.4 as a maintenance release it would work out fine, though we can't retroactively change 2.1. For later releases we could skip the even numbers until it was time for a maintenance release, like 2.5, 2.7, 2.8, then 2.9, 2.11, 2.12 or 3.0 (depending on how major the features are, e.g. exascale stuff). > > Cheers, Andreas I agree. I think there are lots of possibilities for making the numbering more understandable, and for better explaining them to the community. But we should make the numbering scheme fit a development model, not try to change development to fit a numbering scheme. On the question of making the feature-release cadence nine months, my feeling is that is a bit too long. Although I guess that opinion isn't very strong. I believe that the current six month cadence is working fairly well, and I haven't heard anyone suggest that the current six month schedule is too short, so I would tend to prefer that we leave that part of the process unchanged. Chris From nathan_rutman at xyratex.com Tue Jul 17 23:09:34 2012 From: nathan_rutman at xyratex.com (Nathan Rutman) Date: Tue, 17 Jul 2012 16:09:34 -0700 Subject: [cdwg] Lustre maintenance releases In-Reply-To: <5005A52B.7020105@llnl.gov> References: <50006200.1040803@llnl.gov> <50006711.20800@llnl.gov> <45931C56-9AF2-4E6B-848F-CF5E1373CAFA@whamcloud.com> <5005A52B.7020105@llnl.gov> Message-ID: <2FB884D5-F8AA-4332-9699-D583DA31EE15@xyratex.com> On Jul 17, 2012, at 10:47 AM, Christopher J. Morrone wrote: > On 07/17/2012 07:46 AM, Andreas Dilger wrote: >> On 2012-07-17, at 4:42, Peter Bojanic wrote: >>> On 2012-07-17, at 03:55 , John Carrier wrote: >>>> A colleague asked if OpenSFS had considered the Linux kernel model of alternating even numbered feature releases with odd numbered stable maintenance releases. >>> >> >>> Back to the future! Lustre 1.2, 1.4, 1.6, and 1.8 were the stable production releases, reserving the odd numbers for development releases. The practice was decidedly dropped as of Lustre 2.0. >> >> I wasn't necessarily fond of that decision, but at the same time Linux has dropped the "even means stable" convention as well and is just using 3.1, 3.2, 3.3, ... release numbers, and some of them are picked up for longer term stable releases (usually in line with vendor kernels like 2.6.32), others not. >> >>>> If we apply this model to your proposal (18m maintenance releases, 3y total support for each maintenance branch), then we would need to reduce the number of feature releases to just one, 9months after the maintenance release. Is that too long between releases? Does this model prevent additions of new features in the odd numbered releases? >> >> There will still be new features included in each release (whether slated for long term maintenance or not). The difference is that after the designated maintenance release is made, it will continue to get bug fixes over time. >> >> I wouldn't object to going back to the even/odd convention if this helped people to distinguish between the releases. However, I don't think this should be a driver for when and how the releases are made. >> >> If we wanted 2.3 as a feature release, and 2.4 as a maintenance release it would work out fine, though we can't retroactively change 2.1. For later releases we could skip the even numbers until it was time for a maintenance release, like 2.5, 2.7, 2.8, then 2.9, 2.11, 2.12 or 3.0 (depending on how major the features are, e.g. exascale stuff). >> >> Cheers, Andreas > > I agree. > > I think there are lots of possibilities for making the numbering more understandable, and for better explaining them to the community. But we should make the numbering scheme fit a development model, not try to change development to fit a numbering scheme. > > On the question of making the feature-release cadence nine months, my feeling is that is a bit too long. Although I guess that opinion isn't very strong. I believe that the current six month cadence is working fairly well, and I haven't heard anyone suggest that the current six month schedule is too short, so I would tend to prefer that we leave that part of the process unchanged. I agree, the cadence seems to be working and Peter Jones at WC is keeping things on track - if it ain't broke... From uja at ornl.gov Fri Jul 20 14:13:42 2012 From: uja at ornl.gov (James A Simmons) Date: Fri, 20 Jul 2012 10:13:42 -0400 Subject: [cdwg] broader Lustre testing In-Reply-To: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> References: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> Message-ID: <1342793622.7484.6.camel@bohr.ornl.gov> On Thu, 2012-07-12 at 15:37 -0400, Nathan Rutman wrote: > > On Jul 12, 2012, at 7:30 AM, John Carrier wrote: > > > A more strategic solution is to do more testing of a feature release > > candidate _before_ it is released. Even if a Community member has > > no > > interest in using a feature release in production, early testing > > with > > pre-release versions of feature releases will help identify > > instabilities created by the new feature with their workloads and > > hardware before the release is official. ... > So, it seems, we need to test pre-release versions of Lustre, aka > Master, with my applications. To that end, how willing are people to > set aside a day, say once every two months, to be "filesystem beta > day". Scientists, run your codes, users, do your normal work, but > bear in mind there may be filesystem instabilities on that day. Make > sure your data is backed up. Make sure it's not in the middle of a > critical week-long run. Accept that you might have to re-run it > tomorrow in the worst case. Report any problems you have. > What you get out of it is a much more stable Master, and an end to the > question of "which version should I run". When released, you have > confidence that you can move up, get the great new features and > performance, and it runs your applications. More people are on the > same release, so it sees even more testing. The maintenance branch is > always the latest branch, you can pull in point releases with more bug > fixes with ease. No more rolling your own Lustre with Frankenstein > sets of patches. Latest and greatest and most stable. > > > Pipe dream? Since people are now moving to help test out the current master branch for whamcloud I like to purpose posting a general summary of testing results people are seeing. I personally have finished a first run at testing 2.2.91 this last week and would galdly share the results. Anyone else can to share :-) From pjones at whamcloud.com Fri Jul 20 14:20:15 2012 From: pjones at whamcloud.com (Peter Jones) Date: Fri, 20 Jul 2012 07:20:15 -0700 Subject: [cdwg] broader Lustre testing In-Reply-To: <1342793622.7484.6.camel@bohr.ornl.gov> References: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> <1342793622.7484.6.camel@bohr.ornl.gov> Message-ID: <5009691F.6050709@whamcloud.com> James I would love to see the results. The key thing of course is that any anomalies you encounter are tracked by JIRA tickets. It would be a cherry on top if your test results could be uploaded into maloo. :-) Peter On 12-07-20 7:13 AM, James A Simmons wrote: > > Since people are now moving to help test out the current master branch > for whamcloud I like to purpose posting a general summary of testing > results people are seeing. I personally have finished a first run at > testing 2.2.91 this last week and would galdly share the results. Anyone > else can to share :-) > > > _______________________________________________ > cdwg mailing list > cdwg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org -- Peter Jones Whamcloud, Inc. www.whamcloud.com From pjones at whamcloud.com Fri Jul 20 14:38:48 2012 From: pjones at whamcloud.com (Peter Jones) Date: Fri, 20 Jul 2012 07:38:48 -0700 Subject: [cdwg] Lustre 2.3 update - July 20th 2012 Message-ID: <50096D78.3030802@whamcloud.com> Hi there Here is an update on the Lustre 2.3 release. Landings ======== -A number of landings made - see http://git.whamcloud.com/?p=fs/lustre-release.git;a=shortlog;h=refs/heads/master -2.6.38 client support (LU-506) landed -new IO engine (LU-1030) landed -Note that the feature freeze date (June 30th) is approaching Testing ======= -Testing on the 2.2.91 tag is in progress; testing on the 2.2.59 and 2.2.90 tag was completed -Now that we are feature frozen for 2.3, each subsequent tag should be converging on release quality. You can assist in the final 2.3 release being higher quality by testing these tags and reporting any issues found in JIRA. Blockers ======== -Full list available at http://jira.whamcloud.com/secure/IssueNavigator.jspa?mode=hide&requestId=10205 Other ===== -We will continue to provide periodic updates on our progress on this release. In the meantime, you can always see the landings as they happen at http://git.whamcloud.com/?p=fs/lustre-release.git;a=shortlog;h=refs/heads/master and follow the patch reviews and testing at http://review.whamcloud.com/#q,status:open+project:fs/lustre-release+branch:master,n,z Regards Peter -- Peter Jones Whamcloud, Inc. www.whamcloud.com -------------- next part -------------- An HTML attachment was scrubbed... URL: From nathan_rutman at xyratex.com Fri Jul 20 18:20:54 2012 From: nathan_rutman at xyratex.com (Nathan Rutman) Date: Fri, 20 Jul 2012 11:20:54 -0700 Subject: [cdwg] broader Lustre testing In-Reply-To: <1342793622.7484.6.camel@bohr.ornl.gov> References: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> <1342793622.7484.6.camel@bohr.ornl.gov> Message-ID: On Jul 20, 2012, at 7:13 AM, James A Simmons wrote: > On Thu, 2012-07-12 at 15:37 -0400, Nathan Rutman wrote: >> >> On Jul 12, 2012, at 7:30 AM, John Carrier wrote: >> >>> A more strategic solution is to do more testing of a feature release >>> candidate _before_ it is released. Even if a Community member has >>> no >>> interest in using a feature release in production, early testing >>> with >>> pre-release versions of feature releases will help identify >>> instabilities created by the new feature with their workloads and >>> hardware before the release is official. > > ... >> So, it seems, we need to test pre-release versions of Lustre, aka >> Master, with my applications. To that end, how willing are people to >> set aside a day, say once every two months, to be "filesystem beta >> day". Scientists, run your codes, users, do your normal work, but >> bear in mind there may be filesystem instabilities on that day. Make >> sure your data is backed up. Make sure it's not in the middle of a >> critical week-long run. Accept that you might have to re-run it >> tomorrow in the worst case. Report any problems you have. >> What you get out of it is a much more stable Master, and an end to the >> question of "which version should I run". When released, you have >> confidence that you can move up, get the great new features and >> performance, and it runs your applications. More people are on the >> same release, so it sees even more testing. The maintenance branch is >> always the latest branch, you can pull in point releases with more bug >> fixes with ease. No more rolling your own Lustre with Frankenstein >> sets of patches. Latest and greatest and most stable. >> >> >> Pipe dream? > > Since people are now moving to help test out the current master branch > for whamcloud I like to purpose posting a general summary of testing > results people are seeing. I personally have finished a first run at > testing 2.2.91 this last week and would galdly share the results. Anyone > else can to share :-) > > I started a page on the OpenSFS Wiki for everyone to share their test results in a free-form format. Note that the Wiki itself is still in it's infancy - I call on the community to help populate it. -------------- next part -------------- An HTML attachment was scrubbed... URL: From morrone2 at llnl.gov Tue Jul 24 16:44:53 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Tue, 24 Jul 2012 09:44:53 -0700 Subject: [cdwg] OpenSFS CDWG call tomorrow, 7/25, at 8:30am Pacific Message-ID: <500ED105.7050901@llnl.gov> Hello, We will be having the OpenSFS CDWG call tomorrow, July 25th, at 8:30am Pacific. On the agenda: Establish the process for selecting the next maintenance release, including a time line for decision making. Chris From quinn1 at llnl.gov Tue Jul 24 16:55:26 2012 From: quinn1 at llnl.gov (Quinn, Terri) Date: Tue, 24 Jul 2012 09:55:26 -0700 Subject: [cdwg] FW: Lustre Development Community Tree Maintenance Q2 2012 report In-Reply-To: References: Message-ID: <7F337A3E11411E4BB74DE2B8942F2913019DBC188F44@NSPEXMBX-D.the-lab.llnl.gov> CDWG, Perhaps someone could help me better understand this report. I have a few questions but I'll start with this one. What should I glean from the Quality Metrics included in the report on pages 4 and 5? What I see are the results of a number of tests. Some passed some did not. This is all understandable, but what does this say about the quality of the work? I might conclude that since I see less red as time goes on that things are getting better. However I also note that less tests are being performed. Regards, Terri Terri Quinn Principal Deputy Department Head Integrated Computing and Communications Lawrence Livermore National Laboratory quinn1 at llnl.gov cell 925 321-2879 office 925 423-2385 -----Original Message----- From: cdwg-bounces at lists.opensfs.org [mailto:cdwg-bounces at lists.opensfs.org] On Behalf Of Hamilton, Pam Sent: Thursday, July 12, 2012 10:11 PM To: 'cdwg at lists.opensfs.org'; OpenSFS Execs (Execs at lists.opensfs.org) Subject: [cdwg] Lustre Development Community Tree Maintenance Q2 2012 report Hi all, Attached is the Q2 2012 quarterly report from Whamcloud for the Lustre Development Community Tree Maintenance contract. Please send any questions/comments to cdwg at lists.opensfs.org. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA  94551-9900 E-Mail:  pgh at llnl.gov Phone:  925-423-1332          Fax:  925-423-8719 -------------- next part -------------- A non-text attachment was scrubbed... Name: OpenSFS-Whamcloud Tree Report - Q2 2012 FINAL.pdf Type: application/pdf Size: 393998 bytes Desc: OpenSFS-Whamcloud Tree Report - Q2 2012 FINAL.pdf URL: -------------- next part -------------- _______________________________________________ cdwg mailing list cdwg at lists.opensfs.org http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org From morrone2 at llnl.gov Tue Jul 24 17:51:49 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Tue, 24 Jul 2012 10:51:49 -0700 Subject: [cdwg] OpenSFS CDWG call tomorrow, 7/25, at 8:30am Pacific In-Reply-To: <500ED105.7050901@llnl.gov> References: <500ED105.7050901@llnl.gov> Message-ID: <500EE0B5.4090807@llnl.gov> I should have included call-in info: Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# On 07/24/2012 09:44 AM, Christopher J. Morrone wrote: > Hello, > > We will be having the OpenSFS CDWG call tomorrow, July 25th, at 8:30am > Pacific. > > On the agenda: > > Establish the process for selecting the next maintenance release, > including a time line for decision making. > > Chris > From spitzcor at cray.com Tue Jul 24 17:48:33 2012 From: spitzcor at cray.com (Cory Spitz) Date: Tue, 24 Jul 2012 12:48:33 -0500 Subject: [cdwg] OpenSFS CDWG call tomorrow, 7/25, at 8:30am Pacific In-Reply-To: <500ED105.7050901@llnl.gov> References: <500ED105.7050901@llnl.gov> Message-ID: <500EDFF1.80003@cray.com> Chris, I thought that all the CDWG calls were on Thursdays at 8:30 am PT. Will all future meetings be on Wednesdays too, or just tomorrow's? Or, did you really mean that we would meet on July 26? [I have a conflict tomorrow and might be late dialing in] Thanks, -Cory On 07/24/2012 11:44 AM, Christopher J. Morrone wrote: > Hello, > > We will be having the OpenSFS CDWG call tomorrow, July 25th, at 8:30am > Pacific. > > On the agenda: > > Establish the process for selecting the next maintenance release, > including a time line for decision making. > > Chris > _______________________________________________ > cdwg mailing list > cdwg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org From morrone2 at llnl.gov Tue Jul 24 18:18:17 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Tue, 24 Jul 2012 11:18:17 -0700 Subject: [cdwg] CORRECTION: OpenSFS CDWG call Thursday 7/26, at 8:30am Pacific In-Reply-To: <500EDFF1.80003@cray.com> References: <500ED105.7050901@llnl.gov> <500EDFF1.80003@cray.com> Message-ID: <500EE6E9.7090407@llnl.gov> Right, right. Sorry folks. Seriously this time, the call will be on Thursday at 8:30am Pacific: Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# On 07/24/2012 10:48 AM, Cory Spitz wrote: > Chris, > > I thought that all the CDWG calls were on Thursdays at 8:30 am PT. Will > all future meetings be on Wednesdays too, or just tomorrow's? Or, did > you really mean that we would meet on July 26? > > [I have a conflict tomorrow and might be late dialing in] > > Thanks, > -Cory > > On 07/24/2012 11:44 AM, Christopher J. Morrone wrote: >> Hello, >> >> We will be having the OpenSFS CDWG call tomorrow, July 25th, at 8:30am >> Pacific. >> >> On the agenda: >> >> Establish the process for selecting the next maintenance release, >> including a time line for decision making. >> >> Chris >> _______________________________________________ >> cdwg mailing list >> cdwg at lists.opensfs.org >> http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org > From spitzcor at cray.com Tue Jul 24 19:02:31 2012 From: spitzcor at cray.com (Cory Spitz) Date: Tue, 24 Jul 2012 14:02:31 -0500 Subject: [cdwg] Lustre maintenance releases In-Reply-To: <50006200.1040803@llnl.gov> References: <50006200.1040803@llnl.gov> Message-ID: <500EF147.2070700@cray.com> I like the direction of this thread and there are some good ideas for a long term model. Thanks, Chris, for making your proposal to the list. I promised on the last CDWG call that I would send mail to the list to clarify the idea that I had to get us to our future model. Apologies for the lateness. We seemed to reach a consensus that b2_4 ought to be designated as a maintenance branch after the initial 2.4 feature release. What I was trying to say over the phone was that if we do agree on that and if we nominally wait 3 months to spool up bugfixes to 2.4, it will arrive around Q2'13 if everything stays on-track. And that's over a year away from the original 2.1.x maintenance release and an awful lot of change for folks to take between 2.1.x and 2.4.1, which seemed like a concern for people on the call. I was only suggesting that we could create a 2.3.1 release as a quick stability catch-up for the feature stream and as a stepping stone to get to a future 2.4.1. We would have to state up-front that there would be no 2.3.2. It was merely a proposal to mitigate the risk of a big catch-up to 2.4.x. If we don't feel that it is worthwhile to address that issue, then I suggest that we get started on formalizing the long term maintenance branch and release model. To recap, here is what I was thinking for the maintenance stream: 2012 2013 Q1 Q2 Q3 Q4 Q1 Q2 Q3 Q4 2.1.1 -> 2.1.2 -> 2.1.3 2.2 2.3 -> 2.3.1 2.4 -> 2.4.1 -> 2.4.2 2.5 I think that the above agrees with what Chris has proposed for the short term, aside from the quick blip of 2.3.1. Thanks, -Cory On 07/13/2012 12:59 PM, Christopher J. Morrone wrote: > Hi folks, > > Attached is a graphic (lustre_releases.png) that I've made to try to > illustrate the difference between "feature" and "maintenance" releases. > Note that this doesn't fully represent every tag and branch in the > repo, but this should be all that most end users and non-developer folks > should likely need to concern themselves with. > > The horizontal arrow represents the "master" branch of lustre, on which > we land both bug fixes and all new features. We follow a development > cadence of 6 months between major tags on master. We've been calling > those tags "feature" releases. Perhaps we need to change that > terminology in the future to reduce confusion. > > The vertical arrows represent the "maintenance" branches. Code changes > along the vertical branches should ONLY be bug fixes. This is the place > to go if you are an institution or vendor that plans to have a system > that needs to be stable for 18 months, 3 years, whatever, and you are > not planning to do and significant OS upgrades. I know we have systems > like that here at LLNL, and there are other systems out there that > operate that way. > > The "maintenance" branch is where we point the user that just wants a > stable lustre, and isn't interested in the latest bleeding-edge features. > > We are doing this today. The only part of my diagram that has not been > decided is if we should officially declare that a new maintenance branch > will begin on every third feature release. So the "18 months" is not > official yet, and we don't yet know if "b2_4" will be the next > maintenance branch. > > I would like us to decide on that rather soon, so we can all begin > shifting our internal schedules to align with the 2.4 release in March > 2013 (or whatever release we decide to designate). > > After that, an obvious question might be: "how long lived are these > branches?". I am going to propose, to start the conversation, that we > make the stated lifetime of the maintenance branches 3 years. That > would make the timeline for stable maintenance branches look like the > attached timeline graphic (lustre_maintenance_timeline.png). A new > maintenance branch would start every 18 months, and live for 36 months. > > If we maintain our rule that a maintenance branch receives only bug > fixes, after the first 18 months I suspect that a maintenance branch's > required effort will have tapered off quite a bit, so hopefully > overlapping the supported stable maintenance branches will not prove an > unacceptable burden on the development community. > > This proposal would mean that at any point in time there are basically > three major branches that the general public will be aware of: > > Current stable maintenance branch > Previous stable maintenance branch > Development branch (master) > > And frankly, most end users will only look at one of the maintenance > branches. Anyone not already using lustre should always pick the > current stable maintenance branch. Anyone running an older machine that > has no plans to upgrade their OS will stick with the maintenance branch > that they are already on. > > And of course, if you absolutely need some new feature in lustre, you > can always pick up the latest feature release from master. But you do > that knowing that your upgrade path is always horizontally on the master > branch until you reach the next maintenance release that forks off > vertically (as represented in the lustre_releases.png). > > Chris > > > > _______________________________________________ > cdwg mailing list > cdwg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org From pjones at whamcloud.com Tue Jul 24 19:31:09 2012 From: pjones at whamcloud.com (Peter Jones) Date: Tue, 24 Jul 2012 12:31:09 -0700 Subject: [cdwg] FW: Lustre Development Community Tree Maintenance Q2 2012 report In-Reply-To: <7F337A3E11411E4BB74DE2B8942F2913019DBC188F44@NSPEXMBX-D.the-lab.llnl.gov> References: <7F337A3E11411E4BB74DE2B8942F2913019DBC188F44@NSPEXMBX-D.the-lab.llnl.gov> Message-ID: <500EF7FD.3050701@whamcloud.com> Hi Terri The chart you are referring to is an aggregated summary of all the testing results for master in our test results database (maloo) broken down by test suite and git tag (which we create ~ every two weeks). The numbers (x/y) show the number of test runs that passed with no failures (x) and the total number of test runs. Each test suite executes many tests. The number of tests run should increase over time as new tests are added in conjunction with bugfixes. Some test suites have been removed because they were deemed to be obsolete (e.g. liblustre) and some have been added (e.g. mds-survey) The best way to analyze this report is dynamically online at https://maloo.whamcloud.com/reports. There you are able to drill into the individual test runs and see what configuration was run and the details of any failures (with links to JIRA) Our goal is work through all the test failures (which are mostly due to either tests that were written with certain assumptions that are not reliably valid in our present-day testing environment or to certain configurations which the test environment does not fully support yet) so that all tests will show as green unless there is an actual bug introduced. As to how well this correlates to the quality of the work, that is open to debate. Personally, I think that, while I would not look at this data in isolation, I think that it has the advantage of being largely comparable to earlier Lustre releases and less affected by unrelated factors. For example, if we looked at the number of bugs reported against a given release, say, this could be affected by the the total number of sites running Lustre, or the introduction of new hardware that breaks former coding assumptions and so overall quality improvements could be offset by "noise" (or vice versa if people are slow to adopt newer releases and so do not discover and report bugs, say) I'll be happy to answer any further questions that you have either via email or on the upcoming CDWG call on Thursday. Regards Peter On 12-07-24 9:55 AM, Quinn, Terri wrote: > CDWG, > > Perhaps someone could help me better understand this report. I have a few questions but I'll start with this one. What should I glean from the Quality Metrics included in the report on pages 4 and 5? What I see are the results of a number of tests. Some passed some did not. This is all understandable, but what does this say about the quality of the work? I might conclude that since I see less red as time goes on that things are getting better. However I also note that less tests are being performed. > > Regards, Terri > > > Terri Quinn > Principal Deputy Department Head > Integrated Computing and Communications > Lawrence Livermore National Laboratory > quinn1 at llnl.gov > cell 925 321-2879 > office 925 423-2385 > > > -----Original Message----- > From: cdwg-bounces at lists.opensfs.org [mailto:cdwg-bounces at lists.opensfs.org] On Behalf Of Hamilton, Pam > Sent: Thursday, July 12, 2012 10:11 PM > To: 'cdwg at lists.opensfs.org'; OpenSFS Execs (Execs at lists.opensfs.org) > Subject: [cdwg] Lustre Development Community Tree Maintenance Q2 2012 report > > Hi all, > > Attached is the Q2 2012 quarterly report from Whamcloud for the Lustre Development Community Tree Maintenance contract. Please send any questions/comments to cdwg at lists.opensfs.org. > > Regards, > Pam > ___________________________________ > Pam Hamilton > Lawrence Livermore National Lab > P.O. Box 808, L-556 > Livermore, CA 94551-9900 > E-Mail: pgh at llnl.gov > Phone: 925-423-1332 Fax: 925-423-8719 > > > > > _______________________________________________ > cdwg mailing list > cdwg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org -- Peter Jones Whamcloud, Inc. www.whamcloud.com -------------- next part -------------- An HTML attachment was scrubbed... URL: From morrone2 at llnl.gov Tue Jul 24 22:10:04 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Tue, 24 Jul 2012 15:10:04 -0700 Subject: [cdwg] Lustre maintenance releases In-Reply-To: <500EF147.2070700@cray.com> References: <50006200.1040803@llnl.gov> <500EF147.2070700@cray.com> Message-ID: <500F1D3C.3080804@llnl.gov> On 07/24/2012 12:02 PM, Cory Spitz wrote: > We seemed to reach a consensus that b2_4 ought to be designated as a > maintenance branch after the initial 2.4 feature release. What I was > trying to say over the phone was that if we do agree on that and if we > nominally wait 3 months to spool up bugfixes to 2.4, it will arrive > around Q2'13 if everything stays on-track. I think that 2.4.0 at the end of March 2013 should be considered the first maintenance release on the maintenance branch, not 2.4.1. I would also argue that minor point releases should come as-needed instead of every three months, especially at the beginning of the branch's life. We may have 4 point releases in the first 3 months. And that's over a year away > from the original 2.1.x maintenance release and an awful lot of change The first 2.1 maintenance release in my mind was 2.1.0, and 2.4.0 would be exactly 18 months later. > To recap, here is what I was thinking for the maintenance stream: > > 2012 2013 > Q1 Q2 Q3 Q4 Q1 Q2 Q3 Q4 > 2.1.1 -> 2.1.2 -> 2.1.3 > 2.2 > 2.3 -> 2.3.1 > 2.4 -> 2.4.1 -> 2.4.2 > 2.5 > > I think that the above agrees with what Chris has proposed for the short > term, aside from the quick blip of 2.3.1. I'm not sure if it is or not. You seem to be combining all development, features, bug fixes, etc., into a single "maintenance stream". Maybe that isn't what you meant to imply, but I think that would be quite different than what I was suggesting. In the model I write about, maintenance branches progress in parallel with each other and with the development branch. But if I read the picture another way, assuming that the 2.1 and 2.3 don't dead-end, then I think you've just got too many maintenance branches. I don't think we really have the resources to have three active maintenance branches (2.1, 2.3, and 2.4) at the same time. Chris From spitzcor at cray.com Thu Jul 26 14:37:47 2012 From: spitzcor at cray.com (Cory Spitz) Date: Thu, 26 Jul 2012 09:37:47 -0500 Subject: [cdwg] Lustre maintenance releases In-Reply-To: <500F1D3C.3080804@llnl.gov> References: <50006200.1040803@llnl.gov> <500EF147.2070700@cray.com> <500F1D3C.3080804@llnl.gov> Message-ID: <5011563B.8010405@cray.com> Hi, Chris. My responses are snipped below. On 07/24/2012 05:10 PM, Christopher J. Morrone wrote: > > I think that 2.4.0 at the end of March 2013 should be considered the > first maintenance release on the maintenance branch, not 2.4.1. OK. Technically, by the current model, 2.4.0 is a "feature release". I agree that it would be the first release along the b2_4 maintenance branch. > > I would also argue that minor point releases should come as-needed > instead of every three months, especially at the beginning of the > branch's life. We may have 4 point releases in the first 3 months. I agree that we want to introduce some flexibility. By WC's current model, the release branches are tentatively scheduled quarterly. > I'm not sure if it is or not. You seem to be combining all development, > features, bug fixes, etc., into a single "maintenance stream". Maybe > that isn't what you meant to imply, but I think that would be quite > different than what I was suggesting. In the model I write about, > maintenance branches progress in parallel with each other and with the > development branch. That's not exactly what I was implying. I was not trying to close the door completely on parallel maintenance branches. But, I think that we should make it clear that people are encouraged to move to the next stable maintenance branch once it is created. That is, we should not plan for more 2.1.x releases after b2_4 is created. We should however leave the door open (just a crack) to allow us to fix anything egregious on the older release branches. I think it would be too much effort to keep multiple maintenance branches 'alive'. > > But if I read the picture another way, assuming that the 2.1 and 2.3 > don't dead-end, then I think you've just got too many maintenance > branches. I don't think we really have the resources to have three > active maintenance branches (2.1, 2.3, and 2.4) at the same time. > Agreed as stated above. That's why I think that we should be very clear about EOL for 2.3.1 if it were to exist. It would be very short lived; just until 2.4 was available. Thanks, -Cory From morrone2 at llnl.gov Fri Jul 27 19:03:04 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Fri, 27 Jul 2012 12:03:04 -0700 Subject: [cdwg] Summary of telecon on July 26, and review of information for the board Message-ID: <5012E5E8.7050400@llnl.gov> Hi folks, Here is a summary of the telecon from July 26 from my notes: Attendees: Chris Morrone - LLNL Peter Jones - Intel Cory Spitz - Cray Justin Miller - IU Mark Gary - LLNL James Simmons - ORNL John Hammond - TACC My apologies if I missed anyone. Some of the topics were: - Maintenance releases There was general agreement that we will all target our efforts towards making 2.4 the next maintenance branch. See note for the board at the end of this email for more information. Comments should be made by the end of Monday July 30th (lets say 4pm Pacific). - Conversation about testing (and how we need more) - Justin reported that the OpenSFS test cluster at IU is targeting August 1st to allow PAC contract testing to begin. Rough specs are: 24 compute nodes, 4 OSS nodes, 4 MDS nodes, IB connected, Westmere processors. They are working on a wiki that will document the hardware. * CDWG identified as likely group to schedule usage of the cluster. - Need for unified roadmap discussed. CDWG wants to take that on. We need to make sure that our roadmap is well communicated to our community and to our board to avoid us advertising conflicting roadmaps. And now my summary of the Lustre maintenance release plan as I understand it. As I mentioned early, please let me know if you think I've gotten anything wrong as soon as possible. I plan to relay this to the OpenSFS board by the end of the day on Monday. Summary of Lustre Maintenance Release Plan ------------------------------------------ Lustre maintenance branches host the releases of lustre that we shall advertise to the general public as "stable" releases. The goal of a maintenance branch is to include only bug fixes, to ensure a stable releases for a significant period of time. We currently have a development cadence that puts out a "feature" release of Lustre every 6 months. This is going reasonably well, and we plan to continue that process. However, we do not have the resources to ensure that every six month release is entirely bug-free, nor do we have the resources to add a new branch every six months that will receive only bug, and be maintained for years. We are not able to reasonably support and test that many branches in parallel at our current level of investment, nor would we wish to, as the testing requirements rise exponentially as the number of supported branches increases. We have decided that every third feature release, occurring every 18 months, will also be the beginning of a maintenance branch. Tagged maintenance releases along the maintenance branch will occur on an as-needed basis, according to the demands of discovered bugs and their severity. It is likely that tags will occur more frequently early in the branch's lifetime, and taper off in the later months and years. We plan to make the initially advertised lifetime of a maintenance branch three years. We agreed that Lustre 2.4 will begin the next maintenance branch, which has a targeted release date of the end of March, 2013. There was some discussion among folks who intend to use 2.3 in production that it would be nice to have a "mini-maintenance" branch to hold them over until 2.4 is released. Note that this branch would not be supported for a significant period of time, likely only a few month. This will happen if there is sufficient demand and resources applied, but to avoid confusion among the broader community we will not be advertising this release. We will also try to avoid using the term "maintenance" in association with 2.3 to avoid expectations of long term support if that branch becomes a reality. Chris From pjones at whamcloud.com Fri Jul 27 19:14:49 2012 From: pjones at whamcloud.com (Peter Jones) Date: Fri, 27 Jul 2012 12:14:49 -0700 Subject: [cdwg] Summary of telecon on July 26, and review of information for the board In-Reply-To: <5012E5E8.7050400@llnl.gov> References: <5012E5E8.7050400@llnl.gov> Message-ID: <5012E8A9.40302@whamcloud.com> Thanks for writing this up. I think that my only suggestion alteration would be with this section. I think that my characterization of this section would be that any additional maintenance release would be on the roadmap but clearly marked as an ad-hoc maintenance release rather than the codeline itself being the official maintenance release stream (with a succession of pre-scheduled releases). I think that the confusion here is that the expectation is that ad-hoc releases would occur as needed and likely not have much lead time, as opposed to releases on the maintenance release stream, which we would know about sometime ahead. On 12-07-27 12:03 PM, Christopher J. Morrone wrote: > > There was some discussion among folks who intend to use 2.3 in > production that it would be nice to have a "mini-maintenance" branch > to hold them over until 2.4 is released. Note that this branch would > not be supported for a significant period of time, likely only a few > month. This will happen if there is sufficient demand and resources > applied, but to avoid confusion among the broader community we will > not be advertising this release. We will also try to avoid using the > term "maintenance" in association with 2.3 to avoid expectations of > long term support if that branch becomes a reality. From morrone2 at llnl.gov Tue Jul 31 01:25:07 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Mon, 30 Jul 2012 18:25:07 -0700 Subject: [cdwg] Summary of telecon on July 26, and review of information for the board In-Reply-To: <5012E8A9.40302@whamcloud.com> References: <5012E5E8.7050400@llnl.gov> <5012E8A9.40302@whamcloud.com> Message-ID: <501733F3.2070400@llnl.gov> I am not opposed to ad-hoc maintenance releases, but I think it would be problematic to advertise the ad-hoc maintenance releases on a roadmap. We really want to avoid the same confusion we find ourselves in now. For instance, I would have thought that the difference between "maintenance release" and "feature release" would have been clear, but obviously it wasn't. Having two kinds of branches with "maintenance" in the name that have very different levels of support and community involvement planned is probably asking for trouble. But I'll reword my version to tone it down a bit, and work in the ad-hoc bit. On 07/27/2012 12:14 PM, Peter Jones wrote: > Thanks for writing this up. I think that my only suggestion alteration > would be with this section. I think that my characterization of this > section would be that any additional maintenance release would be on the > roadmap but clearly marked as an ad-hoc maintenance release rather than > the codeline itself being the official maintenance release stream (with > a succession of pre-scheduled releases). I think that the confusion > here is that the expectation is that ad-hoc releases would occur as > needed and likely not have much lead time, as opposed to releases on the > maintenance release stream, which we would know about sometime ahead. > > On 12-07-27 12:03 PM, Christopher J. Morrone wrote: >> > >> There was some discussion among folks who intend to use 2.3 in >> production that it would be nice to have a "mini-maintenance" branch >> to hold them over until 2.4 is released. Note that this branch would >> not be supported for a significant period of time, likely only a few >> month. This will happen if there is sufficient demand and resources >> applied, but to avoid confusion among the broader community we will >> not be advertising this release. We will also try to avoid using the >> term "maintenance" in association with 2.3 to avoid expectations of >> long term support if that branch becomes a reality. > > _______________________________________________ > cdwg mailing list > cdwg at lists.opensfs.org > http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org > From morrone2 at llnl.gov Tue Jul 31 01:46:44 2012 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Mon, 30 Jul 2012 18:46:44 -0700 Subject: [cdwg] Lustre Maintenance Release Plan In-Reply-To: <5012E5E8.7050400@llnl.gov> References: <5012E5E8.7050400@llnl.gov> Message-ID: <50173904.4060801@llnl.gov> Lustre Maintenance Release Plan ------------------------------- Lustre maintenance branches host the releases of lustre that we shall advertise to the general public as "stable" releases. The goal of a maintenance branch is to include only bug fixes, to ensure a stable releases for a significant period of time. We currently have a development cadence that puts out a "feature" release of Lustre every 6 months. This is going reasonably well, and we plan to continue that process. However, we do not have the resources to ensure that every six month release is entirely bug-free, nor do we have the resources to add a new branch every six months that will receive only bug, and be maintained for years. We are not able to reasonably support and test that many branches in parallel at our current level of investment, nor would we wish to, as the testing requirements rise exponentially as the number of supported branches increases. We have decided that every third feature release, occurring every 18 months, will also be the beginning of a maintenance branch. Tagged maintenance releases along the maintenance branch will occur on an as-needed basis, according to the demands of discovered bugs and their severity. It is likely that tags will occur more frequently early in the branch's lifetime, and taper off in the later months and years. We plan to make the initially advertised lifetime of a maintenance branch three years. We agreed that Lustre 2.4 will begin the next maintenance branch, which has a targeted release date of the end of March, 2013. None of this precludes ad-hoc maintenance releases on other branches. For instance, there was discussion in the CDWG that some institutions plan to run Lustre 2.3, and may need a "mini-maintenance" release to hold them over until Lustre 2.4. We just note that those "mini-maintenance" releases will be made as the participants have the time and desire to work on them, and that they will be made on an as-needed, and ad-hoc manner. We will continue to refine our terminology to help avoid confusion. To reiterate, out basic plan of record is: - Next official maintenance branch will be 2.4 (end of March 2013) - Official maintenance branches will start every 18 months - Official maintenance branches will be supported for 3 years Chris From Roman_Grigoryev at xyratex.com Mon Jul 16 18:31:32 2012 From: Roman_Grigoryev at xyratex.com (Roman Grigoryev) Date: Mon, 16 Jul 2012 11:31:32 -0700 Subject: [cdwg] [Lustre-devel] broader Lustre testing In-Reply-To: <4FFF3A44.30106@llnl.gov> References: <97EB7132-D1FA-4BFF-9EDC-9AEA4D1807E7@xyratex.com> <4FFF3A44.30106@llnl.gov> Message-ID: Hi Christopher, ..... > > The automated regression suite that we use is great. We should continue > to improve that over time. But I would content that it is not, and > never will be, sufficient to tells us if Lustre is stable. > > I would argue that the regressions tests are, in fact, a very low bar. > And Lustre is just too complicated, networks are too complicated, we > have too few developers, to ever come up with an automated suite with > any thing but a relatively low confidence level in the stability of the > software. > > And human testers are given a very different set of goals then > developers. A developer's job is to make things work. A tester's is to > do whatever they can to break it. And then create a good report of how > they broke it so the developers can fix it. > ............. Just for proving your statement that it is not enough just execute automated regression suite (acc-small) for testing quality I would like to share coverage summary which we got: 958 tests was executed Hit Total Coverage Lines: 79691 128935 61.8 % Functions: 6206 7935 78.2 % Branches: 49287 113914 43.3 % Thanks, Roman From hamilton5 at llnl.gov Tue Jul 31 16:55:48 2012 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Tue, 31 Jul 2012 09:55:48 -0700 Subject: [cdwg] Canceled: OpenSFS Community Development WG call Message-ID: When: Occurs every 2 weeks on Wednesday effective 6/8/2011 from 9:00 AM to 10:00 AM (UTC-08:00) Pacific Time (US & Canada). Where: Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# Note: The GMT offset above does not reflect daylight saving time adjustments. *~*~*~*~*~*~*~*~*~* Hi all, I’m cancelling this invitation in advance of sending a new one which will be based just on the cdwg mail list. General information about the mailing list including how to subscribe is at: http://lists.opensfs.org/listinfo.cgi/cdwg-opensfs.org Regards, Pam -------------- next part -------------- An HTML attachment was scrubbed... URL: -------------- next part -------------- A non-text attachment was scrubbed... Name: not available Type: text/calendar Size: 6173 bytes Desc: not available URL: From hamilton5 at llnl.gov Tue Jul 31 21:22:35 2012 From: hamilton5 at llnl.gov (Hamilton, Pam) Date: Tue, 31 Jul 2012 14:22:35 -0700 Subject: [cdwg] OpenSFS CDWG conference call schedule Message-ID: Hi all, I'd like to do a reset on the OpenSFS Community Development Working Group bi-weekly conference call day/time. Originally, it was scheduled for alternate Wednesdays at 9am Pacific. This worked for me as long as the Wednesday was the opposite week from another bi-weekly meeting I have at 9:30am Pacific. More recently due to requests for a more timely working group call than I could accommodate because of my other Wednesday commitment, there were conference calls held on two alternate Thursdays at 8:30am Pacific. Here is my survey, please respond with your preference: 1) Next working group call will be Thursday, 8/9/12, at 8:30am Pacific (and bi-weekly from then on) OR 2) Next working group call will be Wednesday, 8/15/12, at 9am Pacific (and bi-weekly from then on) FWIW, I had avoided Thursdays in the past because of not wanting to conflict with the TWG calls which are often held at 9:30am Pacific on Thursdays. Please send me your preference by Monday, August 6th. Regards, Pam ___________________________________ Pam Hamilton Lawrence Livermore National Lab P.O. Box 808, L-556 Livermore, CA  94551-9900 E-Mail:  pgh at llnl.gov Phone:  925-423-1332          Fax:  925-423-8719