From ssimms at iu.edu Wed May 1 03:51:32 2013 From: ssimms at iu.edu (ssimms at iu.edu) Date: Tue, 30 Apr 2013 23:51:32 -0400 (EDT) Subject: [cdwg] LUG 2013 Presentations Message-ID: Greetings- This year's LUG was a panoply of perspicacious presentations. All alliteration aside, we were very lucky to have such a terrific set of presenters and presentations. If you haven't found them on the OpenSFS site yet, I want to call your attention to the meeting's slide decks: http://www.opensfs.org/resources/presentations/ Download and enjoy! As part of initial planning for LUG2014, OpenSFS expects to get a survey out on LUG2013 soon. We look forward to getting your feedback and input. Thank you all for helping make LUG a success! Sincerely, Stephen Simms Manager, High Performance File Systems Indiana University 812-855-7211 From andreas.dilger at intel.com Thu May 2 05:57:12 2013 From: andreas.dilger at intel.com (Dilger, Andreas) Date: Thu, 2 May 2013 05:57:12 +0000 Subject: [cdwg] Lustre 2.4 update - April 26th 2013 In-Reply-To: Message-ID: On 2013/30/04 10:40 AM, "Cory Spitz" wrote: >Also, is it reasonable to use the new e2fsprogs with 1.8.9-wc1 too? [how >about earlier 1.8.x releases?] We haven't made a 1.8.x release with the new e2fsprogs, so this hasn't been tested very much. That said, the only potential problems would be in the lfsck code. There were changes in the upstream e2fsprogs, which are always listed in the RELEASE-NOTES file in the source code, or at the e2fsprogs site: http://e2fsprogs.sourceforge.net/e2fsprogs-release.html#1.42.7 The only Lustre-specific fixes since 1.42.6.wc2 are related to building against the newer 2.4 header files, and accommodating some of the changes to the on-disk format (LU-2677), as well as some cleanup and code reorganization. Cheers, Andreas >On 4/30/13 11:30 AM, "Cory Spitz" wrote: > >>Hello, >> >>Is there more detail about the new stable e2fsprogs? Is there a README >>or >>published changelog for e2fsprogs like there is for Lustre? >> >>Nevertheless, Cray will restart 2.4 testing with the new e2fsprogs >>package. >> >>Thanks, >>-Cory >> >> >>On 4/26/13 6:23 PM, "Jones, Peter A" wrote: >> >>>Hi there >>> >>>Here is an update on the Lustre 2.4 release. >>> >>>Landings >>>======== >>> >>>-A number of landings made - see >>>http://git.whamcloud.com/?p=fs/lustre-release.git;a=shortlog;h=refs/head >>>s >>>/ >>>master >>> >>>Testing >>>======= >>> >>>-Testing on the 2.3.64 tag is drawing to a close; a new tag is >>>anticipated early next week >>> >>>Blockers >>>======== >>> >>>-Full list available at https://jira.hpdd.intel.com/issues/?filter=10292 >>>-If there are any issues not presently marked as blockers that you >>>believe should be, please let me know >>> >>>Other >>>===== >>>-New version of e2fsprogs (1.42.7-wc1) released for compatibility with >>>2.4 features >>>-We are in the stabilization period for the release now so this is an >>>ideal time for community members to test tags and open JIRA tickets for >>>any issues encountered >>> >>>Thanks >>> >>>Peter Cheers, Andreas -- Andreas Dilger Lustre Software Architect Intel High Performance Data Division From uja at ornl.gov Thu May 2 19:04:08 2013 From: uja at ornl.gov (James A Simmons) Date: Thu, 02 May 2013 15:04:08 -0400 Subject: [cdwg] ORNL test shot results 4-12-2013 Message-ID: <1367521448.1869.59.camel@bohr.ornl.gov> In the interest to help out the community we have decided it is best to share our test results with the pre Lustre releases after we have our test shot. This is the results from the April 13th test shot. OLCF Portion Lustre 2.4 Testing Information After the last test shot on March 18th several Lustre software bugs were reported. At the smaller scale of the single cage Arthur Cray test bed we managed to replicate those issues. With the recent code base, the large stripe issue appeared to have been resolved, so it was decided to include the large stripe count test set. As has been the case DVS is still not functional with the current LNET version in the 2.3.63 code branch. Currently we are in contact with the Cray DVS developer to resolve this issue as soon as possible. The test shot plan was laid out into the standard three phases that have been used in previous test shots. The first phase was ORNL to test the special 2.3.63 (pre 2.4) Cray clients; the second phase was to have Cray test the 2.3.63 clients with Cray's I/O stress suite; the third phase was ORNL to perform a retesting of the first phase with the default (Cray) 1.8.6 clients. For all three phases the back end storage ran Lustre 2.3.63 on an RHEL6.3 image. The ORNL test set consisted of a S3D run, mdtest benchmarking, and several combinations of IOR runs. For all intents and purposes, ORNL tests were targeted for evaluating the Lustre 2.3.63 functionality on large scale. The file system start up process began at 12:30pm after the bulk of the Maintenance activities were completed by the HPC Operations Infrastructure team. The file system start up process encountered an oops on mounting the MDT. This was from the recent FID format change that was landed in the Lustre code, which required the test file system to reformat. After this detail was revealed, the file system was reformatted. By 3:30 the file system was ready to mount but one OST failed to mount which was traced to an Infiniband Network issue where the Subnet Manger was not allowing new IPoIB connections because of an error in the fabric. The resolution was to restart the Subnet Manger and then the file system was then successfully mounted at 4:30. Next the file system was to taken down to set the fail over node-pairs. In the previous test shot the setting of the file over nids did not work, but this time it was successful using the workaround provided after the last test shot. The fail over pairs were set on the widow-oss5a[1-4]. The file system was successfully mounted on a test node at 5:37. Titan rebooted and mounted the Lustre 2.4 file system at 7:32 PM. Once Titan was successfully booted, issues with the job scheduler were discovered due to the machine being reassembled into 200 cabinets just before the test time. By 9:50PM the job scheduler issues were resolved and we began testing. The first job we ran was the IOR hero run for large stripe. Just like the test shot before the job failed. Again the same job was launched, this time with full debugging turned on. The Lustre debug daemons running on the servers were unable to keep up with the messages generated, but a large volume of logs was collected up to the job failure. The logs on server and client side were then collected on a management server. Once the transfer of the log files was completed, and the debug daemon was disabled, we moved on to the IR recovery test starting at 12:21 AM. To test this setting we ran the hero ssf test run and then powered off widow-oss5a2. The oss recovered successfully in 3:58. We saw a few soft locks, but nothing that prevented recovery from success. Soon after recovery, we saw the single shared file fail to ENOSPC like the wide stripe testing. We began to think it was due to grant space, but latter discussion with Intel engineers suggest this is not the case. After this test, we continued with the final work load of S3, mdtest and IOR scaling. These jobs ran to completion with no problems. CRAY Portion High level summary of the Cray test time on Titan running against Lustre 2.4 client and server on 4/13/2013. Originally the test time was to start at 8pm CST on Friday 4/12, this changed early in the day to Midnight. Time line of events: ~12:30am Got access to the system. John Lewis reported there was an issue with file system, some jobs getting No space errors, he wanted to know how much that would affect my run. I started up a work load for about ~15-30 minutes across the entire machine. I then looked to see how many of the jobs were getting errors. I found about 8% of jobs at that time got the file system errors. John and I determined this was good enough for me to continue with testing. ~12:50am Started the first real workload on the machine. This load consisted of 1regression stream (looking for correctness) and 4 other streams each of which had a different random core count. All together there was more than enough to keep the system full. The job core counts were in the thousands to multiple thousands for each job. The regression stream contains all of our I/O tests plus other functional tests. The 4 other streams had the MDS intensive tests disabled so they would not run. I wanted to focus this part of the ORNL session away from beating only on the MDS. ~1:30am Checked the run, system was loaded. File system was fairly responsive. ~2:30am Checked the system again, most compute nodes were being used. File system was responsive. ~4:30am Again checked the system, most compute nodes were utilized and the file system was a bit sluggish but not bad enough to worry about. ~9:15am Killed off the current jobs and then changed the workload on the machine. Instead of using all the compute nodes, I focused more on a smaller subset about 1/4 of the machine with more smaller jobs running. Each job would be below 1000 cores, still the same threads running and steering away from the MDS intensive tests. ~9:50am Having issues getting the new workload started, few jobs are running. Operations in the Lustre tree are taking tens of minutes. /lustre/routed1/scratch/darason ~10:00-10:30 John looked around to see what was happening. Commands like /bin/ls and mkdir are taking 15+ minutes to /lustre/routed1/scratch/darason John called for the “Server side Calvalry” 11:16am John said they were going to setup to crash dump the MDS 11:20am Time is up. Over all the time went well. Only the one issue was uncovered. I’ve looked over the test results and besides the “No Space” issue hitting about 10% of the jobs executed, I did not see any other issues. ORNL Portion Lustre 1.8 Compatibility Testing Information The first job we ran was the IOR hero run for large stripe. We experienced no problems when running the 1.8 client against our 2.3.64 servers. After the large stripe test was completed we started our next hero run using file per process. This job completed successfully also. After this test, we continued with the final work load of S3, mdtest and IOR scaling. These jobs ran to completion with no problems either. All jobs completed except for the IOR scaling job but that was due to time constraints which required that we kill the job. Once all testing was complete, Titan was shut down to run pre-acceptance diagnostics. The file system was successfully rebooted into the production and returned to service at 12:05am on 4/14/2013. From board at lists.opensfs.org Fri May 3 15:47:42 2013 From: board at lists.opensfs.org (OpenSFS Board of Directors) Date: Fri, 3 May 2013 08:47:42 -0700 Subject: [cdwg] OpenSFS Management Firm Message-ID: <002801ce4815$8c40c090$a4c241b0$@lists.opensfs.org> The Board of OpenSFS is pleased to announce VTM Group as the new management firm for OpenSFS. VTM will support all operations for OpenSFS, including Board, Working Groups, financial management, industry events, LUG, membership, communications, marketing, social media, etc. VTM is now actively engaged in taking on OpenSFS operations, so you will soon be seeing them in various OpenSFS activities. We expect the full transition to be completed by the end of May. Today we are happy to introduce the following individuals from the VTM team supporting OpenSFS: - Jen Franklin, OpenSFS Account Manager and primary point of contact - Molly Nelson, Event Manager - Casey Robins, Senior Accounting Specialist You can reach any member of the VTM team by contacting admin at opensfs.org or calling the OpenSFS Administration office at 503-619-0561. VTM has been providing support to groups like OpenSFS for nearly 20 years. A sampling of their clients includes: Ethernet Alliance, InfiniBand Trade Association, Open Virtualization Alliance, Open Data Center Alliance, The Green Grid, and many more . We look forward to all the benefits of VTM's knowledge and experience as we continue to expand and grow the success of OpenSFS. If you have any questions, please feel free to contact the Board at: board at lists.opensfs.org. Sincerely, OpenSFS Board of Directors -------------- next part -------------- An HTML attachment was scrubbed... URL: From morrone2 at llnl.gov Wed May 8 00:37:24 2013 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Tue, 07 May 2013 17:37:24 -0700 Subject: [cdwg] OpenSFS CDWG Call Message-ID: <51899E44.908@llnl.gov> Hi all, Just a reminder that we will have a Community Development Working Group (CDWG) call on Wednesday, May 8, at 9:00am PDT. Agenda: * 2.4 status updates * Chris's early thoughts on Tree Contract improvement process * Open discussion Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# Regards, Chris From peter.a.jones at intel.com Fri May 10 23:26:46 2013 From: peter.a.jones at intel.com (Jones, Peter A) Date: Fri, 10 May 2013 23:26:46 +0000 Subject: [cdwg] Lustre 2.4 update - May 10th 2013 Message-ID: Hi there Here is an update on the Lustre 2.4 release. Landings ======== -A number of landings made - see http://git.whamcloud.com/?p=fs/lustre-release.git;a=shortlog;h=refs/heads/master -Landed support for RHEL6.4 (LU-3216) -moved to more current version of ZFS (LU-3117) Testing ======= -Testing on the 2.3.65 tag is underway Blockers ======== -Full list available at https://jira.hpdd.intel.com/issues/?filter=10292 -If there are any issues not presently marked as blockers that you believe should be, please let me know Other ===== -We are in the stabilization period for the release now so this is an ideal time for community members to test tags and open JIRA tickets for any issues encountered Thanks Peter From morrone2 at llnl.gov Wed May 22 00:07:39 2013 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Tue, 21 May 2013 17:07:39 -0700 Subject: [cdwg] OpenSFS CDWG Call Message-ID: <519C0C4B.6040801@llnl.gov> Hi all, Just a reminder that we will have a Community Development Working Group (CDWG) call on Wednesday, May 22, at 9:00am PDT. Agenda: * 2.4 status updates * 2.5 early planning * Tree Contract (pick up on previous meeting's discussion) Call-in: 866-914-3976 (925-424-8105) Passcode: 534986# Regards, Chris From morrone2 at llnl.gov Wed May 22 17:41:53 2013 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Wed, 22 May 2013 10:41:53 -0700 Subject: [cdwg] Lustre 2.5 Development Planning Message-ID: <519D0361.5010805@llnl.gov> Hi folks, Please send us your list of planned development items for Lustre 2.5! In two weeks at the next CDWG meeting we will be looking at the list of proposed work and making our first, best attempt at pare it down to a reasonable set of work for Lustre 2.5. If you have a new feature that you believe will be ready to land with in the three month feature-landing window of Lustre 2.5, NOW is the time to tell us! If you have pet bugs that are stuck on the back burner and need to be made a development priority for Lustre 2.5, NOW is the time to tell us! Your input is crucial if we are to properly scope the amount of change that we will introduce for Lustre 2.5, and properly scoping the work is a requirement if we are to produce a solid, on-time software release. So don't delay! Send us your list of desired Lustre 2.5 development items to this mailing list as soon as possible. Then, join us in the next CDWG call on June 5th to help us choose the subset of work that we believe is achievable* for Lustre 2.5. Chris * Not that this list will only be our best attempt at scoping what is reasonable, and is by no means a guarantee that any particular item will appear in Lustre 2.5. From morrone2 at llnl.gov Wed May 22 17:53:33 2013 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Wed, 22 May 2013 10:53:33 -0700 Subject: [cdwg] Make Lustre 2.5 a maintenance release? Message-ID: <519D061D.5050407@llnl.gov> Lustre community: As we all know, HSM is a Lustre feature that is desired by many, but was not fully completed in time for the Lustre 2.4 release. As a result, many folks are wondering if Lustre 2.4 is really the best branch to support for a long period of time (18 or more months). HSM is very likely to be completed for 2.5, and multiple vendors and organizations are planning to move to that version, and desire long term support for that branch. On the other hand, we have for some time been advertising quite widely that Lustre 2.4 will begin the next significant maintenance branch, and some organizations have built there plans around that guidance. I would like to officially propose the following action: 1) Announce that Lustre 2.4 will have maintenance releases as planned, but the maintenance window has been shortened to just 6 months. In other words, we will only expect to see maintenance releases on Lustre 2.4 branch for roughly 6 months. 2) Announce that the Lustre 2.5 branch will be long term maintenance branch. Are there any objections? Any adjustments that should be made to that statement? Chris From morrone2 at llnl.gov Wed May 22 18:20:38 2013 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Wed, 22 May 2013 11:20:38 -0700 Subject: [cdwg] Plea for contract requirements for maintenance branch(es) Message-ID: <519D0C76.2010808@llnl.gov> Hi folks, Work on updating the Lustre Development Community Tree Maintenance contract (SFS-MAINT-001) is slow going. One of sticking points for me is the topic of maintenance branches. A couple of organizations are very interested in extending the contract to cover maintenance branches. I need those of you with interest in doing so to provide me with the contractual requirements that you would like to see added. Without requirements, I have nothing to put into writing, and without that we can't really begin negotiating with Intel over the changes. If this is an area that you feel strongly about, please help propose concrete contract changes that you would like to see! Chris From oleg.drokin at intel.com Fri May 24 20:54:11 2013 From: oleg.drokin at intel.com (Drokin, Oleg) Date: Fri, 24 May 2013 20:54:11 +0000 Subject: [cdwg] new b2_4 tag 2.4.0-RC2 Message-ID: Hello! Just wanted to let everybody know that I tagged 2.4.0-RC2 in b2_4 branch, so those of you who are doing 2.4.0 pre-release testing, please use that tag from now on. Thanks. Bye, Oleg From peter.a.jones at intel.com Sat May 25 00:16:42 2013 From: peter.a.jones at intel.com (Jones, Peter A) Date: Sat, 25 May 2013 00:16:42 +0000 Subject: [cdwg] Lustre 2.4 update - May 24th 2013 Message-ID: Hi there Here is an update on the Lustre 2.4 release. Landings ======== -A number of landings made - see http://git.whamcloud.com/?p=fs/lustre-release.git;a=shortlog;h=refs/heads/b2_4 -Landed support for RHEL6.4 update (LU-3354); this kernel update meant that an RC2 was required Testing ======= -Testing on the 2.4.0-RC1 tag has completed; testing on RC2 is underway -RC1 was tested during dedicated time on Hyperion (17 OSS, 34 OSTs and 737 clients) -ORNL will be testing the latest tag on Titan during a maintenance window this coming week -Cray have been testing with RC1 Blockers ======== -Full list available at https://jira.hpdd.intel.com/issues/?filter=10292 -If there are any issues not presently marked as blockers that you believe should be, please let me know Other ===== -b2_4 branch now created Thanks Peter From morrone2 at llnl.gov Fri May 31 23:00:25 2013 From: morrone2 at llnl.gov (Christopher J. Morrone) Date: Fri, 31 May 2013 16:00:25 -0700 Subject: [cdwg] Lustre 2.5 Development Planning In-Reply-To: <519D0361.5010805@llnl.gov> References: <519D0361.5010805@llnl.gov> Message-ID: <51A92B89.5020204@llnl.gov> Ok, I guess I'll start things off. Here is my first pass at things I would like to see done by 2.5: LU-793 - Reconnections should not be refused when there is a request in progress from this client. LU-2498 - set block device IO scheduler to "deadline" LU-1199 - lustre build system overhaul LU-2182 - New llapi_layout API implementation LU-3338 - IOC_MDC_GETFILESTRIPE can abuse vmalloc() LU-2158 - remove lvfs and fsfilt code LU-2161 - Remove last traces of obdfilter LU-2096 - name ofd device type "ofd" instead of "obdfilter" LU-911 - OFD Supporting Proc Updates LU-2613 - opening and closing file can generate 'unrelaimable slab' space LU-2259 - e2fsprogs: rename liblustreapi.h with lustreapi.h LU-2600 - lustre metadata performance is very slow on zfs LU-2476 - poor OST file creation rate performance with zfs backend LU-2139 - track unstable pages on client LU-1784 - freeing cached clean pages is slow LU-744 - Single client's performance degradation on 2.1 LU-576 - Incorrect build dependencies listed for libcfsutil.a LU-2800 - build: clean out old autoconf options Chris On 05/22/2013 10:41 AM, Christopher J. Morrone wrote: > Hi folks, > > Please send us your list of planned development items for Lustre 2.5! In > two weeks at the next CDWG meeting we will be looking at the list of > proposed work and making our first, best attempt at pare it down to a > reasonable set of work for Lustre 2.5. > > If you have a new feature that you believe will be ready to land with in > the three month feature-landing window of Lustre 2.5, NOW is the time to > tell us! > > If you have pet bugs that are stuck on the back burner and need to be > made a development priority for Lustre 2.5, NOW is the time to tell us! > > Your input is crucial if we are to properly scope the amount of change > that we will introduce for Lustre 2.5, and properly scoping the work is > a requirement if we are to produce a solid, on-time software release. > > So don't delay! Send us your list of desired Lustre 2.5 development > items to this mailing list as soon as possible. > > Then, join us in the next CDWG call on June 5th to help us choose the > subset of work that we believe is achievable* for Lustre 2.5. > > Chris > > * Not that this list will only be our best attempt at scoping what is > reasonable, and is by no means a guarantee that any particular item will > appear in Lustre 2.5. > . >