[Twg] Meeting notes from Feb 13 con call
David Dillow
dillowda at ornl.gov
Fri Feb 4 23:58:28 UTC 2011
So much for COB today. :/
Here's my rough notes, and Shay's much better capture of the discussion
from yesterday's meeting. I'll send my rough draft of the requirements
under separate cover.
--
Dave Dillow
National Center for Computational Science
Oak Ridge National Laboratory
(865) 241-6602 office
-------------- next part --------------
OpenSFS Technical Working Group
Meeting minutes for Feb 3, 2011
Agenda
Requirements gathering
Listing of projects in progress
Schedule for architecture meeting
Are we even ready for the meeting?
If we put it off, can prinicpals still attend?
Attending:
David Dillow ORNL dillowda at ornl.gov
Shea Seager OpenSFS shay at opensfs.org
Kit Westeneat DDN kwesteneat at ddn.com
Chris Morrone LLNL morrone2 at llnl.gov
Justin Miller IU jupmille at indiana.edu
Steve Simms IU ssimms at indiana.edu
Andreas Dilger Whamcloud adilger at whamcloud.com
Eric Barton Whamcloud eeb at whmcloud.com
Cory Spitz Cray spitzcor at cray.com
Nathan Hale IU
Requirements Gathering
Eric -- 10 billion files? what does ls do?
break posix in such a large directory?
what is driving the requirement?
open files per process ==> open files per client?
total open files per system?
* scalable fault managment
fast failover -- O(log n) on service interruption
bound pause + fault handling
replication to acheive fault tolerance
sync vs async replication
quality of service
* per job/cluster/user
* guaranteed share of IOPS/BW
allows priority levels
so simulation doesn't stutter during checkpoint
-- how to handle fairness? keep ls from going away while
jon opens checkpoint files
metadata
-- find has a similar set of requirements
full bandwidth and IOPs for single shared files
Integrity check requirements -- online vs offline
Allowing dynamic configuration of networking vs static config
-- implementation issue-- killing old RPC conneciton
Changing OST/MDT/MGT creation
-- binding targets to specific local networks explicitly
Improved LNET bandwidth and reliability
-- channel bonding
Patchless server and better support for upstream kernels
Administrative server shutdown -- clean shutdown instead of crashing
-- tell clients it is going away, no new actions until
it is back
-- ancillry "rolling upgrades"
-- ability to quiese the clients activity to a server
without unmounting filesystem
space balancing migration
-- empty an OST
-- rebalance over new OST/MDT as added
-- pool management (Use quota/ACLS for access to pool storage)
Mandatory pool authorization for use
Improved security
-- user name mapping
-- propogating authorization from MDT to OST through
trusted channels rather than untrusted clients
ADIO/MPI-IO improvements
collective open?
Improved small file IO performance
Better ability for user to specify expected access patterns
-- fadvise() to say "I will use this range, get locks
appropriately"
Improved analysis and visualization tools?
Linux AIO implementaion
readdir+/statlite/collection open from POSIX HPC ext WG
Better userspacce tools
-- redo llapi
-- lctl
Listing of projects in progress
-- kerberized connections per user to MDT, per client to OST
(Lustre 2)
-- ORNL metadata through 2.x stabilzation
imperative recovery through LNET
-- wide striping -- mostly complete
-- OSD restructuring -- well under way, no current funding
-- IU/NRL uid mapping
-- LLNL stabilaztion and Lustre on ZFS on Linux
-- DDN -- bugfixes for 1.8, possible plans for HPCFS work?
Schedule for architecture meeting
After brief discussion, we don't feel we're ready for the
architecture discussion. We need to refine our requirements
prior to the meeting on Feb 10th, and use that meeting to
generate at least a rough prioritization to guide the arch
discussion.
Currently plan to move the architecure discussion to Feb 17th.
-------------- next part --------------
A non-text attachment was scrubbed...
Name: 2-3-11 TWG concall.doc
Type: application/msword
Size: 50176 bytes
Desc: not available
URL: <http://lists.opensfs.org/pipermail/twg_lists.opensfs.org/attachments/20110204/121ce481/attachment.doc>
More information about the Twg
mailing list