[Twg] Lustre Requirements and Roadmap
John Carrier
carrier at cray.com
Thu Jan 27 17:24:52 UTC 2011
Below is my summary of the outcome from yesterday's meeting at the OpenSFS F2F. I would like to use this as the basis of our meeting today.
--jc
Summary of F2F Requirements Discussion (1/26/2011)
--------------------------------------------------
In the last few weeks, the TWG members gathered the following
requirements from within their organizations and the Lustre community:
- Improved error reporting
- Online filesystem integrity checks
- Integrated performance monitoring
- Completing OSD restructuring/kDMU work
- Size on MDS
- BTRFS investigation
- Imperative recovery
- Improved admin/mgmt tools
- Improved metadata performance
- LNET Channel Bonding
- end-to-end data integrity
- interface to HSM
Further discussion within the TWG organized these requirements into two
high priority categories:
* improve metadata performance
* improve scalability and reliability of backend file system
At our 1/20 meeting, TWG members discussed several options to meet these
goals:
* metadata performance
- clustered metadata servers (CMD)
- network request scheduler (NRS)
- RPC aggregation
- SMP scaling
- subtree lockings
- size on MDS
* backend storage
- OSD restructuring
- online fsck
- btrfs evaluation
Much of the design and implementation for these features was started by
Sun/Oracle. At the F2F meeting in Chicago, the OpenSFS board decided
that there is an opportunity to acknowledge any technical debt in these
initial designs and requested that, instead of an RFI for these
features, the TWG specify our long-term requirements for Lustre and then
engage the Lustre community to discuss architectures to meet these
requirements.
Therefore, the TWG will spend the next two meetings refining the
requirements for metadata and backend storage. We will then spend the
following two weeks leading discussions with the Lustre community to
define the architecture and features that will meet the requirements.
>From these discussions, the TWG will create a roadmap that OpenSFS can
use to direct feature development for Lustre 2.2 and beyond.
The following is an incomplete list of requirements to motivate further
discussion :
* metadata performance
GOAL: improve file system scalability and interactive
performance
requirements: min max
- # files in file system 100 billion 1 trillion
- # files in directory 50 million 10 billion
- file creates / sec 100 thousand 30 thousand
(aggregate) (single client)
- directory lisings / sec
- open files per process - 100 thousand
- file system capacity 30 PB 100 PB
- # clients 30 thousand ?00 thousand
* backend storage
GOAL: provide reliable, scalable backing store for
Lustre servers
requirements:
- large LUNs (min 32 TB)
- end-to-end data integrity (T10 PI or equivalent)
- no performance impact for file system repair
- framework to enable alternatives to ldiskfs
- direct I/O mode
- ??
More information about the Twg
mailing list