{"id":178,"date":"2012-03-16T16:21:24","date_gmt":"2012-03-17T00:21:24","guid":{"rendered":"http:\/\/digitalvampire.org\/blog\/?p=178"},"modified":"2017-07-19T23:47:10","modified_gmt":"2017-07-20T07:47:10","slug":"you-can-never-be-too-rich-or-too-thin","status":"publish","type":"post","link":"https:\/\/digitalvampire.org\/blog\/index.php\/2012\/03\/16\/you-can-never-be-too-rich-or-too-thin\/","title":{"rendered":"You can never be too rich or too thin"},"content":{"rendered":"<p><a href=\"http:\/\/digitalvampire.org\/blog\/wp-content\/uploads\/2012\/03\/thin-mints1.jpg\"><img loading=\"lazy\" decoding=\"async\" title=\"Thin Mints\" src=\"http:\/\/digitalvampire.org\/blog\/wp-content\/uploads\/2012\/03\/thin-mints1.jpg\" alt=\"Thin Mints by by Jesse Michael Nix\" width=\"614\" height=\"439\" \/><\/a><\/p>\n<p>One of the cool things about having a <a title=\"Pure Storage products\" href=\"http:\/\/www.purestorage.com\/products\/\">storage box that virtualizes everything at sector granularity<\/a> is that there&#8217;s pretty much zero overhead to creating as big a volume as you want.\u00a0 So I can do<\/p>\n<pre>    pureuser@pure-virt&gt; purevol create --size 100p hugevol\r\n    pureuser@pure-virt&gt; purevol connect --host myinit hugevol<\/pre>\n<p>and immediately see<\/p>\n<pre>    scsi 0:0:0:2: Direct-Access\u00a0\u00a0\u00a0\u00a0 PURE\u00a0\u00a0\u00a0\u00a0 FlashArray\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 100\u00a0 PQ: 0 ANSI: 6\r\n    sd 0:0:0:2: [sdb] 219902325555200 512-byte logical blocks: (112 PB\/100 PiB)\r\n    sd 0:0:0:2: [sdb] Write Protect is off\r\n    sd 0:0:0:2: [sdb] Mode Sense: 2f 00 00 00\r\n    sd 0:0:0:2: [sdb] Write cache: disabled, read cache: enabled, doesn't support DPO or FUA\r\n    sdb: unknown partition table\r\n    sd 0:0:0:2: [sdb] Attached SCSI disk<\/pre>\n<p>on the initiator side.\u00a0 Being able to create gigantic LUNs makes using and managing storage a lot simpler &#8212; I don&#8217;t have to plan ahead for how much space I&#8217;m going to need or anything like that.\u00a0 But going up to the utterly ridiculous size of 100 petabytes is fun on the initiator side&#8230;<\/p>\n<p>First, I tried<\/p>\n<pre>    # mkfs.ext4 -V\r\n    mke2fs 1.42.1 (17-Feb-2012)\r\n            Using EXT2FS Library version 1.42.1\r\n    # mkfs.ext4 \/dev\/sdb<\/pre>\n<p>but that seems to get stuck in an infinite loop in <tt>ext2fs_initialize()<\/tt> trying to figure out how many inodes it should have per block group. Since block groups are 32768 blocks (128 MB), there are a <em>lot<\/em> (something like 800 million) of block groups on a 100 PB block device, but ext4 is (I believe) limited to 32-bit inode numbers, so the number of inodes per block group calculated ends up being about 6, which the code then rounds it up to a multiple of 8 &#8212; that is, up to 8. It double checks that 8 * number of block groups doesn&#8217;t overflow 32 bits, but unfortunately it does, so it reduces the inodes\/group count it tries, and goes around the loop again, which doesn&#8217;t work out any better.\u00a0 (Yes, I&#8217;ll report this upstream in a better forum too..)<\/p>\n<p>Then I tried<\/p>\n<pre>    # mkfs.btrfs -V\r\n    mkfs.btrfs, part of Btrfs Btrfs v0.19\r\n    # mkfs.btrfs \/dev\/sdb<\/pre>\n<p>but that gets stuck doing a <tt>BLKDISCARD<\/tt> ioctl to clear out the whole device. It turns out my array reports that it can do SCSI UNMAP operations 2048 sectors (1 MB) at a time, so we need to do 100 billion UNMAPs to discard the 100 PB volume. My poor kernel is sitting in the unkillable loop in <tt>blkdev_issue_discard()<\/tt> issuing 1 MB UNMAPs as fast as it can, but since the array does about 75,000 UNMAPs per second, it&#8217;s going to be a few weeks until that ioctl returns.\u00a0 (Yes, I&#8217;ll send a patch to btrfs-progs to optionally disable the discard)<\/p>\n<p style=\"padding-left: 30px;\">[<em>Aside: I&#8217;m actually running the storage inside a VM (with the FC target adapter PCI device passed in directly) that&#8217;s quite a bit wimpier than real Pure hardware, so that 75K IOPS doing UNMAPs shouldn&#8217;t be taken as a benchmark of what the real box would do.<\/em>]<\/p>\n<p>Finally I tried<\/p>\n<pre>    # mkfs.xfs -V\r\n    mkfs.xfs version 3.1.7\r\n    # mkfs.xfs -K \/dev\/sdb<\/pre>\n<p>(where the &#8220;-K&#8221; is stops it from issuing the fatal discard) and that actually finished in less than 10 minutes. So I&#8217;m able to see<\/p>\n<pre>    # mkfs.xfs -K \/dev\/sdb\r\n    meta-data=\/dev\/sda               isize=256    agcount=102401, agsize=268435455 blks\r\n             =                       sectsz=512   attr=2, projid32bit=0\r\n    data     =                       bsize=4096   blocks=27487790694400, imaxpct=1\r\n             =                       sunit=0      swidth=0 blks\r\n    naming   =version 2              bsize=4096   ascii-ci=0\r\n    log      =internal log           bsize=4096   blocks=521728, version=2\r\n             =                       sectsz=512   sunit=0 blks, lazy-count=1\r\n    realtime =none                   extsz=4096   blocks=0, rtextents=0\r\n    # mount \/dev\/sdb \/mnt\r\n    # df -h \/mnt\/\r\n    Filesystem      Size  Used Avail Use% Mounted on\r\n    \/dev\/sdb        100P  3.2G  100P   1% \/mnt<\/pre>\n","protected":false},"excerpt":{"rendered":"<p>One of the cool things about having a storage box that virtualizes everything at sector granularity is that there&#8217;s pretty much zero overhead to creating as big a volume as you want.\u00a0 So I can do pureuser@pure-virt&gt; purevol create &#8211;size 100p hugevol pureuser@pure-virt&gt; purevol connect &#8211;host myinit hugevol and immediately see scsi 0:0:0:2: Direct-Access\u00a0\u00a0\u00a0\u00a0 PURE\u00a0\u00a0\u00a0\u00a0 [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":182,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[4,17],"tags":[],"class_list":["post-178","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-hacking","category-linux"],"_links":{"self":[{"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/posts\/178","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/comments?post=178"}],"version-history":[{"count":7,"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/posts\/178\/revisions"}],"predecessor-version":[{"id":618,"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/posts\/178\/revisions\/618"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/media\/182"}],"wp:attachment":[{"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/media?parent=178"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/categories?post=178"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/digitalvampire.org\/blog\/index.php\/wp-json\/wp\/v2\/tags?post=178"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}