# Rui's Blog/Paper Reading Notes - Introduction

## Introduction

This GitBook mainly hosts [Rui Pan](https://ruipan.xyz)'s paper reading notes. I also keep a non-technical blog [here](/blog/blog-index).

Most of the papers here are about [Machine Learning Systems](/machine-learning-systems/machine-learning-systems-index).

If you spot a mistake or have any suggestions, you are more than welcome to drop me an email at <ruipan@princeton.edu>!

## Topics

{% content-ref url="/pages/-MNWlxwYEWxQm6nw06yq" %}
[Personal Blog](/blog/blog-index)
{% endcontent-ref %}

{% content-ref url="/pages/-MMTv1vvdZRzlssxlvi5" %}
[Machine Learning Systems](/machine-learning-systems/machine-learning-systems-index)
{% endcontent-ref %}

{% content-ref url="/pages/-MMTv1vxMlnF5sjkBZXW" %}
[Big Data Systems - Index](/machine-learning-systems/index)
{% endcontent-ref %}

{% content-ref url="/pages/-MMTv3UnIlaWe0BzD9Ye" %}
[Operating Systems Papers - Index](/earlier-readings-and-notes/index)
{% endcontent-ref %}


# Personal Blog - Index


# How to Create Picture-in-Picture Effect / Video Overlay for a Presentation Video

The picture-in-picture effect (video overlay) allows you to place a video clip in a small frame on top of another so that they play at the same time like this:

![Source: https://www.youtube.com/watch?v=YNJYrM4ecww](/files/-MNWmupIzZff1R4q_IAK)

There are multiple ways to make such videos. Here, I will be introducing my workflow on macOS to create these videos using an easy-to-use, PowerPoint-like application, [ActivePresenter](https://atomisystems.com/activepresenter/) by Atomi Systems.

## Requirements

* [Download and install ActivePresenter](https://atomisystems.com/download/)
* A computer that runs Windows /  macOS
* A webcam
* A microphone (recommended for audio quality)

## Creating Presentation Slides

ActivePresenter is like Microsoft PowerPoint but with additional support for embedding video clips. ActivePresenter supports creating slides and video recordings from within the application. It also supports importing slides from PowerPoint.

## Create a Screen Recording / Import a Pre-Recorded Video

In ActivePresenter, videos can be inserted either by recording a webcam footage in ActivePresenter or doing a drag-and-drop to import pre-recorded videos. Personally, I prefer to do many takes using the macOS-built-in QuickTime Player app, and have one separate clip for each slide in the presentation. The videos can be resized to the same size and placed in a specific corner of the output video.

When recording a video clip, considering using a microphone to record the audio. All built-in microphones in laptops and earphones create static noises that impact the audio quality by a lot. Also, consider adding closed captions for deaf and hard-of-hearing viewers.

## Managing Animations, Closed Captions, etc.

![](/files/-MNWpoC7C7TVIJUeJptZ)

ActivePresenter has a neat functionality that allows users to control the timeline of events in a slide. By dragging the horizontal bars, I can control when elements appear/disappear, adjust the properties of the embedded video clip, and even add closed captions. Closed captions can either be exported as soft subtitles (.SRT) that can be uploaded to platforms like YouTube or as hard subtitles embedded into videos.

## Exporting the video

As easy as a click of a button, final videos can be exported as mp4 files.


# How to Do Your Part to Protect the Environment in Wisconsin

The following content is copied from the Canvas page of the course ENVIR ST 101: Forum on the Environment (sp21). This is a list of personal challenges everyone can do to help with environmental protection.

There are so many things we can do as individuals to help protect our environment.  Here is a compilation of a range of different challenges you might choose from for your weekly personal challenge in the second part of the semester. We've grouped a few ideas into categories (some actions are listed under multiple categories). We don't want the challenges to be a financial burden on anyone, so we have tried to include a diversity of options, many of which should not cost money. You can also consider coming up with your own challenges! It's worth considering in your weekly reflection the situations in which these actions might not result in a more sustainable lifestyle - e.g., if you need to drive for 30 minutes to buy your organic produce, that might reduce the positive impact.

**Food**

• Transporting food long distances can increase its climate footprint. This week, try to reduce the distance your food has been transported by increasing the proportion of locally-grown or raised foods.

• Getting to know the farmers who grow your food is a great way to begin to understand the pathway food takes from the farm to your plate. Visit a farmer's market, such as the [Dane County Farmer's Market (Links to an external site.)](https://dcfm.org/), and say hello to a farmer!

• Agrochemicals used during food production may harm native wildlife, including important pollinators. Organic agriculture reduces the use of these products. This week, try increasing the proportion of organically-produced foods you eat.

• Animal production is one of the most resource- and energy-intensive aspects of our agricultural system. This week, try to reduce the amount of meat or other animal products that you eat.

• A vegan diet can dramatically reduce your impact on the plant. This week, try going vegan.

• Food packaging can add up in our household waste stream. This week, try to reduce the amount of packaging associated with the foods you eat.

• Seafood can have a big environmental impact, but [Seafood Watch (Links to an external site.)](https://www.seafoodwatch.org/) and other organization help us make more sustainable decisions about the seafood we eat. This week, try to choose more sustainable seafood choices.

• [Reducing food waste (Links to an external site.)](https://www.worldwildlife.org/stories/fight-climate-change-by-preventing-food-waste) is a [huge way we can help (Links to an external site.)](https://www.unenvironment.org/regions/north-america/regional-initiatives/minimizing-food-waste) reduce the impact of our diet. This week, try to reduce the amount of food that you end up throwing away or composting.

• When food waste goes to a landfill, those nutrients are lost! This week, look into [composting at home (Links to an external site.)](https://www.cityofmadison.com/streets/compost/) or [on campus.Links to an external site.](https://sustainability.wisc.edu/composting/)

• Buying in bulk is a great way to reduce the impact of food packaging. This week, try reducing your food waste by shopping in the bulk section (at [some stores (Links to an external site.)](https://www.willystreet.coop/about-us/departments), you can bring your own containers, or you can re-use bags you bring yourself).

• Another way to cut down on food packaging is to avoid bagging items unnecessarily. This week, while grocery shopping, try either bringing your own plastic bags to re-use for produce, or don't put the produce in a bag at all. Then, at the cash, bring your own bag instead of using a bag from the store. Many stores even offer a discount for this!

**Waste**

• We often end up with a lot of materials we don't need while shopping. This week, bring a reusable bag with you when you go shopping, or consider refusing a bag at the cash if it is easy to carry your purchase(s).

• We can reduce our waste footprint at the coffee shop by reusing or refusing items. This week, try to reduce your waste impact by bringing your own reusable mug for drinks or choosing a "for here" mug instead of a to-go cup, and declining extra items you might not need (stir stick, heat protector, lid, double-cupping).

• Takeout or food delivery packaging waste can be surprisingly high. This week, try to (1) reduce the amount of unnecessary materials you receive, (2) find ways to re-use what you do receive, and (3) increase the proportion of materials you recycle.

• [Reducing food waste (Links to an external site.)](https://www.worldwildlife.org/stories/fight-climate-change-by-preventing-food-waste) is a [huge way we can help (Links to an external site.)](https://www.unenvironment.org/regions/north-america/regional-initiatives/minimizing-food-waste) reduce the impact of our diet. This week, try to reduce the amount of food that you end up throwing away or composting.

• When food waste goes to a landfill, those nutrients are lost! This week, look into [composting at home (Links to an external site.)](https://www.cityofmadison.com/streets/compost/) or [on campus.Links to an external site.](https://sustainability.wisc.edu/composting/)

• Food packaging can add up in our household waste stream. This week, try to reduce the amount of packaging associated with the foods you eat by choosing food items that have less packaging or packaging that is recyclable.

• Single-use plastics are a major source of plastic pollution to our oceans and lands. This week, try to specifically cut out single-use plastics - that means stir sticks, drinking straws, plastic bags, etc.

• Buying in bulk is a great way to reduce the impact of food packaging. This week, try reducing your food waste by shopping in the bulk section (at [some stores (Links to an external site.)](https://www.willystreet.coop/about-us/departments), you can bring your own containers, or you can re-use bags you bring yourself).

• Another way to cut down on food packaging is to avoid bagging items unnecessarily. This week, while grocery shopping, try either bringing your own plastic bags to re-use for produce, or don't put the produce in a bag at all. Then, at the cash, bring your own bag instead of using a bag from the store. Many stores even offer a discount for this!

• One of the best ways to reduce your impact is to reduce your consumption overall. This week, look for ways you can reduce the amount you buy.

• One of the next ways we can reduce our impact is to re-use the things we do buy. This week, look for ways to re-use things you might have otherwise discarded.

• One way of increasing our re-use of things at a community level is to buy things second-hand. This week, try to look for used versions of items you might have otherwise bought. This could mean shopping for second-hand or vintage clothing, looking for a used version of a tool you need on an online marketplace, or finding something you need through [freecycle (Links to an external site.)](https://groups.freecycle.org/group/MadisonWI/posts/all).

• The flip side of re-using things yourself is looking for ways to make sure the things you're finished with have a second life. This week, try to sell or give away items you might have otherwise thrown out.

• If you can't reduce your consumption, or re-use an item, the next best option is to recycle it. This week, educate yourself on recycling practices at your workplace, on campus, at your house, or in your dorm, and take some kind of step to improve your recycling practices.

**Community Engagement**

• One of the best way to make changes while building connections with others is to be active in your community. This week, volunteer at an environmentally-related event.

• Getting involved on campus can be rewarding! This week, attend a meeting or an event of an [environmentally active campus groupLinks to an external site.](https://sustainability.wisc.edu/student-organizations/).

• One of the most important ways we can make change is through making our voices heard. This week, [reach out to a politician (Links to an external site.)](https://www.nytimes.com/guides/year-of-living-better/how-to-participate-in-government) at any level of government - e.g., municipal, state, federal - to let them know about an environmental issue that is important to you. Here is a guide from the APA on [writing to a member of congress (Links to an external site.)](https://www.apa.org/advocacy/guide/letter-email). If you're feeling really ambitious, consider organizing a letter-writing party with your friends. Or, instead of writing, you can call your representative.

• Another way to make change is by getting involved directly with politics. This week, look into ways to get involved with a political party. This could mean volunteering, attending an event, or even joining the party.

• Knowledge is power. This week, educate yourself about an environmental issue by attending a public event.

• One of the most important things you can do to fight climate change, or other environmental problems, is to [talk about it (Links to an external site.)](https://www.ted.com/talks/katharine_hayhoe_the_most_important_thing_you_can_do_to_fight_climate_change_talk_about_it?language=en). Katharine Hayhoe has a Ted Talk about this idea. This week, make it a priority to talk about an environmental issue with people in your life - this might be coworkers, classmates, friends, or family.

**Energy**

• Transportation can be a big contributor to our energy, climate change, and air pollution impact. This week, choose walking or biking over fossil-fuel powered modes of transport (bus or car) when possible.

• Biking is a great way to decrease our environmental impact while also getting exercise. This week, enhance your bike power by using [Madison BCycle  (Links to an external site.)](https://madison.bcycle.com/nav/rates)or the [RedBike bikeshare program (Links to an external site.)](http://redbikes.org/), or checking out the [UW Bicycle Resource CentreLinks to an external site.](https://transportation.wisc.edu/bicycling/university-bicycle-resource-center/) and maybe tuning up your bike.

• When you can't bike or walk, using public transit can help decrease your impact. This week, take the bus instead of driving when possible.

• After driving less, there are many ways you can decrease the impact from driving a car. This week, if you drive a car, look for ways to [save fuel while driving (Links to an external site.)](https://www.cityofmadison.com/sustainability/Transportation/drivingTips.cfm).

• Another way to reduce the impact from your car is to carpool. This week, look for opportunities to share your ride.

• A direct way of saving energy at home is to [adjust your thermostat (Links to an external site.)](https://www.mge.com/saving-energy/for-homes/heating-and-cooling/thermostat-settings). This week, see if you can keep it a bit cooler (if heating) or warmer (if using air conditioning).

• Laundry can be a big energy and water sink, but there are [lots of things we can do to reduce this impact (Links to an external site.)](https://www.mge.com/saving-energy/for-homes/appliances/laundry). This week, see how you can improve the sustainability of your laundry routine.

• Wisconsin has a Focus on Energy program, where you can get energy-saving tools for free, such as advanced power strips or LED bulbs! This week, [order a pack for your home (Links to an external site.)](https://focusonenergy.com/residential#program-energy-saving-packs).

• Reducing standby power or "phantom power" is a quick way to decrease our energy impact. This week, take [steps to reduce standby power (Links to an external site.)](https://www.energy.gov/energysaver/articles/3-easy-tips-reduce-your-standby-power-loads) in your life.

• Seeing is believing. This week, if you're feeling ambitious, get together a group of at least 6 people and arrange for a free tour of Madison's [Blount Generating Station (Links to an external site.)](https://www.mge.com/about-mge/power-plants/blount-generating-station).

**Biodiversity**

• Spring brings beautiful wildflowers! This week, find and identify [five wildflowersLinks to an external site.](http://wisflora.herbarium.wisc.edu/projects/index.php?pid=8) that grow in this area.

• Lichens are [wacky and wonderful organisms (Links to an external site.)](https://www.theatlantic.com/science/archive/2016/07/how-a-guy-from-a-montana-trailer-park-upturned-150-years-of-biology/491702/). This week, go for a walk in the woods (or somewhere around trees), and use the [Lichens of WI guideLinks to an external site.](https://herbarium.wisc.edu/wp-content/uploads/sites/205/2017/10/lichens-of-wi-web-20170515.pdf) to try to identify lichens.

• Trees are a precious natural resource. This week, go for a walk in your neighborhood or in a natural area, and use one of the [WI DNR-recommended (Links to an external site.)](https://dnr.wi.gov/education/educatorresources/treeid.html) guides to try to identify trees - the [WI Urban Tree key (Links to an external site.)](https://cf-store.widencdn.net/widnr/e/a/9/ea9aff02-7c37-47a6-8edc-25278f0c8028.pdf?response-content-disposition=inline%3B%20filename%3D%22Urban-Tree-Key.pdf%22\&response-content-type=application%2Fpdf\&Expires=1579567110\&Signature=hf9cfmw9siH90Y9UB3ZSuBkKNqDpM4S41oZMIefQlcnKfUP5OHiICS~KifIK5Iyr6B9bKmz8Zxky2brdwWAwsdA9MACizNyFsX7xy6~lnrDZinBsyXBwVOAsflvhFqu3-g10Hmxo6eIgNGNSJrUzZcO2oWsdYTfPRVnW0gUInJptkuED5ZEQPXErEznsQD7-EFyWHB0bHdEXg7a4O7UK9e1j4WEwzx1Q2SbecjcKo~XHYwLOkAq160r6QHyJpCxhAp4pA4xJuU62iIve7ubfgb3Q6HMqKbV0eb4QZ4lpRDd~Fl5Zp2o9QAhYNGYNj3dJwIldk~2waE1n9zVgJGGsfw__\&Key-Pair-Id=APKAJD5XONOBVWWOA65A) is a simple one to start with.

• Get outside! This week, visit a [state park, forest, or recreation (Links to an external site.)](https://dnr.wi.gov/topic/parks/?utm_source=FeatureImage\&utm_medium=Homepage\&utm_campaign=20131212_StateParks) area.

• Learning about invasive species can help us preserve local biodiversity. This week, learn to identify some of the local invasive species. [This guide (Links to an external site.)](https://dnr.wi.gov/topic/Invasives/documents/WI%20inv%20plant%20field%20guide%20web%20version.pdf) is more in depth, or [this one (Links to an external site.)](https://dnr.wi.gov/topic/Invasives/documents/WI_common_inv_Montage\(3-25\).pdf) is a simple starting point.

• We're lucky to have beautiful natural areas right on campus! This week, visit the [Lakeshore Nature PreserveLinks to an external site.](https://lakeshorepreserve.wisc.edu/what-is-the-lakeshore-nature-preserve/). You might [attend one of their eventsLinks to an external site.](https://lakeshorepreserve.wisc.edu/events-calendar/) or there are all sorts of [volunteerLinks to an external site.](https://lakeshorepreserve.wisc.edu/volunteer/) opportunities there.

• The UW-Madison Arboretum is another beautiful natural area. This week, [visit the ArboretumLinks to an external site.](https://arboretum.wisc.edu/visit/), whether it is for [an eventLinks to an external site.](https://arboretum.wisc.edu/visit/events/), to [volunteerLinks to an external site.](https://arboretum.wisc.edu/get-involved/), or just to explore on your own.

• Learning about the natural world around us is a great way to connect with nature. This week, explore [sightings in or around Madison on iNaturalist (Links to an external site.)](https://www.inaturalist.org/observations?nelat=43.171916\&nelng=-89.24645199999998\&place_id=any\&swlat=42.998071\&swlng=-89.56638889999999), and log at least one sighting of an organism on the site.

**Water**

• Water bottle filling stations and fountains make it easy to reduce plastic water bottle waste around campus. This week, only use reusable water vessels (bottles, glasses, mugs).

• Laundry can be a big energy and water sink, but there are [lots of things we can do to reduce this impact (Links to an external site.)](https://www.mge.com/saving-energy/for-homes/appliances/laundry). This week, see how you can improve the sustainability of your laundry routine.

• Read up! This week, read (or start reading) a book about water, as [recommended by the Water@UWLinks to an external site.](https://water.wisc.edu/water-reading-list/) group.

• Turn off the tap! This week, if you're someone who typically lets the sink run continuously while doing dishes, brushing your teeth, washing your hands, etc., [turn off the tap to save water (Links to an external site.)](https://sustainability.ncsu.edu/blog/changeyourstate/6-times-you-should-turn-off-the-tap-to-save-water/).

• Every little bit counts. This week, see if you can reduce water usage while showering or bathing.

• You have a right to clean and safe water. This week, [find and read (Links to an external site.)](https://ofmpub.epa.gov/apex/safewater/f?p=ccr_wyl:102) a recent water quality report for Madison (look for the Madison water utility).

• We are lucky to be living right next to some beautiful lakes and waterways. This week, spend some time enjoying Lake Monona, Lake Mendota, Lake Wingra, Lake Waubesa, Cherokee Marsh, the Yahara River, Wingra Creek, or another nearby waterway.

• Straight to the source! This week, [take a free tour of the Madison Metropolitan Sewerage District’s wastewater treatment plant (Links to an external site.)](https://www.madsewer.org/Education/Take-a-Tour) (they run the first Friday of each month at noon, or by appointment).


# How to Get a Driver's License in Wisconsin

## Step 1: Getting a learner's permit

To get an instruction permit in Wisconsin, you need to provide some documentation and pass the knowledge test. A full list is available [here](https://wisconsindot.gov/Pages/dmv/teen-driver/yr-frst-lcns/permit.aspx). The knowledge test is walk-in for both (east & west) DMVs in Madison.

3 hours of doing practice exams is enough as long as you know the basics of driving. I prepared for the exam by grinding through practice exams on [DMV Genie](https://driving-tests.org/dmv-genie/), and there are a few numbers that need to be memorized in the official [Motorists Handbook](https://wisconsindot.gov/Documents/dmv/shared/bds126-motorists-handbook.pdf).

Once you fill out the application, finish the knowledge test, pass the vision test, and pay the fees, just wait for the permit to be delivered in around two weeks.

## Step 2: Learning to drive

If you plan to take the road test using a Zipcar, it's important to get familiarized with the performance of the vehicle before the exam. Try doing a few warm-up walkthroughs of the testing route a few hours before the road test. If you are not familiar with the vehicle, you might get an auto-fail in the parking lot even before you start the test!

When I learned to drive, I got a ton of help from Weijie Wang (a great friend). If you want to learn to drive, I will be happy to help if you are also willing to help other people after you get your license :)

## Step 3: Passing the road test

There are two DMVs in Madison. Most people I know take the test at the West DMV. To pass a test, you need to:

* Get deducted 25 points or less
* Drive without any disqualifications (auto-fails)

The following is an evaluation sheet used by Wisconsin road test proctors.

![](https://github.com/ruipeterpan/blog/blob/master/content/posts/images/20200206-1.jpg?raw=true)

And here is the usual test route for line 1. Line 2 has a few variations in the residential area, but it should not be too big of an issue. Try to memorize the speed limits of certain road sections before the exam: 35 miles/hr for "big roads", 25 miles/hr for residential areas, and a 30 miles/hr section in the middle. Use Google Map street view for reference.

![](https://s2.ax1x.com/2020/02/20/3ZuA9s.jpg)

## Step 4: Bonus! Zipcar-ing

From my experience, I believe that a Zipcard cannot be delivered to some apartments like Lucky and X01 as they have mailboxes inside the building (???). After three failed attempts, I got my Zipcard delivered to a friend who lives in University Housing.

The rate for on-campus Zipcars is $5.5/hr. It's really awesome for weekend road trips or grocery runs, especially if multiple people are on board!

If you feel this tutorial helped you, please consider dropping me an email (or just try this referral [link](https://refer.zipcar.com/s/RuiPan)) so that I can send you a referral link and we both get $25 of driving credits :) When searching for an organization, choose the option`University of Wisconsin, Madison (UW) - Student Leaders` so you can drive without having to pay an annual fee. I think everyone gets approved, so kudos to whoever reviews these applications.

Enjoy driving!


# How to Travel from the U.S. to China onboard AA127 in June 2021

Shoot me an email at rpan33\@wisc.edu if you have any further questions/need a guide in English!

## 前言

* 路上拍了个[vlog](https://youtu.be/TfZENTQxdkA) :)
* 本文写于2021年6月28日。请参考北美票帝的微信公众号/微博和大使馆官网来获取最新的政策/要求。
* 建议在出行前一个月找到飞友微信群互帮互助，我是在北美票帝的微博评论区里找到了拉我进群的好心陌生人。进群了之后参考了很多飞友总结的资料（最全面的是这篇[AA127情况汇总](https://docs.google.com/document/d/1-m6GvE3ZDos4Mtm27KZhwPAYH0CTZme-Jh3zi_Cygwk/edit)和[从零开始回国指南](https://mp.weixin.qq.com/s/8Z4nrqtVh0IdvMMaaokYwA)，其他的文章我也在文中附了链接），身边也有数不胜数的朋友（Haochen, Yuhan, Yushun等等）提供了帮助，在此道一声感谢！
* AA127这班航班比较特殊，目前仅建议持F/J/M签证或有驻美大使馆认可的“紧急必要”原因的乘客乘坐，非学生签证的可以参考北美票帝的航班表选择其他路线；所有非中国公民，请事先与各个使领馆确认中国签证/居留许可的有效性和回国的“紧急必要”性。
* 购买国内航司的航班之前需要注意，国内为了“公共安全“会每周取消一个国内航司的航班，具体信息可以参考[这篇文章](https://mp.weixin.qq.com/s/HEIGvLELF5OHEIw8kCEX2g)。

![](/files/-MdDQXHnB5uOpVmmucfq)

## **时间线**

* 4/5: 在UHS接种第一针Moderna
* 5/7: 在UHS接种第二针Moderna
* 5/14: 购买6/28 AA127 DFW-PVG
* 5/15: 购买6/25 AA4231 MSN-DFW
* 5/20: 预约RealTime在6/25的PCR\&IgM检测
* 6/8: 经群友提醒Realtime存在IgM-N蛋白假阳的可能性，reschedule了DFW附近仅有的，经大使馆批准的另一家检测机构 (Ayass)
* 6/26 5:50 AM: 从Hyatt Regency DFW出发，乘坐机场shuttle去car rental center
* 6/26 6:00 AM: 到达Car Rental Center, 在Dollar的柜台排队
* 6/26, 7:00 AM: 在Dollar Car Rental checkout车辆
* 6/26, 7:35 AM: 到达Ayass (顺位是第四辆车), 开始排队
* 6/26, 8:34 AM: Drive-through测PCR, 下车排队测IgM
* 6/26, 9:10 AM: 测完IgM
* 6/26, 4:16 PM: 拿到PCR的阴性报告
* 6/26, 11:43 PM: 拿到IgM的报告 (IgM-S Reactive/Positive, IgM-N Non-Reactive/Negative)
* 6/27, 12:30 AM: 提交健康码审核材料
* 6/27, 2:24 AM: 收到绿码
* ???: 申请指尖码
* ???: 顺利登机
* ???: 到达PVG
* ???: 提行李
* ???: 等大巴
* ???: 入住酒店，开始隔离

## 交通和住宿

* MSN -> [DFW](https://goo.gl/maps/Pjqdba8vd3yMQRmVA)
  * 由于现在中国大使馆要求旅客在飞往中国的航班始发地进行检测，乘客需要提前两天到达出发地。我选择了分开购买MSN-DFW和DFW-PVG的机票（非联程，中间相隔两天）。如果检测点不在常住地，请谨慎购买联程机票，否则有可能在前往检测点的这一班航班check in时就会被不熟练业务的地勤要求出示健康码...
* DFW -> [Hyatt Regency DFW](https://goo.gl/maps/ZL7V2AzvH8nTQ8Fp9)
  * 无托运行李：下飞机后可以坐DFW的Skylink小火车到Terminal C。
  * 有托运行李：需要先出terminal提行李，之后可以坐十分钟一班的Terminal Link Shuttle Bus到Terminal C。
  * Hyatt这家酒店虽然地理位置紧靠Terminal C，但它和terminal不是连着的。
  * 不怕热，行李不重的同学可以直接跟着指示牌走这条路线：Gate C19 -> DFW Parking Lot -> Hyatt Regency DFW Parking Lot。这段路程大概只要走五分钟，但是台阶大概要上上下下三四层，而且德州夏天会非常热。另一条路线是给酒店打电话然后直接乘Hyatt Regency到Terminal C的shuttle（20分钟一班）坐到酒店门口，不过这条路线我没有走过不太熟悉...
* Hyatt Regency DFW
  * 用GrubHub可以点外卖送到前台，周围可以送GrubHub的店有Shake Shack, Chipotle, Panda Express, Hooters, Potbelly, Jersey Mike's等等
* Hyatt Regency DFW -> [DFW Car Rental Center](https://goo.gl/maps/r1ZzqfeAJTtpVeP29)
  * 机场有到Car Rental Center的shuttle，频次大约是15分钟一班，车程十分钟左右。注意DFW Terminal出来后有两层，行李转盘的那一层是露天的，shuttle的站是在下面一层，可以通过 {在terminal里坐自动扶梯, 走到terminal parking lot后走楼梯下楼}到达。
* DFW Car Rental Center -> [Ayass Lab](https://goo.gl/maps/TXeuN6Fg7MBBTwPG8):
  * 自己租车的话走toll road全程高速，开过去只要半小时。德州交过路费比较麻烦（要么买pass，要么pay by mail）。为了方便，我在Dollar Car Rental买了一天12刀的toll pass。
  * 如果不会开车/嫌麻烦，可以在飞友群里联系华人司机/和他人拼车！
* Checking in at DFW
  * 我这班AA127是在Terminal D飞。因为AA[独占了大半个DFW](https://www.airport-dallas.com/terminals.php)，理论上在任意一个terminal都可以check in。

## Ayass检测流程

* PCR
  * Drive through: 在指示牌后排队
  * Walk in: 直接走到蓝色棚子里和工作人员说，不过不清楚walk in的需要排多久队
  * 工作人员会给每个人发放两张表格，分别是PCR和IgM的检测表格
  * 填表格：填写个人信息，billing info（虽然PCR免费，但如果有学生保险，还是可以在这里填写保险信息，Ayass大概会bill保险公司？），questionnaire，然后print name + signature。具体怎么填写可以看[这篇文章](https://docs.qq.com/pdf/DTWtFWHZLbWp3Umpn)
  * 工作人员会收走护照和PCR的测试表格
  * 在蓝色棚子里进行鼻咽拭子的样本采集。采集前，工作人员会交还护照和护照首页的复印件
* IgM
  * 测完IgM，停好车，在Ayass八号楼排队测试，准备好护照，护照首页复印件和IgM的测试表格
  * 如果要加测N蛋白，进入测试点后跟医生说，医生会让你填两张表，第一张在划黄色荧光笔的地方填写个人信息，第二张需要签字
  * 准备好450刀现金（如果不测N蛋白，只需要350），工作人员可以提供找零
  * 收钱之后，工作人员会提供receipt，这张收据一定要收好，之后拍照要用
  * 抽血前请和工作人员仔细double check个人信息
* 测完后
  * 拍三张照片，详见“申请健康码“section
* 上厕所：Ayass没有public bathroom，憋得慌可以去一迈之外的[Walgreens](https://goo.gl/maps/pchSCYdAs6PAeUj89)解决（开车五分钟，走路20分钟）
* 周边吃饭：Walgreens那一片有很多饭店，可以点个pick up或者drive through在车里吃
  * 这次来德州最大的收获是发现Walgreens有卖[Nice!](https://www.walgreens.com/store/c/productlist/nice!-snacks/N=360665-362398)这个牌子的零食，好吃又便宜我真的爱了

## 申请健康码

* PCR, IgM-S, IgM-N, IgG解释（参考了[这篇文章](https://docs.qq.com/doc/DSHpwV0NDYkdZSVFT)）：
  * PCR（SARS-CoV-2 by RT-PCR 阴性报告）：鼻咽拭子测试，要拿到绿码这个必须是阴/Not Detected。
  * IgM-S（IgM for “Spike Protein (S-Protein)，S蛋白）：要拿到绿码，没打过疫苗的话S蛋白必须是阴/Non-Reactive，打过疫苗的话S蛋白可以阳，但是N蛋白必须阴。
  * IgM-N（IgM for “Nucleocapsid Protein (N-Protein)，N蛋白）：阴性表明近期没有感染新冠病毒，或者近期没有暴露在新冠病毒的环境下；如果是阳性，表明有近期新冠病毒感染 。这是大使馆最重视的一项结果。
  * IgG：IgG是长期抗体，Ayass没有测，我也不太了解，可以看看上面链接的那篇文章
* 申请健康码的portal是一个微信小程序，全名叫“防疫健康码国际版“。
* 对于打了疫苗的同学，申请健康码所需的图片可能会超过健康码小程序的上传照片数量限制（十张）。我用[Scanner Pro](https://apps.apple.com/us/app/scanner-pro-pdf-scanner-app/id333710667)扫描了文档，然后用[Picsew](https://apps.apple.com/us/app/picsew-screenshot-stitching/id1208145167)对图片进行了拼接。
* 健康码需要的十张图片（**附录里会附上打了码的照片供参考**）：
  1. 护照首页 + 签证页&#x20;
  2. 疫苗接种声明书 + CDC小白卡 + 疫苗接种机构证明
     1. 疫苗接种声明书可以从大使馆官网下载，链接[在这](http://www.china-embassy.org/eng/notices/P020210421787870030822.pdf)。注意网上有好几个版本的接种声明表，现在应该是只认这个有英文的表格
     2. 疫苗接种机构证明：Walgreens应该有自己的documentation。我是在学校的UHS进行的接种，所以在Wisconsin Immunization Registry request了record。全美50州获取疫苗接种记录的方法可以看[这里](<https://www.cdc.gov/vaccines/programs/iis/contacts-locate-records.html&#xA;>)
  3. I20
  4. 回国航班的Itinerary&#x20;
     1. AA官网，confirmation email里都可以找到
  5. Ayass预约邮件
     1. 在[Ayass官网预约](https://ayassbioscience.com/covid-19-testing-for-travelers-to-china/)之后，confirmation email里可以找到，保存成PDF
  6. PCR报告
  7. IgM报告
     1. 如果同时测了IgM-S和IgM-N，可以把Ayass发来的两页PDF转成png再拼图
  8. 在Ayass拍的另外三张照片
     1. 要求：手持护照首页，Ayass的收款receipt，露出静脉抽血伤口，露脸。戴眼镜但是护照照片上没戴的同学们在拍照时建议把眼镜摘了
     2. 照片一：Ayass Bioscience Inc 八号楼楼下，COVID-19 IgM Testing牌子前
     3. 照片二：IgM Sample Collection门前
     4. 照片三：Ayass Bioscience门前
* 可能需要提交的其他材料
  * 做PCR时捅鼻子的照片。注意Ayass要求拍摄时不能拍到工作人员
  * 回国必要性证明（F签坐AA127应该不用写，其他身份/始发地可以看[这里](https://docs.qq.com/doc/DSE9Ga2dudG9jZkZu)做个参考）
* 健康码的小程序做得比较捞，每次审核只能上传十张照片，超过十张可能会在提交申请 -> final confirmation的这一步被吞图。有时图片尺寸过大可能还会上传失败，这种情况的话把图片删了重新上传多试几次即可。注意上传的图片会被压缩，提交之前最好检查一下护照/I20上的字样是否清晰
* 健康码的有效期是48小时，不过还是建议一拿到检测报告就提交申请以防出问题
* 申请海关码
  * 申请的portal也是微信小程序，叫“海关旅客指尖服务“
  * 申请海关码不需要人工审核
  * 这个二维码是在中国海关入关的时候用，有效期是24小时，所以建议在值机前几小时填写
* 健康码在check in的时候柜台工作人员会查，在登机之前会检查两个码，要去检票口敲个章才能登机

## 到中国机场后怎么走

本来打算关于这个topic写一整个section的，但是落地了才发现检测区域不让拍照... 虽然人多，但是工作人员也多，所以落地了把所有材料（护照，健康码，海关码，疫苗接种证明）拿在手上，听从工作人员的指挥就行。

## 隔离

* 上海目前对本地人的政策是14+7（14天集中隔离，7天居家隔离），外省的政策不一
* 我坐的AA127是6/29 14:50落地，解除隔离的时间是7/13 14:00
* 隔离期间第{1, 4, 7, 14, 16, 21}天 (1-based indexing)要做核酸检测，其中第一天/在机场坐的那次是捅鼻子/喉咙，后面的都是鼻子

## 附录

### 申请健康码需要的十张图片（顺序同上）

#### 1. 护照首页 + 签证页，这个就不附图了🐶

#### 2. 疫苗接种声明书 + CDC小白卡 + 疫苗接种机构证明

![](/files/-MdB5AbK0h-VZmQIFa27)

#### 3. I20，也不上图了

#### 4. 回国航班的Itinerary

![](/files/-MdB5n2x1lnVWp080Eb5)

#### 5. Ayass/检测机构预约邮件

![](/files/-MdB65PscxNghJkWJJyo)

#### 6. PCR报告

![Not Detected == Negative == We are good](/files/-MdB6P59wkJvZMit5pYS)

#### 7. IgM报告

![我这里测出来是S阳N阴（左边是S，右边是N，Reactive == Positive，Non-Reactive == Negative）](/files/-MdB6i1BXYAgrhZci3XJ)

#### 8.1. Ayass Bioscience Inc 八号楼楼下，COVID-19 IgM Testing牌子前

![](/files/-MdB7hVWhiacPfiZyffr)

#### 8.2. IgM Sample Collection门前

![](/files/-MdB8bN4ekOH5iUkmMUf)

#### 8.3. Ayass Bioscience门前

![](/files/-MdB8fQeTm69-4lIt-Zo)

### 上传申请材料的位置

![](/files/-MdB3O0T4hSdR75cJ1n-)

### 费用总结

* 来回机票：现在最便宜的直飞机票差不多是两千五朝上，加上在美国飞去检测地/在中国隔离完飞回老家的机票钱，这里一共算6000
* 检测费用：Ayass加测N蛋白一共是450
* 检测地住宿+交通：2-3天酒店，一晚100刀左右；租车/请司机一天100+，因为我有人同行，所以这里的费用加起来算250
* 上海隔离酒店：我住的是400一天包吃住，一共是14天 -> 800刀左右
* 这么看一趟下来是7500刀左右，肉真的疼:\_(，不过回去能陪陪两年没见的爸妈/同学/家里老人，这些是钱买不到的


# How to Transfer Credits Back to UW-Madison

**Update: As of Dec 2020,** [**UW College Online is no longer offering courses**](https://online.uwc.edu/)**. The best alternative at this point seems to be** [**UC Berkeley Extension**](https://extension.berkeley.edu/)**.**

Looking back, I have a mixed feeling about my course plan at UW-Madison. Taking courses that's not in my area of interest is fun and challenging by pushing me out of my comfort zone. The problem is that as of spring 2020, the tuition for taking 12-18 credits as an international student is about $19,500, which means each 3-credit class costs you $3,250.

This sucks, a lot.

The beauty of online classes is that they are **cheap**. I took BIO141 at [UW College Online](https://online.uwc.edu/), and the 3-credit class cost merely $789. Aside from that, the grade you receive will not be calculated into your UW-Madison GPA, and the credit can be transferred as long as you get a C or above, so you don't need to worry about breadth classes hurting your GPA at UW-Madison. With the credits you save, you get to take more major-related classes during a regular term.

Just as a recap, the pros of transferring credits back are:

* It saves you a lot of $;
* It relieves you from worrying about your GPA being held back by a biology class;
* You get to take more major-related classes at your home campus;
* If you take enough classes, you could graduate early and potentially save $$$!

**Note that the information I'm sharing may be outdated, so please refer to the latest policy before you take a class. Always reach out to your advisor for any questions!** As of 2020/01/30, I'm looking at [this page](https://registrar.wisc.edu/transfer-your-credit-to-uw-madison/) as I write down this blog.

## Step 1: Make sure a course transfers

Generally, there are two types of courses you can take: courses **in** the Wisconsin system and courses **not** in the Wisconsin system.

For courses in the Wisconsin system, see the [Credit Transfer Wizard](https://www.wisconsin.edu/transfer/wizards/) to check if a course transfers (and what class it is equivalent to). Most of the classes at [UWC](https://online.uwc.edu/) are transferrable.

For a course that's not in the Wisconsin system, there is a [Transfer Equivalency Database](https://apps.admissions.wisc.edu/apply/transfer/ted/) that you can refer to when taking a class in Illinois and Minnesota. If you don't see your institution (say you are taking a class at UC Berkeley), you need to submit a course equivalency request to the [Course Evaluation Service](https://apps.admissions.wisc.edu/ces/ces_portal.php) before you enroll in the class.

![This is what the course evaluation form looks like](/files/-MNWy5ESduQZjSTtJdIZ)

The CES will be available from March 1 to May 15 for summer term evaluation and from November 1 to December 1 for winter term evaluation. Once you receive a confirmation from the Office of Admissions and Recruitment (like the picture shown above), just sign up for the class and move on to the next part...

## Step 2: Get a C or above

Yeah, it's that simple... As the grade of this class will not show up on your UW-Madison transcript, you just need to pass this class for the credit to transfer. Seriously, don't screw this up. (I almost did XD)

## Step 3: Submit a transcript

After weeks (might be hours though, if you know you know) of binge studying for this class, you got your grade! Now, you just have to request an official transcript and send it to the Office of Admissions and Recruitment at the address below\...

```
Office of Admissions and Recruitment
University of Wisconsin–Madison
702 West Johnson Street, Suite 1101
Madison, WI 53715-1007
```

... and the course will be in your record in a couple of days!

Update: Apparently UW-Madison accepts digital transcripts now, so you can order a transcript in PDF.

Update: You need to send an email to `crediteval@registrar.wisc.edu` once you've submitted the transcript for the registrar people to add the class to your student record.

![Requesting a PDF Transcript](https://github.com/ruipeterpan/blog/blob/master/content/posts/images/20200130-2.jpg?raw=true)

## As we close to an end...

Hopefully, you got your class officially transferred. If this guide helped you (even by just a little bit), please let your friends know about this opportunity so they don't have to tear their hair out over a class they hate. I for one wouldn't have known this without [Shawn Zhong](https://shawnzhong.com/), so kudos to him for sharing this information!


# Resources on Learning Academic Writing (for Computer Science)

* Tips on Writing a Research Paper by Thomas Reps ([talk video @ PLDI](https://www.pldi21.org/prerecorded_plmw.2.html))

{% file src="/files/nDFh9fFQmJu9jb8Ad1zG" %}

* The Most Common Habits from more than 200 English Papers written by Graduate Chinese Engineering Students by Felicia Brittman ([pdf](https://www.chrisyttang.org/assets/misc/The%20Most%20Common%20Habits%20from%20more%20than%20200%20English%20Papers%20written.pdf))


# Towards applying to CS Ph.D. programs

## Introduction

After going through and learning from tens, if not hundreds, of blog posts on applying to Ph.D. programs, it would almost be inappropriate if I didn't write down something and throw in my two cents about this exhausting but ultimately self-enriching and fascinating process.

This blog post will hopefully be a useful guide to the students who are planning to apply to Ph.D. programs. I will try to give you a glimpse of what the application process looks like and offer some advice on how to best approach, embrace, and enjoy this unique journey. I hope this helps you! If so, please consider paying it forward, maybe start by giving a hand to students in your research group who will be applying for Ph.D. programs.

Note that whatever I write down is biased because of my background and experience. For context, I grew up and went to high school in China, and then did my bachelor's in the United States, majoring in computer science and mathematics. During my undergraduate, I worked on systems (scheduling/cluster resource management) for ML starting from the summer of my second year, and I fully committed to doing a Ph.D. in my third year. My research interests fall under the big topic of "systems and networking", and for my Ph.D. application, I applied to professors whose areas of interest range broadly across all system aspects of big data, including systems for ML (training, inference, video analytics, etc.), ML for systems (congestion control, video streaming, etc.), cloud computing/data center resource management (e.g., serverless, scheduling training/inference workloads), etc. Also, my honest opinions can be straight-up wrong, so take everything with a grain of salt.

## Overview

A CS Ph.D. application package should be the culmination of your previous academic career. It usually includes three recommendation letters, a curriculum vitae (CV), a list of publications (if any), a research statement of purpose (and possibly another personal/diversity statement), your college-level transcripts/GPA, standardized test scores, and a bunch of other personal information.

While your technical ability is instrumental to your Ph.D. application, doing the application right is also critical but is often neglected. This blog post aims to point you to some common practices for wrapping your application package nicely with the following chapters:

* **Chapter 1: Why do a Ph.D. at all?**
* **Chapter 2: Narrowing down the programs/professors of interest**
  * §2.1: How do you put up the big list of POI?
  * §2.2: How do you narrow the list down?
* **Chapter 3: Getting in touch with the POI**
  * §3.1: Make a CV and a webpage first
  * §3.2: How and when should you reach out?
* **Chapter 4: Asking for recommendation letters and sending out requests**
  * §4.1: Who should you ask for letters?
  * §4.2: How should you ask for letters?
  * §4.3: What's next?
* **Chapter 5: Writing up the statement of purpose**
  * §5.1: How should you write a statement?
  * §5.2: What should be in your statement?
* **Chapter 6: Preparing for interviews**
  * §6.1: What's a typical interview like?
  * §6.2: What to do before, during, and after an interview
* **Chapter 7: Towards good mental health during the application season**
  * §7.1: Before sending out applications
  * §7.2: After sending out applications
* **Chapter 8: Choosing a Ph.D. program (WIP)**
  * §8.1: Understanding your offer
  * §8.2: What to do on the visiting day
* **Appendix I: My application timeline**
* **Appendix II: Meta-references**

Each chapter includes some of my personal opinions followed by a list of references. Moreover, my application timeline and some meta-level references are provided in the appendix.

## Chapter 1: Why do a Ph.D. at all?

In my opinion, you should do a Ph.D. if you have already done some research, been through the ups and downs (or at least know a bit about what they are like), and still absolutely love doing research. If you are applying for a Ph.D. because your Asian parents are forcing you to do one or if you want to stay in academia because you didn't get an industry job, think twice: it's a huge commitment.

This is a million-dollar question and I don't feel entitled to write more about this right now, so please take a look at all the references below.&#x20;

### References

* [Life after the PhD](http://people.csail.mit.edu/fredo/LifeAfterPhD.pdf) by Prof. Frédo Durand
* [The PhD Grind](https://archive.org/details/the-phd-grind-philip-guo) by Prof. Phillip Guo
* [The CS Assistant Professor Handbook](https://vijay03.github.io/asstprofbook/) by Prof. Vijay Chidambaram
* [Reasons to Pursue a Ph.D.](https://jxyzabc.blogspot.com/2011/12/reasons-to-pursue-phd.html) and [CS Grad School Part 1: Deciding to Apply](https://jxyzabc.blogspot.com/2008/08/cs-grad-school-part-1-deciding-to-apply.html) by Jean Yang
* TODO: Add more references
* [读博，你真的想好了吗？- 张焕晨的文章 - 知乎](https://zhuanlan.zhihu.com/p/372884253) (Are you really sure about doing a Ph.D.? by Prof. [Huanchen Zhang](http://people.iiis.tsinghua.edu.cn/~huanchen/))
* [读博前的思考](https://zhuanlan.zhihu.com/p/548781775) by Xiuyu Li (Some thoughts before doing a Ph.D.)

## Chapter 2: Narrowing down the programs/professors of interest

Most people agree that when applying to Ph.D. programs, the advisor is the most significant factor (even more so than the school/department itself). Thus, picking awesome professors/person of interest (POI) is arguably the most crucial part of the application process: pick good (in terms of research interest match and personality match), and you might be happy for life.

95% of the students I know apply for 5-15 programs and they typically target 1-3 POI for each program. In my case, I checked out \~50 professors in my field of interest and ended up boiling the list down to \~30 professors that I especially liked from \~15 schools. This chapter will try to answer two questions: (1) how do you put up the big list of professors that you are generally interested in working with, and (2) how to narrow the list down to those who you are particularly interested in working with?

### How do you put up the big list of POI?

* Check out your advisor's former lab mates, recent collaborators, and direct connections in their network. These people are likely to share similar interests with your current advisor, plus they know your advisor personally, so these folks should be fun to work with, assuming you love what you are doing right now.
* The math folks have a great thing called [The Mathematics Genealogy Project](https://www.mathgenealogy.org/) where you can see this tree of academic relationships. You can also trace the tree of professors in cs, starting from very senior professors who work in your field of interest, and then go down the genealogy tree to look for those holding tenure-track positions. For me, I started with Ion Stoica. Fun fact: as many as six of my professors of interest have had direct connections with Ion!
* [CSRankings](https://csrankings.org/) is also a great place to visit. You would want to first list the target conferences that you mostly read papers from (for me, it was SOSP, OSDI, MLSys, EuroSys, ATC, SoCC, NSDI, and SIGCOMM). Then, go to CSRankings, select your target conferences, check out the professors from each school who has published in these venues, and go through them one by one. Unless you already have a specific topic that you would like to work on, IMO you should be open-minded in this part of your search and try to check out as many professors as possible. When I was going over csrankings.org, I had also marked professors who publish in venues like SIGMOD & VLDB, SIGMETRICS, HPC conferences, and ML conferences like ICML. Although I still ended up applying to system professors, it was fun to get to know what folks are working on outside of your main areas of interest.
* Take the same list of venues that you like. Then, take a look at the program committee of the recent conferences, and go to their personal sites one by one. Doing so will produce a list that overlaps very much with the one you got from csrankings.org, but csrankings.org may not have the most up-to-date information.
* Talk with other people, e.g. current students and alumnus in your lab, your current advisor, random people you met online, etc. You will genuinely get a lot from this! My personal story is that I didn't consider applying to the program I ended up committing to until after a friend of mine strongly recommended that I shoot this POI an email. I ended up getting in touch with the POI and found that I liked him a lot and that he was hiring. So yeah: talk to people!
* Follow a bunch of professors on Twitter. This generally helps with catching up with the latest news in academia. For example, professors post hiring ads and tweet about their opinions on various things ([example](https://twitter.com/scottniekum/status/1503112889284313088)), and incoming professors tweet about their employment (this information is usually not listed on any official website).
* Although the POI is arguably the most crucial factor for a happy (and successful) Ph.D., the program itself has to be taken into consideration during your application. People usually apply to some schools at their level (match), some schools above their level (lottery), and some schools below their level (safety). To that end, look at students with a similar background as yours and refer to their school list and application results. You can also have your current advisor go over your school/POI list and provide some feedback. It is also a good idea to modify your school list based on your feedback from cold-emailing the professors.

### How do you narrow the list down?

On the one hand, you should work with people whose research interests truly excite you. Although I had a few professors who work on databases/HPC on my first list, I ended up throwing them away because I prefer some topics over others.

On the other hand, you should not work with people who are bad advisors. Advisors who ghosts/abuses students are a big no-no! I used two approaches to filter out these professors:

* Check out their RateMyProfessors reviews. I get that teaching isn't for everyone and some professors put more emphasis on doing research, which is fine by me -- but ultimately, I honestly don't want to be working with someone who is a 1.2, because I think if a professor fails to create a supportive learning environment for their students, then it's unlikely for them to do so for their advisees. IMO a series of reviews that start with a 1 is somewhat of a red flag.
* Talk with their current students or search for their posts online. More practically, talk with people who might have heard about some bad news -- they tend to travel very fast. There has to be something wrong if many Ph.D. students quit a lab halfway through.

And also, you should be careful about applying to a program that only has one professor you are interested in working with, because a lot of things can happen in five years, e.g., your advisor doesn't get tenure and goes to the industry or gets poached by a school you hate. In that case, you will most likely switch to a different advisor in the department. Be prepared for contingencies.

### References

* [finding CS Ph.D. programs to apply to](https://www.youtube.com/watch?v=hOSl3xPmHiQ) by Prof. Phillip Guo
* [How to pick a grad school for a Ph.D. in Computer Science](https://vijayc.medium.com/how-to-pick-a-grad-school-for-a-phd-in-computer-science-a5ce7dceb246) by Prof. Vijay Chidambaram

## Chapter 3: Getting in touch with the POI

Now that you have narrowed down a list of 10\~30 professors you want to work with, it is time to get in touch with them. Some people don't bother to do this at all -- I do not recommend this, as I know professors who will only skim through your application package if you haven't reached out. Plus, getting in touch with them helps you figure out how many students they are taking this season, their ongoing/future interests, how enthusiastic they feel about your background, etc.

### Make a CV and a webpage first

> If you are not able to create a web page, you probably shouldn't be applying to CS graduate programs.  -- David Evans

Before you start emailing people, I strongly recommend making a personal webpage. Quite a few professors have also talked about the importance of personal sites in academia. A good Ph.D. student/research should be visible in the community/on the Internet, and Linkedin/Facebook/department-generated webpages are just not enough for that. Some good templates include [Jon Barron's website](https://jonbarron.info/) and [academicpages.github.io](https://academicpages.github.io/). Prof. Timothy Roscoe recommended including the following information on a personal webpage:

* A picture, preferably a recent one that actually looks like you. If you have a goofy picture that you really like, consider hiding it behind your main picture and make it show on hover.
* A list of publications. No pressure if you haven't published, especially for undergraduates who do research in systems.
* Some biographical information:
  * How long have you been a student? And how long have you got left?
  * Who do you work with? And what do you work on?
  * Random (but non-embarrassing) details for color.

Although a personal website makes you easily accessible on the internet, you should also spend some time putting up a high-quality CV. Looking at other people's CVs and trying to emulate their formats/highlights most certainly helps. Your POI is not the only person who will look at your CV, especially if the Ph.D. admission committee has a huge influence on the admission decision: other professors and senior graduate students (who might be focusing on different research topics) will also go over your CV, and it needs to impress them enough during the first round of the admission process for a professor to even see your package (?). When you cold-email the professors, attach your CV.

### How and when should you reach out?

The de facto way to get in touch with professors is to cold email them. If you are lucky to have the opportunity to meet with professors during a conference or have already known the professor, that would be great! But most people send emails anyway unless they know a POI very well, so getting the email (and the first impression) right is vital. Here are a couple of tips.

* **Figure out the timing.** My recommendation is to start as early as you feel comfortable once the fall semester begins and professors start to check their inboxes regularly so they will hopefully have enough time to get back to you. If you do this a few weeks after you submit the application, the POI might already have a batch of good candidates in mind (but still, better late than never). On a more fine-grained level, a good trick is to send a timed email scheduled for something like 8 AM, so that your email will be on the top of the inbox when professors check their inboxes.
* **Be concise but on-point.** There is a reference below that covers how to send out cold emails, but take it with a grain of salt, as that's more for emailing professors for undergraduate research opportunities. My suggestions are:
  * You must mention who you are working with right now and your research interests.
  * You must mention why you would like to work with the POI, and you should probably mention a bit of their existing work and why you like them. Better yet, go in-depth and ask a technical question/ask if xxx is an interesting idea for possible follow-up work.
  * You must attach your CV. I had attached a draft research statement just in case the POI has some time to take a look, but in retrospect, I don't think any of them had the time to read it (?).
  * You should probably put a link to your webpage in your email signature. I had also included a link to my calendar for the easier scheduling of meetings.
* **Know the email etiquette** (there are two articles in the references that cover this). To that end, triple-check your email for missed attachments and spelling mistakes before sending it out! Better yet, have your roommate proofread it.
* **Within a school, reach out to professors one at a time** and only move on to the next professor if the previous one doesn't get back to you in a week or so. Please don't reach out to five professors whose offices are right next to each other at the same time -- they talk.
* Please **don't send three emails in a day** to try to catch someone's eye. Professors are busy, but they will get back to you if they see a good fit. They might forget about things, and in that case, send a kind reminder after some time (say a week?) of not hearing back.
* **Use your institutional email account.** Gmail/outlook addresses are fine but don't use your QQMail.

### References

* [How to write an Academic CV for a PhD Application](https://www.discoverphds.com/advice/applying/cv-for-phd-application)
* [Resumes & Cover Letters for PhD Students](https://ocs.fas.harvard.edu/files/ocs/files/phd_resume_cover_letters.pdf)
* [Curriculum Vitae Tips and Samples](https://grad.illinois.edu/sites/default/files/pdfs/cvsamples.pdf)
* [Software developer resume template in Latex](https://github.com/sb2nov/resume)
* [How to Cold Email a Professor](https://research.berkeley.edu/how-cold-email-professor)
* [How to Email a Professor](https://academicpositions.com/career-advice/how-to-email-a-professor)
* [Advice (on Writing Emails) for Prospective Research Students](https://uvasrg.github.io/prospective/) by Prof. David Evans
* [千万别犯写邮件的大忌](https://www.zhihu.com/question/68514971/answer/469896862) (The DON'TS when writing emails)

## Chapter 4: Asking for recommendation letters and sending out requests

Most CS Ph.D. programs ask for three letters of recommendation.

First of all, IMO it is necessary to understand what the other side looks like about this recommendation system. Except for the actual letter itself, your recommenders will also be asked about your clarity of goals for graduate study, English skills, creativity, etc. The "scores" recommenders can give are chosen from truly exceptional (top 1%), outstanding (top 10%), above average (top 25%), not applicable/unable to respond, etc. Some other questions include ([source](https://www.1point3acres.com/bbs/forum.php?mod=viewthread\&tid=581428)):

* How long have you known the applicant?
* What group are you using for comparison?
* Admission recommendation to the program is {strongly recommended, recommended, recommended with reservation, etc.}

### Who should you ask for letters?

Moving on. Before sending out requests for letters, you have to figure out who your letter writers will be. It would be best if your letter writers are:

* **Well-known in the research community.** Having a senior professor/big name write a strong recommendation letter for you helps *immensely*. Either that or some junior professors who are actively publishing in a relevant field. These professors have a strong network of connections, which is valuable considering how critical connections are. In contrast, a letter from a postdoc is probably less helpful -- but if a postdoc is writing a letter for you, try to have them co-write a letter with their advisor.
* **Someone who knows you well.** Whoever has worked with you extensively will have plenty of chance to know you and see that bright side of yours. If you mainly worked with Ph.D. students/postdocs during your research and a professor who wasn't very hands-on toward your research is writing the letter, I would suggest coordinating with those people so that the professor can put some insights from the others into the letter.
* **Preferably someone in academia** instead of the industry, and if they are from the industry, they should be involved in doing research (e.g., leading a research team. At least they should have a Ph.D. degree?). I'm not too sure about this, but rumors say that letters from people in the industry are pretty much useless. This makes some sense because the qualifications of a good Ph.D. student are somewhat different from those of a good intern, and letter writers from the industry would have no idea of what the program committee is hoping to see. Note that when I say people from the industry, I'm referring more to an average SWE intern's mentor -- if Lidong Zhou or Amar Phanishayee writes you a letter, then obviously it's very good!

### How should you ask for letters?

Now that you know who you will be asking, it would be best to let them know about it. Here are some tips for doing that.

* **Ask early.** You don't need to know your exact school list when you send your first email request, but you should at least give people a rough idea of the number of letters and when they will be due. Professors are busy, like really busy. Please ask for letters as early as possible so they can plan accordingly. Two months in advance is better than one, and one month is better than two weeks. Imagine being a professor who's going through finals week and a couple of conference deadlines, and boom, five students show up to ask for a total of 100 letters that are due in one week. Just thinking about that makes me uncomfortable. Besides that, if you apply to 30 (?!) programs, some professors may not be too happy about submitting all those letters and will only agree to submit 10 of them. In that case, you should, of course, ask for some other professors to fill in the gap, but anyway, you don't want to know about this a week before the application deadline.
* **Communicate clearly.** It would suck if professors finish your letter but don't know where to submit them. My advice is: (1) It is preferable to send official requests through application portals in batches, so those emails don't get lost in the professor's email inbox. (2) Create a table in Google Sheet to keep track of the programs you are applying to, the application deadlines, and the status of the requests. For example, knowing the date of when you sent a request will be helpful when a professor looks up that email in their inbox. Of course, share this google sheet with your letter writers. (3) Send email reminders to remind professors about an upcoming application deadline, preferably a week in advance, if not more. Note that for most programs, the deadline for people to submit letters is some time after the application deadline, so don't panic if a professor uploads a bit late.
* **Provide enough information.** Once a professor agrees to write you a letter, attach many files when you get back to them, e.g., your cv, draft SoP, transcript, and final project report. Some professors will also ask you to provide a list of relevant personal information, e.g., professors you have taken classes with and the grades you got, major accomplishments, etc. Besides those, remind professors about how your opportunity with them helped you grow as a scholar.
* **Ask for strong letters.** It goes without saying that strong letters from professors say a lot about who you are as an applicant -- they will make a difference in your application. When you ask for letters, explicitly ask for strong letters, so that professors who will otherwise write average letters can give you the chance to pivot.

### What's next?

And last, you should write them a thank you letter. Send a small gift (something under $20 should be fine depending on the department policy, e.g., a box of chocolate provided that they are not allergic to chocolate :P). When you commit to a program a few months later, also remember to send another email to let your writers know about the good news -- they will like it very much!

### References

* [How to get a great letter of recommendation](https://matt.might.net/articles/how-to-recommendation-letter/) by Prof. Matt Might
* [Requesting a letter of recommendation](https://homes.cs.washington.edu/~mernst/advice/request-recommendation.html) by Prof. Michael Ernst
* [How to write a letter of recommendation](https://homes.cs.washington.edu/~mernst/advice/write-recommendation.html) by Prof. Michael Ernst

## Chapter 5: Writing up the statement of purpose

The statement of purpose (SOP) is arguably less important than some other things in your application package, but still, it is one of your first files that get looked at. First impressions are critical: IMO, an exceptional SOP will not make your application stand out as much as one would do in undergraduate admissions, but a bad SOP might ruin your application.

### How should you write a statement?

* Before you start writing your statement, make sure to **read through many other people's (good) SOP.** Prof. Phillip Guo had shared a few statements that are really nice examples. Although he has taken down most of the content and has asked people not to distribute them, some of these statements are scattered across the Internet, and I'm sure you are good at Googling stuff. There are also some other good examples online.
* **Start early to write up the first draft.** I was applying to a pre-doctoral summer program so I was very lucky to have my first draft ready in the summer before the application season, but still, I didn't finalize my statement until a few days before Dec 15: it takes a lot of time to do the endless revisions. Another thing is that except for the research statement of purpose, different schools might ask for additional materials such as diversity statements, short answers to random questions, etc., so make sure to figure out those requirements way before the deadline.
* **Go through a lot of iterations.** Just like writing up a paper, the first draft is guaranteed to be bad, and good papers go through O(10) rounds of revisions. To that end, try to get a lot of feedback by having as many people read your statement as possible: writing center, (former) lab mates, high school classmates who are majoring in computer science, roommates, etc. Five people can probably offer you 50 suggestions, and if you go by 20 of them, your sop will rise to a new level.

### What should be in your statement?

Keep in mind that professors will use this to judge your writing skills. If you are good at writing, it's time to show off. Otherwise, at least make sure there are no grammatical/spelling mistakes: that would look terrible. Consider using something like Grammarly to fix those mistakes and clarify your writing.

A significant portion of the statement should be on your past research. People usually spend one paragraph for each project, and the ranking is either by relevance or by date. IMO if your projects connect well, by date is a more natural way since you get to tell the story behind your motivation to get a Ph.D., but don't worry if you rank them by relevance.

In each paragraph, describe the project (e.g., collaborators, short background & motivation, major technical contribution, your contributions & takeaways). Keep in mind that the people reading your statement might be working in a different subfield (I don't suppose the bioinformatics people will know anything about Cuckoo hashing), so don't just copy-paste the abstract of your past papers.

Depending on how much space you have left, IMO you should talk about your research interest using at least one sentence and at most one paragraph. If possible, IMO you should identify some rising research topics in the next few years and describe why they pique your interests: the POI might resonate with you if you come up with some good ones.

It's a good idea to mention which professors you would like to work with. Try to aim for 1-3 professors in the statement, and briefly discuss why you are interested in working with them.

### References

There are also a bunch of blog posts by people who know more about writing up statements than I do, so make sure to take a look at these articles, including but not limited to:

* [Inside Ph.D. admissions: What readers look for in a Statement of Purpose](https://nschneid.medium.com/inside-ph-d-admissions-what-readers-look-for-in-a-statement-of-purpose-3db4e6081f80) by Prof. Nathan Schneider
* [How to Write a Statement of Purpose for Grad School](https://swapneelm.github.io/how-to-write-a-statement-of-purpose-for-grad-school) by Swapneel Mehta
* [Tips for Writing a Statement of Purpose](https://users.ece.cmu.edu/~mabdelm/statement-of-purpose-tips.html) by mabdelm
* [Ph.D. Statement of Purpose](https://blog.nelsonliu.me/2020/11/11/phd-personal-statement/) by Nelson Liu
* [Writing a Statement of Purpose](https://djunicode.github.io/2018/10/16/writing-a-statement-of-purpose.html) | The Unicode Blog
* [What to avoid and what to mention in your SOP, targeted especially at international students](https://twitter.com/vj_chidambaram/status/933388419589459969) by Prof. Vijay Chidambaram
* [Jean Yang's Statement of Purpose](https://github.com/jeanqasaur/academic-application-materials/blob/master/phd-application-2007/personal_statement.pdf) from her blog [CS Grad School Part 4: Applications](https://jxyzabc.blogspot.com/2008/08/cs-grad-school-part-4-applications.html)
* [\[Graduate School\] Personal Statements](http://cwfletcher.net/Pages/SoP.php) by Prof. Christopher Fletcher
* [What’s a Good Statement of Purpose?](https://ed.stanford.edu/sites/default/files/statement-of-purpose_u.d_2013.pdf) by Eamonn Callan
* [How to Write a Bad Statement for a Computer Science Ph.D. Admissions Application](https://www.cs.cmu.edu/~pavlo/blog/2015/10/how-to-write-a-bad-statement-for-a-computer-science-phd-admissions-application.html) by Prof. Andy Pavlo

## Chapter 6: Preparing for interviews

Congratulations on sending out all of your applications! Take a little break first, both physically & mentally. Then, it's time to start preparing for interviews!

### What's a typical interview like?

* **Duration:** A typical interview lasts an hour or so. The ones I had been in lasted as short as 20 minutes and as long as two hours.
* **Content:** You usually start with some chit-chat, followed by a quick self-introduction. Then, the POI will likely ask you to talk about your research project(s), during which they will evaluate both your hard and soft skills. Afterward, you can ask the POI some questions, including the lab culture & dynamics, the graduate program, their research, etc.
* **Will there be coding/technical questions?** (???) It depends, but most professors don't ask these kinds of questions. From what I've heard, there are professors who ask about things like page fault handling in operating systems or how web indexing works (those are extreme outliers though). But there will of course be technical questions for your past research projects!

### What to do before, during, and after an interview

* **Before an interview: Go over your statement and resume**, since the professors will likely refer to them if they ask questions about you. If you included a topic that you are less familiar with in your future research interests, it doesn't hurt to delve a little bit deeper into that topic.
  * For every project listed on your resume, prepare the following:
    * One sentence that summarizes the project. This is like the title of your project but maybe with some more info for context.
    * A 3-minute introduction that expands a little bit more, say on the background, motivation, technical contributions, and results.
    * A 10-minute overview of the project that can be expanded into a 30-minute discussion. Totally write stuff down beforehand if you feel like it.&#x20;
    * Your contributions to the project. Undergraduates often get carried by Ph.D. students/postdocs in their research, so it's important to highlight what you did and what you got out of a project.
  * Note that the interviews vary in duration, and you might have multiple projects to talk about (I only focused on the most significant one), so be flexible about the timing.
* **Before an interview: Read your POI's work.** It's ok to prioritize the POI who you are really interested in or who showed great interest in you. They won't ask you about the technical questions in their past research projects, but still, getting to know about what a POI used to work on shows your seriousness and enthusiasm. Different students spend various amounts of time on this phase of the preparation, but I think you should at least do the following. For each POI:
  * Do a quick pass through the title/abstract/collaborators/venues of all their past work.
  * Pick 2-4 of their projects to dive deeper into. I had focused on (1) their most highly-cited paper, (2) their most highly-cited first-author paper, (3) a highly-cited paper in the recent two years, and (4) a recent work that you are particularly interested in, either because you can relate to the project regarding motivations/techniques or because you are genuinely captivated. And by diving deeper into it, I meant going over all the figures, learning about the background/motivation/nuggets (high-level contributions and techniques), etc.
  * If you like this POI very much, you can totally go over the technical details of some of their papers. Because why not? Reading papers are fun! If not, then you should probably think twice about your application.
  * Take a quick skim at their Ph.D. thesis, especially the acknowledgments section to know more about them as a person.
* **Before an interview: Do mock interviews** (or research presentation talks) with your friends/labmates. I didn't, so I totally screwed up my first interview, but it was a good practice and I got the chance to learn from my mistakes. In retrospect, I really should have done an actual mock interview with some labmates and had them ask all kinds of questions.
* Before an interview: Make sure you have a quiet environment, a stable internet connection, and a good microphone. If you have noisy roommates or bad routers, you might want to reserve a quiet study room in a library in advance.
* Before an interview: Dress properly. FYI, the chats are mostly casual unless the interviewer specifically mentioned a (multi-person) serious interview.
* **During an interview: Chill out, and be yourself.** IMO a significant purpose of interviews is so that you can get a sense of the vibe/chemistry between you and the POI, so don't force things like saying you are interested in something that you are not.
* **After an interview: Send the POI a thank-you email.** If you had a good chat, maybe it's time for some follow-ups. Anyway, you should let the professor know your thankfulness and reassure your enthusiasm for collaborating with them.

### References

* [What is a typical interview (informal chat) for a PhD in computer science like?](https://www.quora.com/What-is-a-typical-interview-informal-chat-for-a-PhD-in-computer-science-like-What-do-the-professors-generally-ask-Do-I-need-to-have-concrete-research-ideas-of-my-own-Should-I-read-a-lot-of-research-papers-by-the-professor)
* [How To Prep For A Grad School Interview](https://jakec007.github.io/2021-04-02-cs-grad-school-interview/) by Jake Chanenson
* [TOP校CS Ph.D.面试经验教训总结](https://www.1point3acres.com/bbs/thread-628184-1-1.html) (Tips for interviewing at top-tier CS Ph.D. programs)
* [关于我自己观察到的中国学生在PhD面试过程中的一点特点](https://www.1point3acres.com/bbs/thread-590002-1-1.html) (My observations on Chinese students' traits in Ph.D. interviews)
* TODO: add more references that are in English

## Chapter 7: Towards good mental health during the application season

My mental health was surprisingly good during my application season (probably because I only took three credits in the fall and spent most of my time polishing up a paper & applying to Ph.D. programs). Although I am no expert, here is some advice for staying positive during the six months. If things get too tough, please talk to the professionals.&#x20;

### Before sending out applications

* **Make plans to abide by** so that you don't stay up and rush things. "Rushing is the path to the dark side. Rushing leads to staying up. Staying up leads to bad health. Bad health leads to suffering." - Master Yoda
* **Maintain a consistent, healthy sleeping schedule.**
* **Talk with supportive people** around you and be supportive of each other. You are not alone, and every applicant is fighting the same battle. Surround yourself with people who can help relieve your anxiety, not trigger them.
* **Workout.** Participate in team sports, work on bodybuilding, take a random walk outside, etc.
* If things go well, this will be your second last semester as an undergraduate. Since most people go to a different school for Ph.D., it will also likely be your last fall/winter in your current city. On weekends, spend some quality time with your friends here to create some enjoyable memories for future reminiscing. If you are studying at UW-Madison, [here is an article](https://zhuanlan.zhihu.com/p/425849399) I wrote on the 20 must-dos before you graduate.

### After sending out applications

After sending out applications, you will likely have huge chunks of free time since the semester is over and Christmas is coming up. Although the interviews will be coming shortly, IMO you should first **take a week-long mental/physical break.** Congratulations on submitting all those applications!&#x20;

* Once you get back from your mental break, you should get back to studying. You likely won't have a lot of things to work on, and your motivation might be low -- after all, your application is already out, and there is not much you can do to make it drastically better. Instead of spending all your free time being anxious about the applications, **try to divert your anxiety**, say by developing a new hobby. Read a book or something, or learn to cook.
* [1point3acres](https://www.1point3acres.com/bbs/). [zhihu.com](https://www.zhihu.com/question/379814619), and [The GradCafe](https://www.thegradcafe.com/) have a lot of good information and stuff, but please consider restraining yourself from visiting these sites too often. Social media takes a toll on people.
* Also, it might be worthwhile to **turn off instant notifications** for your email inbox and check it a few times a day at regular times.
* **Compare to yourself, not others. This is in general a great suggestion on how to live a happy life.**
* *Spider-Man: No Way Home* was in theatres during my application season, and it had a great line: "**If you expect disappointment, then you can never really get disappointed**". Don't get too hyped up if a POI reached out to you or if you did well in an interview. Otherwise, you will feel really bummed when you get rejected.

### References

* [How to effectively deal with Imposter Syndrome and feelings of inadequacy](https://academia.stackexchange.com/questions/11765/how-to-effectively-deal-with-imposter-syndrome-and-feelings-of-inadequacy-ive) | StackExchange Academia

## Chapter 8: Choosing a Ph.D. program (WIP)

You now have multiple offers in hand! Very nice.&#x20;

By the time you sent out the applications, you should have a rough idea of your preferences for all the programs. But, it's not a good idea to rush your decision. Grad schools usually request you to make a decision by Apr 15, and you should spend at least some time making up your mind. After all, this is one of the biggest decisions in your life. That being said, if you are set to commit to a program, withdraw/decline your other offers as soon as possible so the other POI can extend the offer to other students on the waitlist.

The first thing you would want to do is to do your own research. My personal priorities are: personality/advising style match with the POI >> research interest match with the POI > department ranking in your area of studies > department overall rankings > location/weather > amount of financial support & overall quality of life > overall ranking of the university. There are also other factors to consider depending on your future plan. For example, if you are looking to go into the industry after graduating Ph.D., you might want to go to a program where doing summer internships is easy and encouraged.

Then, talk with a lot of people regarding your offers and decisions: parents, relatives who work in academia, significant others, friends, people on the internet, labmates, current advisor(s), labmates, POI, POI's students, current students in the program but not in your POI's lab, and current students at other places. You should look for a diverse set of opinions on the pros and cons of each program. Keep in mind that there is no one best choice, and you will need to make tradeoffs.

### Understanding your offer

From my knowledge, using financial support as a criterion, all offers can be classified into the following:

* Guaranteed financial support: The funding is guaranteed throughout the first five years through a combination of fellowship, research, teaching, or external awards.
  * Fellowship/external awards: Fellowships usually require faculty nominations during the admissions process. To get external awards like the NSF Graduate Research Fellowships Program, you will need to go through an application process. Note that most external awards I know of require you to be a U.S. citizen, national, or permanent resident.
  * Research/teaching assistantships: For RAs, you work with a professor on research and get paid (!!!). For TAs, you TA a class, and the time commitment varies between a few hours and 20 hours per week depending on the program/course/instructor. The monthly stipend falls somewhere between $2500 and $4000.
* No guaranteed financial support: Finding graduate assistantships to support your studies is possible, but you will need to rely on yourself to find them and there is no guarantee. In theory, if no professors are willing to take you as an RA and you weren't matched to a class that needs a TA, you will need to pay out of your pocket for the tuition and insurance fees. Yikes

### What do to on the visiting day

* **Chat with as many current students/POI as possible.** The POI could very much be your future (co-)advisor/collaborator, and you will also likely share an office with the current students who work in similar research areas, so it's good to know ahead of time who they are and what they are working on. You should also use this opportunity to see if the students are happy.
* **Connect with the other prospective students.** It's an excellent opportunity to build your network and meet new people even before the program starts.
* **Walk around.** I didn't attend my visiting days in person, but if I had, I would have spent a lot of free time walking around both the campus and the adjacent city. You will be spending the next few years in (somewhat of) a bubble around the campus, so a vibe match is crucial.
* Enjoy the free food and the free trip! You totally earned it.

### References

* [How should I choose between multiple Ph.D. programs I was admitted to?](https://academia.stackexchange.com/questions/66926/ive-been-admitted-to-multiple-phd-programs-how-should-i-choose-between-them) | StackExchange Academia Community Wiki
* [How to Choose Your Grad School](https://timdettmers.com/2022/03/13/how-to-choose-your-grad-school/) by Tim Dettmers
* [The Definitive ‘what do I ask/look for’ in a PhD Advisor Guide](https://www.cs.columbia.edu/wp-content/uploads/2019/03/Get-Advisor.pdf)
* [Some notes on picking grad schools/advisors](https://jxyzabc.blogspot.com/2009/02/some-notes-on-picking-grad.html) and [CS Grad School Part 5: School Visits](https://jxyzabc.blogspot.com/2008/08/cs-grad-school-part-5-school-visits.html) by Jean Yang
* [All About Graduate School Visits (for CS PhD programs)](https://koronkevi.ch/posts/grad-school-visits.html) by Paulette Koronkevich
* [How to pick a grad school for a PhD in Computer Science](https://vijayc.medium.com/how-to-pick-a-grad-school-for-a-phd-in-computer-science-a5ce7dceb246) by Prof. Vijay Chidambaram
* [地表最全奖学金攻略：我的二十二万刀经验分享](https://www.1point3acres.com/bbs/thread-763415-1-1.html) (Negotiating with Ph.D. programs for more fellowships: How I got a total of $220K from all offers)
* [如何选择博士导师](https://iphyer.github.io/blog/2017/09/10/choosingpi/) (How to choose a Ph.D. advisor)

## Appendix I: My application timeline

Don't feel obliged to copy my timeline exactly -- this is just for your reference.

* **Mid-August**: First draft of the school list and the SOP
* **Mid-September - Early-October**: Confirmation of recommendation letter from 3 professors
* **Early-October - Mid-November**: Send out cold emails to professors of interest. On average, I sent \~2.5 letters per week.
* **Early-November**: Finalized school list
* **Early-December**: Sent out all rec letter requests
* **Mid-December**: Finalized SoP; Sent out all applications!
* **Late-December - Mid-February**: Interviews
* **Late-January**: First unofficial offer
* **Mid-February**: First official offer -- the offers came in all the way to early March, although I withdrew/turned most of them down.
* **Late-March**: Committed to Princeton!

## Appendix II: Meta-references

* [Grad School Resources](https://martiansideofthemoon.github.io/2018/05/29/grad-resources.html) by Kalpesh Krishna
* [Applying to Ph.D. Programs in Computer Science](http://www.cs.cmu.edu/~harchol/gradschooltalk.pdf) by Prof. Mor Harchol-Balter
* [Matt Might's HOWTO: Apply for and get into grad school for science, engineering, math, and computer science](https://matt.might.net/articles/how-to-apply-and-get-in-to-graduate-school-in-science-mathematics-engineering-or-computer-science/)
* [Reflecting on CS Graduate Admissions](https://da-data.blogspot.com/2015/03/reflecting-on-cs-graduate-admissions.html) by Prof. David Anderson
* [Graduate Study Survival Guide](https://cs.uwaterloo.ca/~thachisu/survival.pdf) by Prof. Toshiya Hachisuka
* [EPFL EPIC Guide](https://epic-guide.github.io/applying) by students at EPFL
* [Chris Liu's list of resources for CS grad school application](https://chrisliu298.io/posts/grad-school-application.html)
* [USA Computer Science PhD Application Advice](https://www.youtube.com/watch?v=IprN9fPV2LI) by Prof. Shriram Krishnamurthi
* [My CS Ph.D.: Whether, why, and how to get a Ph.D. in CS](https://mycsphd.org/) by UCSD CSE
* [Advice on Research Communication Skills | Computer Science Department at Princeton University](https://www.cs.princeton.edu/grad/advice-on-research-communications-skills)
* [Reflections on my CS PhD Application Process](https://www.bodunhu.com/blog/posts/reflections-on-my-cs-phd-application-process/) | [Bodun Hu](https://www.bodunhu.com/)'s Blog
* [Machine Learning PhD Applications — Everything You Need to Know](https://timdettmers.com/2018/11/26/phd-applications/) by Tim Dettmers
* [CS Professors - Drafty](https://drafty.cs.brown.edu/csprofessors): Database of CS professors
* [computer science open rankings](https://drafty.cs.brown.edu/csopenrankings/): Choose and combine existing rankings to generate your preferred meta ranking for computer science programs in the United States and Canada, also by Drafty

For those of you who read Mandarin Chinese or don't bother to use Google Translate:

* [Top tier CS PhD招生官--我是如何审材料的](https://www.1point3acres.com/bbs/thread-585435-1-1.html) (How I review application packages as a student volunteer in the application committee of a top-tier CS Ph.D. program)
* [从审材料的角度谈谈研究生申请](https://www.1point3acres.com/bbs/thread-463109-1-1.html) (CS grad school application from an application reviewer's point of view)
* [也从审材料的角度讲讲如何准备cs phd申请！](https://www.1point3acres.com/bbs/thread-585851-1-1.html)(CS Ph.D. application from an application reviewer's point of view)
* [对于未来的CS PhD申请者来说，暑假应该做实习还是去实验室更好呢？ - Bihan Wen的回答 - 知乎](https://www.zhihu.com/question/308026281/answer/566961945) (Some insights on what professors look for from Ph.D. applicants by [Bihan Wen](https://personal.ntu.edu.sg/bihan.wen/))
* [关于PhD的思考](https://iphyer.github.io/blog/2017/06/21/phd/) (Some thoughs on doing a Ph.D.)
* [CS Ph.D. 2019 Fall 申请季总结 - James.Qiu的文章 - 知乎](https://zhuanlan.zhihu.com/p/60961921) (Review of my CS Ph.D. application in Fall 2019 by [Haoran Qiu](https://haoran-qiu.com/))
* [2019 Fall 你都申请了哪些学校的 MS/PhD，录取结果如何？ - 拎-yin的回答 - 知乎](https://www.zhihu.com/question/290670460/answer/586795783) (Review of my CS Ph.D. application in Fall 2019 by [Yin Lin](https://niceirene.github.io/))
* [我的PhD申请总结与经验分享](https://tylergu.com/summary.html) (Review of my CS Ph.D. application in Fall 2020 by [Jiawei Gu](https://tylergu.com/index.html))
* [CS Ph.D. 申请总结 (2021 Fall) - Romero的文章 - 知乎](https://zhuanlan.zhihu.com/p/362189295) (Review of my CS Ph.D. application in Fall 2021 by [Xiangfeng Zhu](https://xzhu27.me/))
* [2022 Fall你都申请了哪些学校的MA/MS/Ph.D.？- ruipeterpan的回答 - 知乎](https://www.zhihu.com/question/379814619/answer/2325160660) (My personal review of my CS Ph.D. application in Fall 2022)
* [美国CS PhD 申请经验总结 - Calpico的回答 - 知乎](https://zhuanlan.zhihu.com/p/533187747) (Takeaways from my CS Ph.D. application in Fall 2022 by [Yinwei Dai](http://yinwei-dai.com/))
* [申请季回望](https://zhiqiang.site/2022/07/01/retro.html) (Retrospect of my Ph.D. application season by [Zhiqiang Xie](https://zhiqiang.site/))


# Machine Learning Systems - Index

### Distributed Training & Parallelism Paradigms

* [\[OSDI '14\] Scaling Distributed Machine Learning with the Parameter Server](/machine-learning-systems/machine-learning-systems-index/scaling-distributed-machine-learning-with-the-parameter-server) ([pdf](https://web.eecs.umich.edu/~mosharaf/Readings/Parameter-Server.pdf))
* \[SoCC '18] Parameter Hub: a Rack-Scale Parameter Server for Distributed Deep Neural Network Training ([pdf](https://dl.acm.org/doi/pdf/10.1145/3267809.3267840))
* [\[OSDI '20\] BytePS: A High Performance and Generic Framework for Distributed DNN Training](/machine-learning-systems/machine-learning-systems-index/byteps-a-high-performance-and-generic-framework-for-distributed-dnn-training) ([pdf](https://www.usenix.org/system/files/osdi20-jiang.pdf))
* [\[VLDB '20\] PyTorch Distributed: Experiences on Accelerating Data Parallel Training](/machine-learning-systems/machine-learning-systems-index/pytorch-distributed-experiences-on-accelerating-data-parallel-training) ([pdf](https://dl.acm.org/doi/pdf/10.14778/3415478.3415530))
* \[MLSys '20] Resource Elasticity in Distributed Deep Learning ([pdf](https://proceedings.mlsys.org/paper/2020/file/006f52e9102a8d3be2fe5614f42ba989-Paper.pdf))
* \[NSDI '23] Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs ([pdf](https://arxiv.org/pdf/2204.12013.pdf))
* Parallelism Paradigms & Strategies ([Overview by Hugging Face](https://huggingface.co/docs/transformers/v4.16.2/en/parallelism))
  * [\[NIPS '19\] GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism](/machine-learning-systems/machine-learning-systems-index/gpipe-efficient-training-of-giant-neural-networks-using-pipeline-parallelism) ([pdf](https://papers.nips.cc/paper/2019/file/093f65e080a295f8076b1c5722a46aa2-Paper.pdf))
  * [\[SOSP '19\] PipeDream: Generalized Pipeline Parallelism for DNN Training](/machine-learning-systems/machine-learning-systems-index/pipedream-generalized-pipeline-parallelism-for-dnn-training) ([pdf](https://www.microsoft.com/en-us/research/uploads/prod/2019/08/fiddle_pipedream_sosp19.pdf))
  * [\[arXiv '19\] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism](/machine-learning-systems/machine-learning-systems-index/mlsys-papers-short-notes#2019-arxiv-megatron-lm-training-multi-billion-parameter-language-models-using-model-parallelism) ([pdf](https://arxiv.org/pdf/1909.08053.pdf))
  * \[MLSys '19] FlexFlow: Beyond Data and Model Parallelism for Deep Neural Networks ([pdf](https://proceedings.mlsys.org/paper/2019/file/c74d97b01eae257e44aa9d5bade97baf-Paper.pdf))
  * [\[SC '20\] ZeRO: memory optimizations toward training trillion parameter models](/machine-learning-systems/machine-learning-systems-index/2019-sc-zero-memory-optimizations-toward-training-trillion-parameter-models) ([pdf](https://arxiv.org/pdf/1910.02054.pdf))
  * \[ATC '20] HetPipe: Enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data Parallelism ([pdf](https://www.usenix.org/system/files/atc20-park.pdf))
  * [\[SC '21\] ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learning](/machine-learning-systems/machine-learning-systems-index/2019-sc-zero-memory-optimizations-toward-training-trillion-parameter-models#zero-infinity-and-zero-offload) ([pdf](https://dl.acm.org/doi/pdf/10.1145/3458817.3476205))
  * \[SC '21] Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM ([pdf](https://dl.acm.org/doi/pdf/10.1145/3458817.3476209))
  * [\[SC '21\] Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines](/machine-learning-systems/machine-learning-systems-index/mlsys-papers-short-notes#2021-sc-chimera-efficiently-training-large-scale-neural-networks-with-bidirectional-pipelines) ([pdf](https://dl.acm.org/doi/pdf/10.1145/3458817.3476145))
  * \[ICML '21] Memory-Efficient Pipeline-Parallel DNN Training ([pdf](http://proceedings.mlr.press/v139/narayanan21a/narayanan21a.pdf))
  * [\[ATC '21\] ZeRO-Offload: Democratizing Billion-Scale Model Training](/machine-learning-systems/machine-learning-systems-index/2019-sc-zero-memory-optimizations-toward-training-trillion-parameter-models#zero-infinity-and-zero-offload) ([pdf](https://www.usenix.org/system/files/atc21-ren-jie.pdf))
  * \[PPoPP '21] DAPPLE: A Pipelined Data Parallel Approach for Training Large Models ([pdf](https://dl.acm.org/doi/pdf/10.1145/3437801.3441593))
  * \[OSDI '22] Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning ([pdf](https://www.usenix.org/system/files/osdi22-zheng-lianmin.pdf))
  * \[OSDI '22] Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and Parallelization ([pdf](https://www.usenix.org/system/files/osdi22-unger.pdf))
  * \[EuroSys '22] Varuna: Scalable, Low-cost Training of Massive Deep Learning Models ([pdf](https://dl.acm.org/doi/pdf/10.1145/3492321.3519584))
  * \[arXiv '22] Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model ([pdf](https://arxiv.org/pdf/2201.11990.pdf))
  * \[PPoPP '22] BaGuaLu: Targeting Brain Scale Pretrained Models with over 37 Million Cores ([pdf](https://keg.cs.tsinghua.edu.cn/jietang/publications/PPOPP22-Ma%20et%20al.-BaGuaLu%20Targeting%20Brain%20Scale%20Pretrained%20Models%20w.pdf))
  * \[NeurIPS '22] AMP:Automatically Finding Model Parallel Strategies with Heterogeneity Awareness ([pdf](https://arxiv.org/pdf/2210.07297.pdf))
  * \[VLDB '23] MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud ([pdf](https://arxiv.org/pdf/2205.00119.pdf))

### Workload Scheduling, Cluster Resource Management

* [\[NSDI '11\] DRF: Fair Allocation of Multiple Resource Types](/machine-learning-systems/machine-learning-systems-index/dominant-resource-fairness-fair-allocation-of-multiple-resource-types) ([pdf](https://www.usenix.org/legacy/events/nsdi11/tech/full_papers/Ghodsi.pdf))
* [\[OSDI '18\] Gandiva: Introspective Cluster Scheduling for Deep Learning](/machine-learning-systems/machine-learning-systems-index/gandiva-introspective-cluster-scheduling-for-deep-learning) ([pdf](https://www.usenix.org/system/files/osdi18-xiao.pdf))
* \[EuroSys '18] Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters ([pdf](https://dl.acm.org/doi/pdf/10.1145/3190508.3190517))
* [\[ATC '19\] Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads](/machine-learning-systems/machine-learning-systems-index/analysis-of-large-scale-multi-tenant-gpu-clusters-for-dnn-training-workloads) ([pdf](https://www.usenix.org/system/files/atc19-jeon.pdf))
* [\[NSDI '19\] Tiresias: A GPU Cluster Manager for Distributed Deep Learning](/machine-learning-systems/machine-learning-systems-index/tiresias-a-gpu-cluster-manager-for-distributed-deep-learning) ([pdf](https://www.usenix.org/system/files/nsdi19-gu.pdf))
* [\[NSDI '20\] Themis: Fair and Efficient GPU Cluster Scheduling](/machine-learning-systems/machine-learning-systems-index/themis-fair-and-efficient-gpu-cluster-scheduling) ([pdf](https://www.usenix.org/system/files/nsdi20-paper-mahajan.pdf))
* [\[MLSys '20\] Salus: Fine-Grained GPU Sharing Primitives for Deep Learning Applications](/machine-learning-systems/machine-learning-systems-index/2020-sigcomm-reducto-on-camera-filtering-for-resource-efficient-real-time-video-analytics/salus-fine-grained-gpu-sharing-primitives-for-deep-learning-applications) ([pdf](https://www.mosharaf.com/wp-content/uploads/salus-mlsys20.pdf))
* [\[OSDI '20\] Gavel: Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads](/machine-learning-systems/machine-learning-systems-index/gavel-heterogeneity-aware-cluster-scheduling-policies-for-deep-learning-workloads) ([pdf](https://www.usenix.org/system/files/osdi20-narayanan_deepak.pdf))
* [\[OSDI '20\] AntMan: Dynamic Scaling on GPU Clusters for Deep Learning](/machine-learning-systems/machine-learning-systems-index/2020-osdi-antman-dynamic-scaling-on-gpu-clusters-for-deep-learning) ([pdf](https://www.usenix.org/system/files/osdi20-xiao.pdf))
* \[OSDI '20] HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees ([pdf](https://www.usenix.org/system/files/osdi20-zhao_hanyu.pdf))
* \[EuroSys '20] Gandiva-Fair: Balancing efficiency and fairness in heterogeneous GPU clusters for deep learning ([pdf](https://dl.acm.org/doi/pdf/10.1145/3342195.3387555))
* \[EuroSys '20] AlloX: Compute Allocation in Hybrid Clusters ([pdf](https://www.mosharaf.com/wp-content/uploads/allox-eurosys20.pdf))
* [\[MLSys '21\] Wavelet: Efficient DNN Training with Tick-Tock Scheduling](/machine-learning-systems/machine-learning-systems-index/wavelet-efficient-dnn-training-with-tick-tock-scheduling) ([pdf](https://proceedings.mlsys.org/paper/2021/file/c81e728d9d4c2f636f067f89cc14862c-Paper.pdf))
* [\[OSDI '21\] Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning](/machine-learning-systems/machine-learning-systems-index/pollux-co-adaptive-cluster-scheduling-for-goodput-optimized-deep-learning) ([pdf](https://www.usenix.org/system/files/osdi21-qiao.pdf))
* \[ATC '21] Zico: Efficient GPU Memory Sharing for Concurrent DNN Training ([pdf](https://www.usenix.org/system/files/atc21-lim.pdf))
* \[SoCC '21] Chronus: A Novel Deadline-aware Scheduler for Deep Learning Training Jobs ([pdf](https://dl.acm.org/doi/pdf/10.1145/3472883.3486978))
* \[NSDI '21] AFS/CoDDL: Elastic Resource Sharing for Distributed Deep Learning ([pdf](https://www.usenix.org/system/files/nsdi21-hwang.pdf))
* \[NSDI '22] MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters ([pdf](https://www.usenix.org/system/files/nsdi22-paper-weng.pdf))
* [\[OSDI '22\] Synergy: Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters](/machine-learning-systems/machine-learning-systems-index/mlsys-papers-short-notes#2022-osdi-looking-beyond-gpus-for-dnn-scheduling-on-multi-tenant-clusters) ([pdf](https://www.usenix.org/system/files/osdi22-mohan.pdf))
* \[SIGCOMM '22] Multi-Resource Interleaving for Deep Learning Training ([pdf](https://dl.acm.org/doi/pdf/10.1145/3544216.3544224))
* \[arXiv '22] Deep Learning Workload Scheduling in GPU Datacenters: Taxonomy, Challenges and Vision ([pdf](https://arxiv.org/pdf/2205.11913.pdf))
* \[NSDI '23] Shockwave: Proactive, Fair, and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning
* \[NSDI '23] ModelKeeper: Accelerating DNN Training via Automated Training Warmup

### Serving/Inference

* \[NSDI '17] Clipper: A Low-Latency Online Prediction Serving System ([pdf](https://www.usenix.org/system/files/conference/nsdi17/nsdi17-crankshaw.pdf))
* \[NIPS '17 MLSys workshop] TensorFlow-Serving: Flexible, High-Performance ML Serving ([pdf](http://learningsys.org/nips17/assets/papers/paper_1.pdf))
* \[arXiv '18] Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications ([pdf](https://arxiv.org/pdf/1811.09886.pdf))
* [\[NIPS '18\] Dynamic Space-Time Scheduling for GPU Inference](/machine-learning-systems/machine-learning-systems-index/2018-nips-dynamic-space-time-scheduling-for-gpu-inference) ([pdf](http://learningsys.org/nips18/assets/papers/102CameraReadySubmissionGPU_Virtualization%20\(8\).pdf))
* [\[SOSP '19\] Parity Models: Erasure-Coded Resilience for Prediction Serving Systems](/machine-learning-systems/machine-learning-systems-index/2019-sosp-parity-models-erasure-coded-resilience-for-prediction-serving-systems) ([pdf](https://www.cs.cmu.edu/~rvinayak/papers/sosp2019parity-models.pdf))
* \[SOSP '19] Nexus: A GPU Cluster Engine for Accelerating DNN-Based Video Analysis ([pdf](https://dl.acm.org/doi/pdf/10.1145/3341301.3359658))
* \[arXiv '19] No DNN left behind: Improving inference in the cloud with Multi-Tenancy ([pdf](https://arxiv.org/pdf/1901.06887.pdf))
* \[ATC '19] MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving ([pdf](https://www.usenix.org/system/files/atc19-zhang-chengliang.pdf))
* \[SoCC '20] GSLICE: controlled spatial sharing of GPUs for a scalable inference platform ([pdf](https://dl.acm.org/doi/pdf/10.1145/3419111.3421284))
* \[SoCC '20] InferLine: Latency-Aware Provisioning and Scaling for Prediction Serving Pipelines ([pdf](https://dl.acm.org/doi/pdf/10.1145/3419111.3421285))
* \[OSDI '20] Serving DNNs like Clockwork: Performance Predictability from the Bottom Up ([pdf](https://www.usenix.org/system/files/osdi20-gujarati.pdf))
* \[OSDI '20] PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications ([pdf](https://www.usenix.org/system/files/osdi20-bai.pdf))
* \[ATC '21] INFaaS: Automated Model-less Inference Serving ([pdf](https://www.usenix.org/system/files/atc21-romero.pdf))
* [\[EuroMLSys '21\] Interference-Aware Scheduling for Inference Serving](/machine-learning-systems/machine-learning-systems-index/2021-euromlsys-interference-aware-scheduling-for-inference-serving) ([pdf](https://dl.acm.org/doi/pdf/10.1145/3437984.3458837))
* \[arXiv '21] Serving DNN Models with Multi-Instance GPUs: A Case of the Reconfigurable Machine Scheduling Problem ([pdf](https://arxiv.org/pdf/2109.11067.pdf))
* \[arXiv '21] Gati: Accelerating Deep Learning Inference via Learned Caches ([pdf](https://arxiv.org/pdf/2101.07344.pdf))
* [\[ICML '21\] Boosting the Throughput and Accelerator Utilization of Specialized CNN Inference Beyond Increasing Batch Size](/machine-learning-systems/machine-learning-systems-index/mlsys-papers-short-notes#2021-icml-boosting-the-throughput-and-accelerator-utilization-of-specialized-cnn-inference-beyond-in) ([pdf](http://proceedings.mlr.press/v139/kosaian21a/kosaian21a.pdf))
* \[ICML '22] DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale ([pdf](https://proceedings.mlr.press/v162/rajbhandari22a/rajbhandari22a.pdf))
* \[OSDI '22] Achieving μs-scale Preemption for Concurrent GPU-accelerated DNN Inferences ([pdf](https://www.usenix.org/system/files/osdi22-han.pdf))
* \[OSDI '22] Orca: A Distributed Serving System for Transformer-Based Generative Models ([pdf](https://www.usenix.org/system/files/osdi22-yu.pdf))
* \[ATC '22] Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing ([pdf](https://www.usenix.org/system/files/atc22-choi-seungbeom.pdf))
* \[SIGMOD '22] Serverless Data Science - Are We There Yet? A Case Study of Model Serving ([pdf](https://dl.acm.org/doi/pdf/10.1145/3514221.3517905))

### Optimizing Networks/Communications for ML

* \[ATC '17] Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters ([pdf](https://www.usenix.org/system/files/conference/atc17/atc17-zhang.pdf))
* [\[MLSys '19\] BlueConnect: Decomposing All-Reduce for Deep Learning on Heterogeneous Network Hierarchy](/machine-learning-systems/machine-learning-systems-index/mlsys-papers-short-notes#2019-mlsys-blueconnect-decomposing-all-reduce-for-deep-learning-on-heterogeneous-network-hierarchy) ([pdf](https://mlsys.org/Conferences/2019/doc/2019/130.pdf))
* [\[MLSys '19\] TicTac: Accelerating Distributed Deep Learning with Communication Scheduling](/machine-learning-systems/machine-learning-systems-index/2019-sosp-bytescheduler-a-generic-communication-scheduler-for-distributed-dnn-training-...#comparisons-with-p3-and-tictac) ([pdf](https://mlsys.org/Conferences/2019/doc/2019/199.pdf))
* [\[MLSys '19\] P3: Priority-Based Parameter Propagation for Distributed DNN Training](/machine-learning-systems/machine-learning-systems-index/2019-sosp-bytescheduler-a-generic-communication-scheduler-for-distributed-dnn-training-...#comparisons-with-p3-and-tictac) ([pdf](https://proceedings.mlsys.org/paper/2019/file/d09bf41544a3365a46c9077ebb5e35c3-Supplemental.pdf))
* [\[SOSP '19\] ByteScheduler: A Generic Communication Scheduler for Distributed DNN Training Acceleration](/machine-learning-systems/machine-learning-systems-index/2019-sosp-bytescheduler-a-generic-communication-scheduler-for-distributed-dnn-training-...) ([pdf](https://dl.acm.org/doi/pdf/10.1145/3341301.3359642))
* [\[NetAI '20\] Is Network the Bottleneck of Distributed Training?](/machine-learning-systems/machine-learning-systems-index/2020-netai-is-network-the-bottleneck-of-distributed-training) ([pdf](https://dl.acm.org/doi/pdf/10.1145/3405671.3405810))
* [\[MLSys '20\] Blink: Fast and Generic Collectives for Distributed ML](/machine-learning-systems/machine-learning-systems-index/mlsys-papers-short-notes#2020-mlsys-blink-fast-and-generic-collectives-for-distributed-ml) ([pdf](https://proceedings.mlsys.org/paper/2020/file/43ec517d68b6edd3015b3edc9a11367b-Paper.pdf))
* \[MLSys '20] PLink: Discovering and Exploiting Datacenter Network Locality for Efficient Cloud-based Distributed Training ([pdf](https://proceedings.mlsys.org/paper/2020/file/182be0c5cdcd5072bb1864cdee4d3d6e-Paper.pdf))
* \[SoCC '20] Network-accelerated Distributed Machine Learning for Multi-Tenant Settings ([pdf](https://dl.acm.org/doi/pdf/10.1145/3419111.3421296))
* [\[NSDI '21\] SwitchML: Scaling Distributed Machine Learning with In-Network Aggregation](/machine-learning-systems/machine-learning-systems-index/2021-nsdi-switchml-scaling-distributed-machine-learning-with-in-network-aggregation) ([pdf](https://www.usenix.org/system/files/nsdi21-sapio.pdf))
* \[NSDI '21] ATP: In-network Aggregation for Multi-tenant Learning ([pdf](https://www.usenix.org/system/files/nsdi21-lao.pdf))
* \[SIGCOMM '21] Efficient Sparse Collective Communication and its application to Accelerate Distributed Deep Learning ([pdf](https://dl.acm.org/doi/pdf/10.1145/3452296.3472904))
* \[MLSys '21] In-network Aggregation for Shared Machine Learning Clusters
* [\[NSDI '23\] Synthesizing Collective Communication Algorithms for Heterogeneous Networks with TACCL](/machine-learning-systems/machine-learning-systems-index/mlsys-papers-short-notes#2021-arxiv-synthesizing-collective-communication-algorithms-for-heterogeneous-networks-with-taccl) ([pdf](http://arxiv-export-lb.library.cornell.edu/abs/2111.04867v2))
* \[arXiv '21] Cloud Collectives: Towards Cloud-aware Collectives for ML Workloads with Rank Reordering ([pdf](https://arxiv.org/pdf/2105.14088.pdf))
* \[PPoPP '21] Synthesizing Optimal Collective Algorithms ([pdf](https://dl.acm.org/doi/pdf/10.1145/3437801.3441620))
* \[NSDI '22] Accelerating Collective Communication in Data Parallel Training across Deep Learning Frameworks ([pdf](https://www.usenix.org/system/files/nsdi22-paper-romero.pdf))
* \[NSDI '23] Better Together: Jointly Optimizing ML Collective Scheduling and Execution Planning using SYNDICATE
* Optical Networks for ML
  * \[SIGCOMM '21] SiP-ML: High-Bandwidth Optical Network Interconnects for Machine Learning Training ([pdf](https://people.csail.mit.edu/ghobadi/papers/sipml_sigcomm_2021.pdf))
  * \[SIGCOMM '21 OptSys workshop] IOI: In-network Optical Inference ([pdf](https://people.csail.mit.edu/zhizhenzhong/papers/2021_OptSys_IOI.pdf))
  * \[OFC '22] Emerging Optical Interconnects for AI Systems ([pdf](https://people.csail.mit.edu/ghobadi/papers/optics_for_ai_ofc_2022.pdf))
  * \[NSDI '23] TOPOOPT: Optimizing the Network Topology for Distributed DNN Training ([pdf](https://arxiv.org/pdf/2202.00433.pdf))

### ML for Systems, Video Analytics & Streaming

* [Kuntai Du's overview on video analytics](https://kuntai.notion.site/Video-analytics-literature-review-90947b73637f427da7d8adc82e764c77)
* [CS34702 @ UChi: Machine Learning for Networking and Systems](https://people.cs.uchicago.edu/~junchenj/34702-fall21/)
* \[SIGCOMM '17] Pensieve: Neural Adaptive Video Streaming with Pensieve
* \[HotNets '17] Congestion-Control Throwdown
* [\[SIGCOMM '18\] Chameleon: Scalable Adaptation of Video Analytics via Temporal and Cross-camera Correlations](/machine-learning-systems/machine-learning-systems-index/2018-sigcomm-chameleon-scalable-adaptation-of-video-analytics-via-temporal-and-cross-camera-...)
* \[NSDI '18] PCC Vivace: Online-Learning Congestion Control
* \[NSDI '18] Salsify: Low-Latency Network Video through Tighter Integration between a Video Codec and a Transport Protocol
* \[HotEdge '19] Edge-based Transcoding for Adaptive Live Video Streaming ([pdf](http://web.cs.ucla.edu/~dogga/publications/hotedge19.pdf))
* [\[SIGCOMM '20\] Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video Analytics](/machine-learning-systems/machine-learning-systems-index/2020-sigcomm-reducto-on-camera-filtering-for-resource-efficient-real-time-video-analytics)
* \[SIGCOMM '20] DDS: Server-Driven Video Streaming for Deep Learning Inference
* \[MobiCom '20] OnRL: Improving Mobile Video Telephony via Online Reinforcement Learning
* \[NSDI '20] Learning in situ: a randomized experiment in video streaming
* \[OSDI '21] Polyjuice: High-Performance Transactions via Learned Concurrency Control ([pdf](https://www.usenix.org/system/files/osdi21-wang-jiachen.pdf))
* \[NSDI '22] Ekya: Continuous Learning of Video Analytics Models on Edge Compute Servers
* \[HotMobile '22] Understanding the Potential of Server-Driven Edge Video Analytics
* \[SIGCOMM '22] Genet: automatic curriculum generation for learning adaptation in networking
* \[NSDI '23] GEMEL: Model Merging for Memory-Efficient, Real-Time Video Analytics at the Edge

### Tricks and Relaxations in Learning and Systems: Compression, Pruning, Freezing, and many more

* \[NIPS '13] More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server
* \[arXiv '16] Training Deep Nets with Sublinear Memory Cost
* \[ICLR '16] Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding
* \[NIPS '17] Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent
* \[ICLR '18] Mixed precision training
* \[ICLR '19] The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
* \[arXiv '21] AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning
* \[PVLDB '21] BAGUA: Scaling up Distributed Learning with System Relaxations
* \[arXiv '22] BagPipe: Accelerating Deep Recommendation Model Training
* \[arXiv '22] Efficient DNN Training with Knowledge-Guided Layer Freezing
* Hongyi Wang's talk: [On the Utility of Gradient Compression in Distributed Training Systems](https://www.youtube.com/watch?v=gprhrinr3I4)
  * \[NIPS '18] ATOMO: Communication-efficient Learning via Atomic Sparsification ([pdf](https://proceedings.neurips.cc/paper/2018/file/33b3214d792caf311e1f00fd22b392c5-Paper.pdf))
  * [\[MLSys '21\] Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification](/machine-learning-systems/machine-learning-systems-index/accordion-adaptive-gradient-communication-via-critical-learning-regime-identification) ([pdf](https://proceedings.mlsys.org/paper/2021/file/1d7f7abc18fcb43975065399b0d1e48e-Paper.pdf))
  * \[MLSys '21] Pufferfish: Communication-efficient Models At No Extra Cost ([pdf](https://arxiv.org/pdf/2103.03936.pdf))
  * \[SOSP '21] Gradient Compression Supercharged High-Performance Data Parallel DNN Training ([pdf](https://dl.acm.org/doi/pdf/10.1145/3477132.3483553))
  * \[MLSys '22] On the utility of gradient compression in distributed training systems ([pdf](https://proceedings.mlsys.org/paper/2022/file/cedebb6e872f539bef8c3f919874e9d7-Paper.pdf))
  * \[arXiv '22] Cuttlefish: Factorized Model Training without All the Tuning
  * \[arXiv '22] ByteComp: Revisiting Gradient Compression in Distributed Training ([pdf](https://arxiv.org/pdf/2205.14465.pdf))

### Misc: Storage, Hyperparameter Tuning, Federated Learning, DL Compilers, Green Datacenters

* \[NIPS '16 workshop] Federated Learning: Strategies for Improving Communication Efficiency
* \[ICML '18 workshop] Tune: A research platform for distributed model selection and training
* \[OSDI '18] TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
* \[MLSys '19] Bandana: Using Non-Volatile Memory for Storing Deep Learning Models
* \[MLSys '19] Towards Federated Learning at Scale: System Design
* \[SOSP '19] TASO: Optimizing Deep Learning Computation with Automatic Generation of Graph Substitutions
* \[MLSys '20] A System for Massively Parallel Hyperparameter Tuning
* \[ICLR '20] Federated Learning with Matched Averaging
* \[OSDI '20] Ansor: Generating High-Performance Tensor Programs for Deep Learning
* \[OSDI '20] Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks
* \[EuroSys '21] RubberBand: Cloud-based Hyperparameter Tuning
* [\[FAST '21\] CheckFreq: Frequent, Fine-Grained DNN Checkpointing](/machine-learning-systems/machine-learning-systems-index/checkfreq-frequent-fine-grained-dnn-checkpointing)
* [\[VLDB '21\] Analyzing and Mitigating Data Stalls in DNN Training](/machine-learning-systems/machine-learning-systems-index/analyzing-and-mitigating-data-stalls-in-dnn-training)
* \[MLSys '21] Fluid: Resource-aware Hyperparameter Tuning Engine
* \[OSDI '21] Oort: Efficient Federated Learning via Guided Participant Selection ([pdf](https://www.usenix.org/system/files/osdi21-lai.pdf))
* \[OSDI '21] PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated Corrections
* \[SoCC '21] Elastic Hyperparameter Tuning on the Cloud ([pdf](https://dl.acm.org/doi/pdf/10.1145/3472883.3486989))
* \[NSDI '22] Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models
* \[ICML '22] FedScale: Benchmarking Model and System Performance of Federated Learning at Scale
* \[HotCarbon '22] Treehouse: A Case For Carbon-Aware Datacenter Software
* \[NSDI '23] Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training


# MLSys Papers - Short Notes

## \[2019 arXiv] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

This work proposes tensor parallelism (TP), where tensors are partitioned across devices and are only aggregated for operations that require the whole tensor. A key insight of TP is that matrix multiplication can be split between multiple GPUs to parallelize computation and save memory.

&#x20;![](/files/IQjx8YwxgkFPDy3DC8QT)

Each transformer layer consists of a self-attention block followed by a two-layer, multi-layer perceptron (MLP). To parallelize an MLP, column parallelism can be used to split the matrix multiplication, and synchronizations are not needed until the very end of the computation. Parallelizing the multi-headed attention layers is even easier since they are already inherently parallel. As a result, each transformer layer requires two allreduce during the forward pass and two allreduce during the backward pass.

![](/files/pqPjXDLkfKmGWS3pwJ9F)

Note that using TP requires a super fast network for near-theoretical-optimal performance, and in real life, TP is usually used in conjugation with other forms of parallelism.

## \[2019 MLSys] BlueConnect: Decomposing All-Reduce for Deep Learning on Heterogeneous Network Hierarchy

BlueConnect adapts to the hierarchy of communication bandwidths by leveraging topology-awareness to fully utilize the heterogeneous network architecture. It decomposes all-reduce (reduce-scatter + all-gather) into multiple stages of parallelizable reduce-scatter & all-gather, which provides more granularity and flexibility to map operations to the heterogeneous underlying network hierarchy.

![](/files/w99uTjXfkzjN7oaomk0Z)

## \[2020 MLSys] Blink: Fast and Generic Collectives for Distributed ML

This paper address the problem of link under-utilization due to **topology heterogeneity** in distributed ML training. Topology heterogeneity mainly comes from (1) differing server configurations (e.g, different NVLink topologies across generations of DGX nodes) and (2) scheduler’s topology-agnostic placements/allocations (e.g., an 8-GPU job uses 3 GPUs in an 8-GPU DGX node and 5 GPUs from another). To handle topology heterogeneity from hardware generations or partial allocations from cluster schedulers, Blink **dynamically generates optimal communication primitives for a given topology**. Blink models collective communication operations as flows on a directed graph and uses a spanning-tree packing algorithm to maximize link bandwidth utilization.

![](/files/thB2ijxIkXHEyCO3ItMi)

![](/files/kQLzOY7SNYqaOB75ZfWH)

## \[2021 ICML] Boosting the Throughput and Accelerator Utilization of Specialized CNN Inference Beyond Increasing Batch Size

Serving specialized CNNs (e.g., for offline video analytics) have low arithmetic intensity, leading to the severe under-utilization of server-grade accelerators. Increasing the batch size is a popular technique to boost the arithmetic intensity, utilization, and application-level throughput by amortizing the cost of loading a CNN’s weights from memory. However, it suffers from diminishing returns. This paper proposes a technique to **redesign specialized CNNs** with the purpose of **boosting the inference utilization and throughput**. The key insight is that, once arithmetic intensity has plateaued due to increased batch size, reading/writing activations accounts for most of the memory traffic in specialized CNNs. The authors show that this memory traffic can be significantly reduced, while performing the same number of FLOPs, by jointly decreasing the size of the batch of input/output activations for a layer and increasing the layer’s width.

![](/files/ccFvkRZD03ojNcNYUfoz)

Compared to vanilla CNNs, FoldedCNNs have improvements on the throughput and the accelerator utilization while suffering slight accuracy loss.

## \[2021 arXiv] Synthesizing Collective Communication Algorithms for Heterogeneous Networks with TACCL

TACCL encodes a profiled topology and input size into a synthesis problem to generate optimized communication algorithms.

NCCL uses the topology of GPU connections and NIC placement along with buffer size to decide between two main types of communication algorithms — Ring and Tree, but it is agnostic to the exact performance profile of the links, and thus is often multiple times slower than TACCL’s custom collectives.

![](/files/5DhNcb4X8Oi7u1X4JHbH)

## \[2021 SC] Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines

Chimera is yet another pipeline parallelism paradigm. Compared with the other STOA systems, it reduces more compute idleness and has a more balanced activation memory consumption.

![](/files/dIMBbQR5dLer6C2FoSRk)

![](/files/dVSypCg4VPkj9cujqXj8)

## \[2022 OSDI] Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters

Nowadays, DNN workload schedulers in shared GPU clusters consider GPU as the dominant resource and only allocate other types of resources (e.g., CPU and memory) proportional to the number of GPUs. However, different jobs have various sensitivity to these other types of resources, which leads to sub-optimal allocation results by current schedulers. &#x20;

![](/files/MZHx9lyuDoO5XnpJ3O8W)

Synergy is an idea that applies to all existing scheduling policies: It uses profiling to infer a workload's sensitivity to different resources and performs multi-resource workload-aware resource allocation. The key nugget is to co-locate two jobs on the same server, one of which is CPU-sensitive and the other is not, so that while the CPU-insensitive job does not hurt from the reduced resource allocation, the CPU-sensitive job can gain a higher throughput, benefiting the cluster-wide aggregate throughput and metrics like avg JCT, makespan, fairness, etc.

![](/files/DxvmMvtYH7nu8h5fsEH6)

The main technical contributions of this paper are two-fold:&#x20;

* Profiling the workloads: Naively profiling all possible resource configurations can be expensive due to the large combination space. Synergy introduces an optimistic profiling technique that exploits the predictability in the relationship between job throughput and memory allocation. As for the CPU allocation, Synergy empirically profiles the job for varying, discrete CPU allocations at full memory allocation. The profiling time is tens of minutes, which is reasonable considering most DNN jobs are long-running.
* Encorporating resource-sensitivity-awareness into existing scheduling algorithms.


# \[2011 NSDI] Dominant Resource Fairness: Fair Allocation of Multiple Resource Types

## One-line Summary

DRF is a generalization of max min fairness for multiple resources. Each user has a demand vector (containing multiple resources; for one task of the user) and a dominant resource (xxx-bound) -- calculated using the demand vector and total resources available -- and DRF tries to equalize the dominant share (fraction of the dominant resource user is allocated) for all users.

## Paper Structure Outline

1. Introduction
2. Motivation
3. Allocation Properties
4. Dominant Resource Fairness (DRF)
   1. An Example
   2. DRF Scheduling Algorithm
   3. Weighted DRF
5. Alternative Fair Allocation Policies
   1. Asset Fairness
   2. Competitive Equilibrium from Equal Incomes
   3. Comparison with DRF
6. Analysis
   1. Fairness Properties
      1. Properties Violated by Asset Fairness
      2. Properties Violated by CEEI
      3. Resource Monotonicity vs. Sharing Incentives and Pareto efficiency
   2. Discrete Resource Allocation
7. Experimental Results
   1. Dynamic Resource Sharing
   2. DRF vs. Alternative Allocation Policies
   3. Simulations using Facebook Traces
8. Related Work
9. Conclusion and Future Work
10. Acknowledgments

## Background & Motivation

### Background

* Desired properties for resource allocation policies
  * Sharing incentive: Each user should be better off sharing the cluster, than exclusively using her own partition of the cluster. Consider a cluster with identical nodes and n users. Then a user should not be able to allocate more tasks in a cluster partition consisting of 1/n of all resources.
  * Strategy-proofness: Users should not be able to benefit by lying about their resource demands. This provides incentive compatibility, as a user cannot improve her allocation by lying.
    * The paper included two really interesting examples of users benefitting from lying
  * Pareto efficiency: It should not be possible to increase the allocation of a user without decreasing the allocation of at least another user. This property is important as it leads to maximizing system utilization subject to satisfying the other properties.
  * Envy freeness: A user should not prefer the allocation of another user.
  * Four other nice properties that are nice-to-have
    * Single resource fairness: For a single resource, the solution should reduce to max-min fairness.
    * Bottleneck fairness: If there is one resource that is percent-wise demanded most of by every user, then the solution should reduce to max-min fairness for that resource.
    * Population monotonicity: When a user leaves the system and relinquishes her resources, none of the allocations of the remaining users should decrease.
    * Resource monotonicity: If more resources are added to the system, none of the allocations of the existing users should decrease.
* Existing (fair) allocation policies
  * Equal share: Not work conserving
  * Max min fairness: Maximize the allocation for most poorly-treated users
  * Asset fairness: Equalizes each user's sum of resource shares (violates sharing incentive)
  * (Microeconomic theory) Competitive Equilibrium from Equal Incomes (CEEI): (not strategy-proof)

### Motivation

* Previous work on fair allocation focuses on one single resource type

## Design and Implementation

![Pseudo-code for DRF](/files/-MkpJrgc2CaxtEjACwq4)

## Evaluation

## Links & References

* [Paper PDF](https://cs.stanford.edu/~matei/papers/2011/nsdi_drf.pdf)


# \[2014 OSDI] Scaling Distributed Machine Learning with the Parameter Server

## One-line Summary

This paper presents the design, implementation, and evaluation of an implementation of the parameter server framework for distributed machine learning problems.

## Paper Structure Outline

1. Introduction
   1. Contributions
   2. Engineering Challenges
   3. Related Work&#x20;
2. Machine Learning
   1. Goals
   2. Risk Minimization
   3. Generative Models
3. Architecture
   1. (Key, Value) Vectors
   2. Range Push and Pull
   3. User-Defined Functions on the Server
   4. Asynchronous Tasks and Dependency
   5. Flexible Consistency
   6. User-defined Filters
4. Implementation
   1. Vector Clock
   2. Messages
   3. Consistent Hashing
   4. Replication and Consistency
   5. Server Management
   6. Worker Management
5. Evaluation
   1. Sparse Logistic Regression
   2. Latent Dirichlet Allocation
   3. Sketches
6. Summary and Discussion

## Background & Motivation

ML jobs and model sizes are getting bigger, thus we distributed the data/model across multiple worker machines. The parameter server model is a framework for distributed machine learning problems.

This paper presents a third-generation parameter server model which has five key features:

1. **Efficient communication**: The asynchronous communication model does not block computation
2. **Flexible consistency models**: Relaxed consistency further hides synchronization cost and latency. The algorithm designers are allowed to balance the algorithmic convergence rate and system efficiency
3. **Elastic Scalability**: New nodes can be added w/o restarting the running framework
4. **Fault Tolerance and Durability**: Recover from non-catastrophic failures w/o interrupting computation
5. **Ease of Use**: The globally shared parameters are represented as (potentially sparse) vectors and matrices to facilitate the development of machine learning applications. The linear algebra data types come with high-performance multi-threaded libraries.

## Design

![](/files/-MQ0RSUciRjT-dYGnPFW)

A server node in the server group maintains a partition of the globally shared parameters. The server manager node maintains a consistent view of the metadata (liveness, assignment of partitions) of the servers. Server nodes communicate with each other to replicate and/or to migrate parameters for reliability and scaling. Worker groups communicate with the server groups to pull the latest parameters, then compute the gradients locally and push them back.

The model shared among nodes can be represented as a set of (key, value) pairs.

![](/files/-MQ0Zpc3Ix_FfNhUCo1B)

An issue with having independent tasks (is this the same as async training?) is that inconsistency may arise. For example, in this case, iteration 11 is started before the parameters are pulled back, so it uses the old params from iter 10 and thus obtains the same gradients as iter 10. This is namely a tradeoff between system efficiency and algorithm convergence rate, and the best tradeoff depends on a variety of factors including the algorithm’s sensitivity to data inconsistency, feature correlation in training data, and capacity difference of hardware components. PS gives the algorithm designer the flexibility in defining consistency models. There are three main consistency models:

1. **Sequential**: All tasks are executed sequentially. The next task can only start when the previous one has finished.
2. **Eventual**: All tasks may start simultaneously. This is only recommendable if the underlying algorithms are robust to delays.
3. **Bounded Delay**: A knob, τ, the maximal delay time, shifts bounded delay between the previous two policies (τ=0 is sequential consistency model, τ=∞ is the eventual consistency model). When a maximal delay time τ is set, a new task will be blocked until all previous tasks τ times ago have been finished. The idea is to deliver as many updates as possible w/o missing any updates older than a given age. For more info, see this paper ([More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server](http://www.cs.cmu.edu/~seunghak/SSPTable_NIPS2013.pdf)).

![](/files/-MQ0bPMpqXcEACGK5aWd)

## Implementation

The servers store the parameters (key-value pairs) using consistent hashing (Sec. 4.3). For fault tolerance, entries are replicated using chain replication (Sec. 4.4). Different from prior (key, value) systems, the parameter server is optimized for range based communication with compression on both data (Sec. 4.2) and range based vector clocks (Sec. 4.1).

1. **Vector Clock**: In the naive implementation, each key-value pair is associated with a vector clock (VC) which records the time of each individual node on this key-value pair. This requires O(nm) space complexity, where n = #nodes and m = #parameters. To optimize this, the authors observe that parameters share the same timestamp due to the range-based communication pattern of the PS. As a result, they can be compressed into a single range VC. This requires O(nk) vector clocks, where n = #nodes and k = #unique ranges communicated by the algorithm. k is usually much smaller than m.
2. **Messages**: Messages sent between nodes/node groups consist of a list of (key, value) pairs in the key range R and the associated range vector clock. Both shared parameters and tasks (taskID, args or return results) can be communicated. Training data often remains unchanged between iterations (same key lists are sent again), and values may contain many zero entries. Hence, the key lists are cached (**key-caching**, so the sender only needs to send a hash of the list rather than the list itself), and we only need to use **value-compression** to send nonzero (key, value) pairs (by using a compression library to compress messages and remove zeros).
3. **Consistent Hashing**: Keys and server node IDs are both inserted into the hash ring (see Fig. 7).
4. **Replication and Consistency**: Each server node holds a replica of the k counterclockwise neighbor key ranges relative to the one it owns. The nodes holding the extra copies are denoted as slaves of the appropriate key range.
5. **Server Management**: When a server joins, a key range is assigned by the server manager. The new server fetches the range of data and k additional ranges to keep as slave. Fetching the data requires two phases. Finally, the server manager broadcasts the node changes. The departure is similar to a join.
6. **Worker Management**: When a worker joins, the task scheduler assigns a range of data. The worker loads the range of training data (w/o a two-phase fetch), and pulls the parameters from servers. Finally, the task scheduler broadcasts the change.

![What constitutes a message](/files/-MQ3SKPmBzsDlFYpWJM7)

![Each key range set may split the range and create at most 3 new vector clocks](/files/-MQ3STQNG_-WE7lhq-sy)

![](/files/-MQ3SGg26RpIbnaB_k6T)

## Evaluation

### Sparse Logistic Regression

![System-B outperforms system-A because of a better algorithm. The PS outperforms system-B because of the efficacy of reducing the network traffic and the relaxed consistency model. The relaxed consistency model also greatly improves worker node utilization.](/files/-MQ3SwoyMoX4mNaK5SET)

![Reduction of network traffic by each system component & the best tradeoff achieved by the bounded delay consistency model.](/files/-MQ3ThnbWL2AXwsvJXiu)

### Latent Dirichlet Allocation

![\~4x speedup is achieved when increasing the #machines from 1000 to 6000](/files/-MQ3U-d0Unvzw9aUUuJt)

### Distributed Sketching

![The good performance is due to (1) bulk communication reducing the communication cost and (2) message compression reducing the average key-value size. Also, the system can recover from failures well.](/files/-MQ3UXjNXVlszbQdWwRn)

## Links

* [Paper PDF](http://www.cs.cmu.edu/~muli/file/parameter_server_osdi14.pdf)
* [Parameter Server for Distributed Machine Learning](https://www.cs.cmu.edu/~muli/file/ps.pdf), the same work at a different venue (NIPS '14)
* [Presentation Video by the author at Tsinghua](https://www.youtube.com/watch?v=SHu5qHTDai8\&ab_channel=ASEStreamLine)
* [Presentation Slides at OSDI '14](https://www.cs.cmu.edu/~muli/file/osdi14_talk.pdf)
* [Course notes on PS from CS 4787 @ Cornell](https://www.cs.cornell.edu/courses/cs4787/2019sp/notes/lecture22.pdf)
* [Course notes on PS from CS 294 @ Berkeley](https://bcourses.berkeley.edu/courses/1413454/files/65798745/download?verifier=kFM8TYOCEAoLPkVzJCDDr8f0oRaUZ03RYgpKlbYg\&wrap=1)
* [Course notes on PS from CS 744 @ UW-Madison](http://pages.cs.wisc.edu/~shivaram/cs744-fa19-slides/cs744-paramserver-notes.pdf)
* [ps-lite on GitHub](https://github.com/dmlc/ps-lite)
* [Xiangfeng Zhu](https://xzhu27.me/)'s [paper reading notes](https://xzhu0027.gitbook.io/blog/ml-system/sys-ml-index/parameter-servers)
* [parameterserver.org by the Wayback Machine](https://web.archive.org/web/20150212084849/http://parameterserver.org/)


# \[2018 OSDI] Gandiva: Introspective Cluster Scheduling for Deep Learning

## One-line Summary

The authors present Gandiva, a cluster scheduling framework that employs techniques like time-slicing, migration, intra-job elasticity, and dynamic priority.

## Paper Structure Outline

1. Introduction
2. Background
3. DLT Job Characteristics
   1. Sensitivity to locality
   2. Sensitivity to interference
   3. Intra-job predictability
4. Design
   1. Mechanisms
   2. Scheduling Policy
      1. Reactive Mode
      2. Introspective Mode
5. Implementation
   1. Scheduler
   2. Modifications to DL toolkits
6. Evaluation
   1. Micro-benchmarks
   2. Model exploration in a multi-job
   3. Cluster experiments: time-slicing and packing
   4. Cluster experiments: time-slicing and migration
7. Related Work
8. Conclusion

## Background & Motivation

Today's DNN schedulers (e.g., YARN, Kubernetes) treat deep learning jobs naively (as if they are traditional big-data jobs): A job is scheduled on a set of GPUs exclusively, and the job holds the GPUs until completion. There are some problems:

1. High Latency (head-of-line blocking): Long DNN jobs have runtimes of hours and days, so we need time-slicing of jobs. However, GPUs are not efficiently virtualizable.
2. Low Efficiency (fixed decision at the job-placement time): Need the ability to migrate jobs, and the sensitivity to locality varies across jobs.

DLT jobs have the following characteristics:

1. Sensitivity to locality: Different models have various levels of sensitivity to intra-server and inter-server locality that a DLT scheduler needs to take into account.
2. Sensitivity to interference: Similarly, different models demonstrate different levels of sensitivity to interference between jobs.
3. Intra-job predictability: DLT jobs' GPU memory usage reveals a pattern (goes up during forward pass of a minibatch and goes down during backward pass). Gandiva leverages this in three ways:
   1. A job can be split into mini-batch iterations
   2. If suspend/resume is performed during the nadir, less amount of memory needs to be copied from GPU to CPU
   3. The progress rate can be profiled to evaluate the effectiveness of mechanisms

![](/files/-MRRZm5iyd_B5aDjjsvo)

![When suspending a job, as GPUs are not efficiently virtualizable, the state needs to be moved from GPU to CPU before suspension.](/files/-MRR_0fe-UVbJEfDGntI)

## Design and Implementation

![](/files/-MRRaVKYFOnrKBuasFIE)

Gandiva employs the following mechanisms:

1. Suspend-Resume and Packing
   1. Suspend-Resume: Intra-job predictability is leveraged to suspend/resume DLT jobs when their GPU usage is at the lowest.
   2. Packing: Run multiple jobs on a GPU simultaneously and let the GPU time-share the jobs, with the premise that the packed jobs do not interfere with each other. It is only considered during overload.
2. Migration: The set of GPUs assigned to a job can be changed for (1) moving time-sliced jobs to vacated GPUs, (2) moving interfering jobs away from each other, and (3) doing de-fragmentation of the cluster. The migration overhead is as little as a second or two.
3. Grow-Shrink: # GPUs available for a job can be increased during idle times and shrank when the load goes up.
4. Profiling: Gandiva profiles each job's time for one forward/backward pass over a minibatch. With this, Gandiva introspects DLT jobs to estimate the rate of progress, e.g. to check if packing helped.

Gandiva's scheduler works in two modes: reactive and introspective. The reactive mode handles events such as job arrivals/departures and machine failures, while the introspective mode monitors and optimizes job placement to improve the overall utilization and the completion time.

![](/files/-MRRh5VGLEQG_kb_bawA)

## Evaluation

![Microbenchmark: Time-slicing](/files/-MRRn67uzbDyeUHy6sT5)

![Microbenchmark: Packing](/files/-MRRnBoNvZOzgJHqSkSV)

![Microbenchmark for AutoML: Gandiva provides much faster hyper-parameter exploration](/files/-MRRoi9r40CiAcPsL0kK)

![Cluster utilization](/files/-MRRoLTJG4ROukcJ0OON)

## New Vocabulary

* Introspection (反省): The examination of one's own conscious thoughts and feelings.

## Links

* [Paper PDF](https://www.usenix.org/system/files/osdi18-xiao.pdf)
* [Presentation audio at OSDI '18](https://www.usenix.org/conference/osdi18/presentation/xiao)
* [Presentation slides at OSDI '18](https://www.usenix.org/sites/default/files/conference/protected-files/osdi18_slides_sivathanu.pdf)
* [Presentation video by Muthian Sivathanu, one of the authors and a UW-Madison alumni](https://www.youtube.com/watch?v=i4YOKOLsyFI\&ab_channel=MicrosoftResearch)


# \[2018 SIGCOMM] Chameleon: Scalable Adaptation of Video Analytics via Temporal and Cross-camera ...

...Correlations

## Summary

Chameleon is a video analytics system that optimizes the tradeoff between resource consumption and accuracy by continuously adapting an application's configurations in real time.

## Background & Motivation

Video analytics pipelines consist of several video processing modules (e.g., count vehicles: decoder -> resize & sample frames -> object detection), each of which has a few configuration knobs (e.g., frame resolution, frame sampling rate, the model used for object detection) that collectively determine both the resource consumption and accuracy of the video analytics application. Our target objective is thus to strike the best tradeoff between resources and accuracy.

The thing is, the best configuration for a video analytics pipeline may vary over time. For example, when there is congestion going on and cars are moving slowly, a lower frame sampling rate saves a huge amount of resources without hurting the accuracy much. This kind of optimization technique is a constant theme in system research, IMO -- for example, see the bit on pixel-level frame differencing in [Reducto](/machine-learning-systems/machine-learning-systems-index/2020-sigcomm-reducto-on-camera-filtering-for-resource-efficient-real-time-video-analytics#background-and-motivation). Anyway, the challenge comes down to how we can continuously adapt to different configurations.

## Design & Implementation

A straw man approach is to periodically profile the configurations, but this is super expensive, because (1) the configuration search space is exponential in size, and (2) executing certain candidate configurations may be orders of magnitude more costly than executing the optimal one. How can we reduce the resource cost of periodic configuration profiling?

![Chameleon's periodic reprofiling pipeline](/files/9rNH8kpcif0PqTfcByZ5)

Due to the non-stationary setting of video analytics applications, traditional modeling approaches like Bayesian optimization are similarly expensive. To address the problem, the authors exploited the domain-specific characteristics of the configurations, namely the temporal and spatial correlations.

* Temporal correlation (good configurations are always good): Although the best configuration varies over time, the top-k best configurations are relatively stable over time, and vice versa. This allows the search space to be significantly pruned.
* Cross-camera correlation (cameras near each other are similar): For example, if there is congestion on the highway, two cameras on that same highway share the same properties (e.g., the velocities and sizes of objects) that affect the optimal configuration. This allows the profiling cost to be amortized across multiple cameras.
* Independence of configuration knobs: An expensive, golden configuration is used to establish the ground truth. To reduce the cost of running this golden config, the authors rely on an empirical observation that (unlike say in DBMS tuning) the knobs are typically independent.

![](/files/l2Tf74SoJqXby7brgHaI)

## Evaluation

![Chameleon good!](/files/4TL3TypMh4GRSxzgqtqB)

![Contribution breakdown](/files/LTfGwTTfizWqSUBrH3WG)

## Links & References

* [Paper PDF](https://people.cs.uchicago.edu/~junchenj/docs/Chameleon_SIGCOMM_CameraReady_faceblurred.pdf)
* [Presentation slides at SIGCOMM '18](https://conferences.sigcomm.org/sigcomm/2018/files/slides/paper_5.2.pptx)


# \[2018 NIPS] Dynamic Space-Time Scheduling for GPU Inference

## Summary

The authors evaluated different multiplexing (time & space) techniques for ML inferences on GPUs and proposed ideas to achieve the best tradeoff across criterias.

## Background & Motivation

Almost all cloud inference service providers/frameworks assign each model an exclusive GPU. This, combined with the small batch sizes used in an online setting, results in low hardware utilization. Current approaches that multiplex workloads have different tradeoffs, and there is no single solution that wins on all criteria.

<table><thead><tr><th>Approach</th><th width="150">Utilization</th><th width="179.18672199170126">Performance (throughput/latency)</th><th width="223">Predictability/Performance Isolation</th></tr></thead><tbody><tr><td>Exclusive access</td><td>Poor</td><td>Good</td><td>Good</td></tr><tr><td>Time multiplexing (CUDA context switching)</td><td>Average</td><td>Poor</td><td>Good</td></tr><tr><td>Spatial multiplexing</td><td>Good</td><td>Average</td><td>Poor</td></tr></tbody></table>

## Design & Implementation

![](/files/ppSB52QnYqqRbFCUogwv)

The authors proposed software-level fusion of kernel operators across multiple inference jobs to get the best of all worlds.

## Links & References

* [Paper PDF](http://learningsys.org/nips18/assets/papers/102CameraReadySubmissionGPU_Virtualization%20\(8\).pdf)


# \[2019 ATC] Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads

## One-line Summary

This paper presents a characterization study of large-scale GPU clusters for DNN training. It uncovers some inefficiencies in cluster utilization and presents some lessons for better cluster manager decisions.

This paper means a lot to me in that this is the first paper that I read through thoroughly :)

## Paper Structure Outline

1. Introduction
2. Philly: System Overview
   1. Workloads
   2. Cluster Architecture
   3. Job Scheduling and Execution Workflow
   4. Data Collection and Analysis
3. Impact of Locality Awareness
   1. Queueing Delays
      1. Impact of Locality-Driven Scheduling
   2. GPU utilization
      1. Impact of Distributed Learning
4. Training Progress and Completion
   1. Effectiveness of Training Iterations
   2. Job Failures
      1. Failure Classification
      2. Failure Frequency
      3. Runtime to Failure
      4. Impact of GPUs Allocated
5. Design Implications for Future Schedulers
6. Related Work
7. Conclusion

## Background & Motivation

DNN-based workloads are different from traditional big data analytics workloads in two ways:

1. Cluster utilization: GPUs represent a monolithic resource that cannot be shared at a fine granularity across users
2. Workload: Deep learning frameworks require gang scheduling, reducing the flexibility of scheduling and job's elasticity of runtime failures

The authors first present an overview of Philly, a large, multi-tenant GPU-based cluster for production-scale deep learning tasks. Then, they present a detailed workload characterization and study how factors like gang scheduling, locality requirements, and failures might affect cluster utilization.

## Microsoft Philly

![](/files/-MPukF-TbFd7hbFD7jjJ)

![](/files/-MPuoh8jz5P1ZfyDS8B8)

The three main steps in Fig. 1 are:

1. **Incoming jobs and queuing**: The scheduler needs to perform gang scheduling while being locality-aware. Each production group is provided with a virtual cluster and a quota (in terms of #GPUs to each virtual cluster).
2. **Job placement and utilization**: The scheduler aims to maximize locality and minimize fragmentation of resources (from smaller jobs, e.g. 1-GPU jobs). There is a trade-off between colocation and distribution, though, as placing different jobs on the same server could lead to lower GPU utilization (because of interference in shared resources like RDMA and PCIe).
3. **Training progress and completion**: Jobs can finish with three statuses: passed, killed, or unsuccessful. Failed jobs are retried a few times to overcome non-deterministic failures.

The logs are collected over a 75-day period and it consists of 96260 jobs over 14 virtual clusters. There are three main sources of the logs:

1. YARN scheduler log: job arrival time, # GPUs requested, GPU allocation status, job finish status
2. stdout & stderr logs from ML frameworks
3. Ganglia monitoring system log: Per-minute statistics on hardware usage (CPU, memory, network, GPU utilization)

## Impact of Locality Awareness

### Queueing Delays

![The scheduler works in practice to trade-off locality for lower scheduling delay](/files/-MPvAxJzsfmSyGZ8CRu_)

![There are two types of delays: Fair-share denotes fairness (which is common in conventional data analytics clusters), while fragmentation denotes locality requirement and resource fragmentation (which is more prevalent in DL clusters).](/files/-MPvB2GtGEOGSh9D31Ps)

### GPU Utilization

![GPU utilization is low (lower in distributed training) because of (1) distribution across servers and (2) intra-server interference](/files/-MPv9GRo7PgA0wcwOG_E)

![](/files/-MPvDeDkLDsPND50Rt6O)

![GPU utilization when running 8 and 16 GPU jobs on dedicated servers](/files/-MPvDknJwZAMz46qi8bW)

![In general, DL training jobs underutilize GPU processing cycles regardless of their job sizes.](/files/-MPvDuz9nycrlg1xRzFU)

Relaxing locality constraints:

* High intra-server locality
  * Pros
    * High communication efficiency
  * Cons
    * Long queueing time
* Low intra-server locality
  * Pros
    * Low queueing time
  * Cons
    * Contention in the use of network
    * Risk of intra-server interference (across jobs)

## Training Progress and Completion

![A significant fraction (30.7%) of jobs are either termintated by users or are unsuccessful. These jobs constitute \~55% of the total GPU time.](/files/-MPvEB8_f8S9JYz6lkRu)

As observed in the above table, it is important to understand the reasons behind these failures, as fewer unsuccessful jobs would mean more resources for successful jobs.

### Training Iterations

![](/files/-MPvF56pJRPEVgSWY_tX)

\~80% of passed jobs require all epochs executed to reach the lowest loss. However, an average of 62% and 56% (for passed jobs and killed jobs, respectively) GPU times for each job are used to improve the convergence accuracy by merely 0.1%. This suggests that jobs can be terminated early to save considerable resources.

### Job Failures

![](/files/-MPvG3xk2H41HjMR7jcM)

The failures happen across the whole stack: Infrastructure (GPU, HDFS, resource scheduler), ML frameworks (PyTorch, TensorFlow), and user programs (shitty code :P). A failure classifier is used to analyze the causes of the job failures. The most important failures are from these classifications:

1. **Incorrect inputs**: Model files or input data stored in the external HDFS storage cannot be read
2. **Semantic error**: Errors that happen due to library version mismatch or other dependencies of the user training program not being setup correctly
3. **Model checkpoint error**: The job is not able to successfully create a model checkpoint after a certain number of epochs complete. This is usually due to either transient error in HDFS or HDFS name node recovery
4. **MPI runtime failure**: This is usually due to either a failure of network connection to peer MPI process, or possibly an internal failure of the MPI daemon itself
5. **Job preempted**: YARN reclaims any GPU currently in use to schedule another job
6. **Invalid memory access**: Training job dies because of violating access on memory address space e.g., using an invalid pointer value, or having race condition while copying data. This failure is observed in both CPU memory and memory allocated for GPU access

An analysis on the failure frequency show that:

1. Failures repeat for the same job/user
2. User/programming errors lead to a lot of failures

An analysis on the runtime to failure (RTF) shows that:

1. RTF exhibits high variability with many short RTFs
2. Infrastructure failures occur infrequently but have much longer RTF

The authors also find that large jobs with programming semantic errors tend to fail a while after execution.

## Guidelines

1. **Prioritize locality**: As the lack of locality impacts both utilization and job runtime, and because DNN training jobs are long-running, schedulers should trade queueing delay for adhering to locality constraints.
2. **Mitigate interference**: As different jobs on a single server might interfere with each other, schedulers should aim to isolate the jobs on dedicated servers while implementing techniques like migration for defragmentation to support the locality constraints of jobs that need more GPUs.
3. **Improve failure handling**: To catch failures early before they are scheduled on a cluster and thus prevent resources from being wasted, each incoming job should be scheduled on a small dedicated pool of servers/a single GPU to catch simple programming and configuration errors from multi-GPU jobs. Another possible improvement is for clusters to predictively mitigate failures by proactively observing related failures. For example, the scheduler should stop retrying for failure categories like incorrect data input and continue retrying for network timeouts.

## Links

* [Paper PDF](https://www.usenix.org/system/files/atc19-jeon.pdf)
* [Full presentation video at USENIX ATC '19](https://www.youtube.com/watch?v=FoA1M7wAZ3I\&ab_channel=USENIX)
* [Lightning talk at USENIX '19](https://www.youtube.com/watch?v=ClEpCcZru_Q\&ab_channel=MyeongjaeJeon)
* [Full presentation slides](https://www.usenix.org/sites/default/files/conference/protected-files/atc19-slides-jeon.pdf)
* [Lightning talk slides](https://www.usenix.org/sites/default/files/conference/protected-files/atc19_slides_lt_jeon.pdf)
* [Philly traces on GitHub](https://github.com/msr-fiddle/philly-traces)
* [Project Fiddle](https://www.microsoft.com/en-us/research/project/fiddle/)


# \[2019 NSDI] Tiresias: A GPU Cluster Manager for Distributed Deep Learning

## One-line Summary

Tiresias is a cluster manager that uses (1) a Two-Dimensional Attained Service-Based Scheduler to minimize the average JCT and (2) a placement algorithm to relax the consolidation constraints for some models.

## Paper Structure Outline

1. Introduction
2. Background and Motivation
   1. Distributed Deep Learning (DDL)
   2. Challenges
   3. Potential for Benefits
3. Tiresias Design
   1. Overall Architecture
   2. Scheduling
      1. Why Two-Dimensional Scheduling?
      2. Two-Dimensional Attained Service-Based Scheduler (2DAS)
      3. Priority Discretization
   3. Placement
      1. Profiler
      2. The Placement Algorithm
   4. Summary
4. Implementation
5. Evaluation
   1. Experimental Setup
   2. Tiresias in Testbed Experiments
      1. JCT Improvements
      2. Cluster-Wide GPU Utilization
      3. Sources of Improvements
      4. Overheads
   3. Tiresias in Trace-Driven Simulations
      1. Simulator Fidelity
      2. JCT Improvements
   4. Sensitivity Analysis
      1. Impact of Queue Thresholds
      2. Impact of K (number of priority queues)
      3. Impact of PROMOTEKNOB
6. Discussion and Future Work
7. Related Work
8. Conclusion

## Background & Motivation

More and more deep learning jobs are being trained on GPU clusters. The design objectives of GPU managers include:

1. Minimizing cluster-wide average job completion time (JCT)
2. Achieve high resource (GPU) utilization

There are some challenges:

### Unpredictable Training Time

Algorithms like SJF and SRTF (despite good in minimizing the avg. JCT) require the prior knowledge of a job's (remaining) execution time, which is often unknown for DL training jobs. Existing solutions include predicting the remaining execution time using the smooth loss curve. However, not all jobs have smooth loss curves & run to completion IRL. Thus, state-of-the-art managers are naive.

![Expectation vs. reality. Here, job 2 gets terminated early.](/files/-MQiL6oQXQDRKbLNfpZJ)

### Over-Aggressive Job Consolidation

Existing cluster managers try to send jobs onto as few as possible number of servers to improve locality and reduce the network bottleneck. This leads to fragmented free GPUs in the cluster and longer queueing delays for jobs that require a large number of GPUs. In this work, the authors found that only some of the models have structures that are sensitive to placement.

### Preemption is Costly

Existing clusters do not preempt jobs because of the large time overhead.

![Motivation: None of the existing solutions handles the two problems well!](/files/-MQiLivMI3o-BGTOMYuC)

## Design and Implementation

Tiresias addresses the two aforementioned issues by:

1. Using an age-based scheduler to minimize JCT w/o complete knowledge of jobs
2. Doing model profile-based placement to place jobs w/o additional information from users

![](/files/-MQms8RD9N7bxzE6BcSS)

A job lifecycle is as follows:

* 1: As soon as a job is submitted, its GPU requirements are known, and the job is appended to a WAITQUEUE
* 2: Scheduler
  * 2a: The scheduler schedules jobs from the WAITQUEUE
  * 2b: The scheduler preempts running jobs from the cluster to the WAITQUEUE
* 3: The placement module accepts starting/resuming jobs for GPU allocation
* 4: For new jobs, the profiler decides if they should be consolidated or not

### Scheduling

In DDL job scheduling, both the spatial (#GPUs) and temporal (time) aspects of the jobs need to be considered. In Tiresias, the authors present a Two-Dimensional Attained Service-Based Scheduler (2DAS) that generalizes:

1. The classic least-attained-services (LAS) scheduling discipline (2D-LAS)
2. The Gittins index policy (2D-Gittins Index)

...to consider both spatial and temporal aspects of the jobs. LAS prefers jobs that received less service, while the Gittins index value represents how likely the job that has received some amount of service can complete within the next service quantum.

![Each job is assigned a priority based on its attained service. If no job duration information is provided, the LAS algorithm is applied where the priority is inverse to its attained service. If the distribution of job duration is provided, a job's priority equals its Gittins index value.](/files/-MQmuVo6hPrpWFoPJSuY)

![The avg JCTs are 9.3, 10, and 11.7 for SRSF, 2D-Gittins, and 2D-LAS. For a more detailed walkthrough of this example, see the NSDI presentation video.](/files/-MQmwDAkJGg2o-fwt1Ns)

#### Priority Discretization

Using continuous priorities lead to preemptions and resumptions (which are costly), and continuous preemption degenerates 2DAS to fair sharing by time-division multiplexing. In Tiresias, a MLFQ is used for priority discretization.&#x20;

![K queues are maintained](/files/-MQmyYhPv70UXC04-fw7)

### Placement

![](/files/-MQn-AFvbQGF4Zy_2tUU)

![](/files/-MQn-VEarekiB_7m6Toz)

The skew level of a model is a good predictor of whether a job benefits from consolidation, as the message size distribution depends on the tensor size distribution of the model. Observing the network communications sent out by the PS can inform us of the skew. The authors built a RDMA-level traffic monitoring tool for Tiresias as most production DL jobs use RDMA (e.g., InfiniBand in Microsoft) for PS-worker communication. The placement algorithm compares the model skew with a threshold, and if the skew is larger than the threshold, consolidation is performed.

## Evaluation

![](/files/-MQn1cim8kq3k5REHnQt)

![](/files/-MQn1oTTL86mpFE94sl3)

![](/files/-MQn23LJaQw0_iqtQs0h)

## New Vocabulary

* [Apache Hadoop YARN](https://hadoop.apache.org/docs/current/hadoop-yarn/hadoop-yarn-site/YARN.html)
* SRSF (Shortest Remaining Service First): The multiplication of a job's remaining time and the number of GPUs.
* Preemption: The act of temporarily interrupting a task being executed (w/o requiring its cooperation) and with the intention of resuming the task later. Such changes are known as context switches.
* ILP formulation: The mathematical formulation of an optimization problem in which variables are restricted to integer values and the constraints and objective function are linear. Mixed integer linear programming (MILP) refers to optimization problems in which some of the variables are continuous.

## Links

* [Paper PDF](https://www.usenix.org/system/files/nsdi19-gu.pdf)
* [Presentation video at NSDI '19](https://www.youtube.com/watch?v=-RtcM0oz1lQ)
* [Presentation slides](https://www.usenix.org/sites/default/files/conference/protected-files/nsdi19_slides_gu.pdf)
* [Tiresias on GitHub](https://github.com/SymbioticLab/Tiresias)
* [Xiangfeng Zhu's paper reading notes](https://xzhu0027.gitbook.io/blog/ml-system/sys-ml-index/tiresias-a-gpu-cluster-managerfor-distributed-deep-learning)
* [\[EuroSys 18'\] Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters](https://i.cs.hku.hk/~cwu/papers/yhpeng-eurosys18.pdf)


# \[2019 SOSP] ByteScheduler: A Generic Communication Scheduler for Distributed DNN Training ...

...Acceleration

## Summary

Priority-based communication scheduling + tensor partitioning: acceleration! Fig. 2 is a good toy example that showcases why the default order of communication (FIFO) in current ML frameworks is suboptimal. However, prior systems ([P3](https://arxiv.org/abs/1905.03960) and [TicTac](https://arxiv.org/abs/1803.03288)) that try to tackle this are not generic, in the sense that each of them only targets one combination of DL framework & network stack. Moreover, existing work does not adapt well to different system setups.

In contrast, ByteScheduler is generic (framework/communication method-agnostic), which required some intricate engineering efforts/techniques. Also, ByteScheduler proposes a BO-based auto-tuning algorithm to search for the best system parameters (e.g., tensor partition sizes) under different environments (DNN models, communication paradigms, bandwidth, etc.).

![The training speedup of priority scheduling is 44%!](/files/tFhTDgh8KTtHFULAl2YX)

## Background & Motivation

In distributed DNN training using data parallelism, the default ML framework engines execute communication operations in a FIFO order, as the underlying communication stack (PS/all\_reduce, TCP/RDMA) is inherently based on FIFO queues. However, this is suboptimal: if some communication operations are prioritized, the training can be sped up.

Tensor partitioning is a technique that enables more flexible priority-based scheduling. Without partitioning, a large, low-priority tensor might block high-priority tensors. Instead, the tensors can be partitioned before being en-queued, and high-priority tensor partitions can jump ahead of the queue after they arrive.

## Design & Implementation

### Which layer should ByteScheduler be implemented in to make it more general?

![](/files/dEcZDYxeOsxDqa4Ql7bU)

The five original layers are shown above. After some thoughtful thinking, the authors placed ByteScheduler at the high-level API implementation layer in the framework. For each ML framework, a shim layer ("plugin") is designed to wrap the original operation into a unified "CommTask" abstraction.

### Unified abstraction for communication tasks

A single interface, `Core.enqueue(CommTask)`, is exposed to the plugins. Once a communication tensor arrives, it is first wrapped into a CommTask. Then, the Core partitions it into SubCommTasks and decides when to send each. Four CommTask interfaces are implemented:

* CommTask.partition(size): Partitions a CommTask into multiple SubCommTasks with tensors no larger than the specified size. This invokes a callback in the plugin, as tensor partitioning is framework-dependent. This has a low overhead, as DL frameworks provide zero-copy APIs for tensor partitioning.
* CommTask.notify\_ready(): The engine uses this interface to notify the Core about a tensor being ready, so it can be actually scheduled.
* CommTask.start(): The Core calls this to let engines and the underlying communication stacks send the tensor.
* CommTask.notify\_finish(): The framework engines notify the Core once the communication (push/pull/all\_reduce) finishes so that the Core can continue scheduling more Tasks.

### Interaction with framework engines and crossing the global barrier

![](/files/sN0YTSuT5lyOxldvqTy3)

![](/files/7TH9KPlLVWQxjMrtnYhK)

### Auto-tuning partition size and credits using Bayesian Optimization

![](/files/tdx4SDVItKbh4HchjMYF)

## Comparisons with P3 and TicTac

P3 and TicTac, both in MLSys '19, employ similar ideas and techniques (transmission prioritization via tensor partitioning & reordering). However, both systems target specific training setups (e.g., P3 targets MXNet PS + TCP), while ByteScheduler devotes a significant chunk of engineering efforts on the system design so that it not only outperforms prior systems but also works well with different training configurations.

## Evaluation

![](/files/VXu7yVnMItimB5VSN8pt)

The paper provided the reasoning for different speedups in different setups.

## Links & References

* [Paper PDF](https://i.cs.hku.hk/~cwu/papers/yhpeng-sosp19.pdf)
* [Presentation video at SOSP '19](https://www.youtube.com/watch?v=UL1_69lI9BE)
* [bytescheduler on GitHub](https://github.com/bytedance/byteps/tree/bytescheduler/bytescheduler)


# \[2019 SOSP] PipeDream: Generalized Pipeline Parallelism for DNN Training

## One-line Summary

PipeDream uses pipeline parallelism to combine intra-batch parallelism with inter-batch parallelization and reduce the communication overhead in DNN training.

## Paper Structure Outline

1. Introduction
2. Background & Related Work
   1. DNN Training
3. Parallel Training in PipeDream
   1. Pipeline Parallelism
   2. Partitioning Layers Across Machines
   3. Work Scheduling
   4. Effective Learning
   5. GPU Memory Management
4. Implementation
5. Evaluation
   1. Experimental Setup
   2. PipeDream vs. Data Parallelism
   3. Value of Data Parallelism in stages
6. Conclusion

## Background & Motivation

Existing parallelization techniques (data, model, hybrid) use intra-batch parallelization where each iteration of the optimization algorithm is parallelized across a set of workers. In data parallelism, there exists a **high communication overhead** at a large scale, while in vanilla model parallelism, **system resources are under-utilized**.

![](/files/-M_bDv2Vbg2vVZfVz2SI)

![](/files/-M_bMXB8lGgwMvvqXHkg)

## Design and Implementation

### Pipeline parallelism

With pipeline parallelism, multiple minibatches are injected into the pipeline one after another (Fig. 4).

Pipeline parallelism is superior because of 2 reasons:

1. **Pipelining communicates less**: Compared to DP where the whole model is communicated in an all-to-all fashion, PipeDream reduces the amount of inter-worker communication by only communicating intermediate inputs & outputs across consecutive layers on different workers peer-to-peer.&#x20;
2. **Pipelining overlaps computation and communication**: Asynchronous communication of intermediate results in significant overlap of communication with the computation of a subsequent minibatch (Fig. 5).

![](/files/-M_bL034W5EhJEA_E_Km)

![](/files/-M_bL2jXsKOFoMqEOZ9f)

### Shenanigan 1: Work/model partitioning

PipeDream partitions the training of a model into stages in a pipeline. Ideally, each stage would have the same throughput rate, otherwise, there would be a load imbalance. The communication between stages should also be minimized as much as possible to improve the overall throughput. PipeDream uses an optimizer to output a balanced pipeline. To minimize the time taken by the slowest stage, PipeDream uses a partitioning algorithm that takes in:

1. Do a short profiling run of 1000 minibatches. For each layer,
   1. The total computation time across the forward & backward passes
   2. The size of the output activations & input gradients
   3. The size of weight parameters

and outputs {a partitioning of layers into stages, the number of workers for each stage, and the optimal number of in-flight minibatches to keep the training pipeline busy}.

### Shenanigan 2: Work scheduling

Each stage should decide whether to perform a forward pass or a backward pass on different minibatches. During the startup state, PipeDream admits enough minibatches to keep the pipeline full in steady state. Then, each stage alternates between a forward pass for a minibatch & a backward pass for a different minibatch (1F1B schedule). For a data parallel configuration, an 1F1B-RR (round robin) schedule is used.

![](/files/-M_bW9Ty_WBvK7eFy4XQ)

### Shenanigan 3: Weight version mismatch

Naive pipelining leads to mismatch in weight versions. For example, in Fig. 4, for minibatch 5, on stage 1, the forward pass is performed after the backprop of minibatch 1, whereas the backprop is performed after the backprop of minibatches 2-4. This discrepancy results in invalid gradients and may prevent model convergence. PipeDream's workaround is to use weight stashing to store multiple versions of the weights, one for each active minibatch. Even with the extra memory footprint of stashing the weight versions, PipeDream's peak per-worker memory usage is on par with data parallelism.

![](/files/-M_bWhC9KqBehVAG2tbi)

## Evaluation

![](/files/-M_bYvVCPmCaE5qT3akq)

![](/files/-M_bZEWav8hvD9e5xmIj)

## Links

* [Paper PDF](https://arxiv.org/pdf/1806.03377.pdf)
* [Presentation video at DATA + AI Summit Europe](https://databricks.com/session_eu20/generalized-pipeline-parallelism-for-dnn-training)
* [Presentation video at SOSP '19](https://sosp19.rcs.uwaterloo.ca/videos/D1-S1-P1.mp4)
* [PipeDream on GitHub](https://github.com/msr-fiddle/pipedream)


# \[2019 SOSP] Parity Models: Erasure-Coded Resilience for Prediction Serving Systems

## Summary

This work uses erasure codes for reducing tail latency in ML inference.

![](/files/2DuWFtqRXL1Qj7oqG2At)

## Background & Motivation

ML inference, typically done in large-scale clusters, is latency-sensitive. However, slowdowns (network/compute contention) and failures in clusters might cause inference queries to miss their SLOs. This work aims to alleviate the effects of slowdowns and failures to reduce tail latency.

Erasure codes is a technique widely deployed in systems (e.g., storage systems, communication systems) for resource-efficient data corruption prevention. The difference between erasure codes for ML serving and for traditional settings is the need to handle computation over inputs. In other words, the encoding and decoding must hold over computation F.&#x20;

![](/files/niRIOJsiCVQCqevZbDEn)

![](/files/aEuDmZqOJPji2DcrrMCD)

The problem boils down to: How do we design the erasure codes for ML inference?

## Design & Implementation

Current approaches hand-craft erasure codes, which is relatively straightforward for a linear computation F, but is far more challenging for non-linear computations like ML serving. The authors overcome this challenge by taking a learning-based approach. Slap a NN, problem solved!

![](/files/CFKefb2PgzOeqPQGuWJv)

But wait! Using NNs for encoders/decoders is computationally expensive. Instead, the authors use simple, fast encoders/decoders and operate over parities using a new computation model, namely the parity model. In this diagram, the parity model takes as input parity queries P = X1 + X2 and outputs Fp(P) = F(X1) + F(X2), which can later be used to reconstruct F(X2).&#x20;

![](/files/lBDLUuBaJuEYPp5LKYSf)

What a brilliant idea. We can also tweak the settings of this process, e.g. using a larger degree of query multiplexing (erasure codes parameter), or using different encoders/decoders instead of the simple summation encoder.&#x20;

![For example, for image tasks, we can downsample multiple queries and concatenate them into a single query](/files/UUDLxHjSEhVY7NiWvH2f)

## Evaluation

Note that although there is an accuracy loss, the inaccuracy only comes into play when predictions are otherwise slowed down or straigh up failed, which violate the latency requirements. This still sounds like a pretty good tradeoff, although I am curious about the accuracy loss on larger datasets and models with more complex architectures.

![Evaluation of the accuracy loss](/files/Sy8AW0zIDT7YNSU0AxA0)

![Tail latency reduction in the presence of resource contention](/files/5ljEOkD7B3uHPdpMHZwp)

## Links & References

* [Paper PDF](https://www.cs.cmu.edu/~rvinayak/papers/sosp2019parity-models.pdf)
* [Presentation video at SOSP '19](https://www.youtube.com/watch?v=NlrH_4HNJI4) (one of my favorite talks)
* [parity-models on GitHub](https://github.com/Thesys-lab/parity-models)


# \[2019 NIPS] GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

## One-line Summary

GPipe presents pipeline parallelism on top of model parallelism for better hardware utilization.

## Paper Structure Outline

1. Introduction
2. The GPipe Library
   1. Interface
   2. Algorithm
   3. Performance Optimization
3. Performance Analyses
   1. Performance Overhead Breakdown
4. Image Classification
5. Massive Massively Multilingual Machine Translation
6. Design Features and Trade-Offs
7. Conclusion

## Background & Motivation

Generally, the larger (# parameters) a model is, the higher accuracy it yields. However, we are hitting the bottleneck on the memory of a single accelerator.&#x20;

![](/files/-M_ly2jT-dayPXcUxWkv)

Traditional approaches to resolve this include:

* Recompute forward activations during the backprop calculations
  * Trades compute for memory: Activations are not stored but have to be recomputed
* Memory swap: Copy activations back to the CPU or main memory, and then copy them back
  * Communication between the CPU & accelerator becomes the bottleneck
* Parallelism: Split the computation between multiple "workers"
  * Data parallelism: Works well when there are a few parameters and lots of data and computations
  * Model parallelism: Works well when the number of parameters is large compared to data

## Design and Implementation

Vanilla model parallelism is not time efficient because of the serialized dependencies, and it also leads to system underutilization (bubbles between compute blocks).

GPipe uses pipeline parallelism to integrate data and model parallelism by dividing a minibatch into smaller microbatches so that accelerators can operate in parallel on different microbatches.

![](/files/-M_lyEFVxHvLwaszKUoA)

The user is required to define (1) the number of model partitions, (2) the number of micro-batches, and (3) the sequence/definition of the layers that define the model.&#x20;

GPipe uses re-materialization to reduce the activation memory requirements. During the forward pass, only the output activations at the partition boundaries are stored. During the backward pass, the composite forward function is recomputed at each accelerator. The authors found that the bubble overhead (idle time on every accelerator) is negligible when M, the number of micro-steps, is bigger than 4 \* K, the number of accelerators, as the recomputations during the backward pass can be scheduled w/o waiting for the gradients from earlier layers.

## Evaluation

![](/files/-M_m5d92OvOj7gFMdY8x)

GPipe scaled up AmoebaNet in both the number of channels and the size of the input image. The giant models report competitive results on all target datasets.

![](/files/-M_m5fdf7Q97jdiWd-ul)

![Overhead breakdown](/files/-M_m4r0uf3e22_gzUzxT)

## Links

* [Paper PDF](https://papers.nips.cc/paper/2019/file/093f65e080a295f8076b1c5722a46aa2-Paper.pdf)
* [Presentation video by Kartik Nanda](https://www.youtube.com/watch?v=9s2cum25Kkc)
* [A GPipe implementation in PyTorch on GitHub](https://github.com/kakaobrain/torchgpipe)
* [The integration of GPipe is done on tensorflow/lingvo](https://github.com/tensorflow/lingvo)


# \[2019 SC] ZeRO: memory optimizations toward training trillion parameter models

## Summary

ZeRO is a new distributed training paradigm that vastly improves the memory efficiency of large-scale model training.

![](/files/S024JDDG7cecgkoqrl5N)

## Background & Motivation

Existing solutions for distributed training include data parallelism (DP), model parallelism (MP), pipeline parallelism (PP), 3D parallelism, CPU offloading, etc., each with a couple of catches:

* DP has good compute/communication efficiency but poor memory efficiency. In DP, each parallel worker holds a full copy of the model, so each device quickly runs out of memory (e.g., for models with > 1.4B parameters on a GPU with 32 GB GRAM).
* MP has good memory efficiency but poor compute/communication efficiency. It splits the model vertically, requiring significant communications between each layer on different devices. As a result of that, the efficiency quickly degrades when devices become far apart from each other: for example, when training a 40B-parameter model across two DGX nodes, each GPU's computing efficiency (tflops) is only 5% of the hardware peak.
* Different implementations of PP have different issues. For example, [G-pipe](/machine-learning-systems/machine-learning-systems-index/gpipe-efficient-training-of-giant-neural-networks-using-pipeline-parallelism) requires a batch size proportional to the number of pipeline partitions, and a large batch size might hurt convergence. [PipeDream](/machine-learning-systems/machine-learning-systems-index/pipedream-generalized-pipeline-parallelism-for-dnn-training) is very memory inefficient due to weight stashing.

Is there a way to achieve the best of all worlds? The Microsoft folks first take a look at the spectrum of memory consumption in large-model training, and they classify the memory consumption into two parts:

1. Model states: parameters, gradients, and optimizer states (e.g., momentum and variances in Adam). These take the majority of the memory in large-model training.
2. Residual states: activations, temporary buffers, and unusable fragmented memory.

The authors develop ZeRO-DP to address (1) and ZeRO-R to address (2).

## Design & Implementation

A key insight of ZeRO-DP is that both DP and MP keep all the model states needed over the entire training process, but not everything is required all the time. For example, parameters corresponding to each layer are only needed during the forward/backward propagation of the layer.

### ZeRO-DP Stage 1: Optimizer State Partitioning

The optimizer states are grouped into N equal partitions (N is the DP degree), such that worker i only stores and updates the optimizer states of partition i (1/N of the total optimizer states and parameters).

### ZeRO-DP Stage 2: Gradient Partitioning

As each DP process only updates its corresponding parameter partition, it only needs the reduced gradients for the corresponding parameters. Therefore, as each gradient of each layer becomes available during the backward propagation, we only reduce them on the DP process responsible for updating the corresponding parameters. After the reduction, we no longer need the gradients and their memory can be released. The total communication volume is the same as vanilla DP (all-reduce = reduce-scatter + all-gather), because in this setup, a reduce-scatter is performed on the gradients and an all-gather on the parameters.

### ZeRO-DP Stage 3: Parameter Partitioning

In this stage, each process only stores the parameters corresponding to its partition. When the parameters outside of its partition are required for forward/backward propagation, they are received from the appropriate DP process through broadcast. This technique only increases the total communication volume to 1.5x of a DP baseline (two all-gathers of the parameters are required for forward/back prop, and the gradients need to be reduce-scattered), while reducing the per-worker memory consumption by N times. With all these optimizations, trillion-parameter models can be trained on thousands of modern-day GPUs.

![](/files/tF4IIIJ4w35vv3cUebo4)

### ZeRO-R

A lot of follow-up works of ZeRO focuses on ZeRO-DP, so I'll come back to read this section later

## Evaluation

![](/files/LQr9hq6Ye64uJluDC3eo)

## ZeRO-Infinity and ZeRO-Offload

ZeRO-Infinity and ZeRO-Offload are follow-up systems that offload data and compute to CPUs and NVMe.

![](/files/XWAIrSCwNS4S7oXpdTwQ)

![](/files/uTqEjSVNdLh4gADefdsY)

## Links & References

* [Paper PDF](https://arxiv.org/pdf/1910.02054.pdf)
* [Presentation video at a webinar](https://www.youtube.com/watch?v=zqsOEzKZX2Y)
* Blog: [ZeRO & DeepSpeed: New system optimizations enable training models with over 100 billion parameters](https://www.microsoft.com/en-us/research/blog/zero-deepspeed-new-system-optimizations-enable-training-models-with-over-100-billion-parameters/). This includes a nice video that explains how ZeRO-DP works.
* [DeepSpeed on GitHub](https://github.com/microsoft/DeepSpeed)


# \[2020 OSDI] Gavel: Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads

## One-line Summary

Gavel is a scheduler that takes into account the performance heterogeneity of underlying accelerators (GPUs, TPUs, etc.) when training DNN jobs. Gavel makes existing policies heterogeneity-aware and incorporates them as optimization problems.

## Paper Structure Outline

1. Introduction
2. Background
   1. DNN Training
   2. Performance Optimizations
3. System Overview
   1. Heterogeneity-Aware Policies
   2. Round-based Scheduling Mechanism
   3. Throughput Estimator
   4. Limitations and Non-Goals
4. Scheduling Policies
   1. Max-Min Fairness as an Optimization Problem
   2. Other Policies as Optimization Problems
   3. Hierarchical Scheduling Policies
   4. Properties of Gavel's Policies
5. Scheduling Mechanism
6. Implementation
7. Evaluation
   1. Experiment Setup
   2. End-to-End Results on Physical Cluster
   3. End-to-End Results in Simulation
   4. Scalability of Heterogeneity-Aware Policies
   5. Efficacy of Scheduling Mechanism
   6. Impact of Throughput Estimation
8. Related Work and Discussion
9. Conclusion

## Background & Motivation

For a cluster scheudler, choosing the optimal accelerator types is difficult for three reasons:

1. **Performance heterogeneity**: Common deep learning training workloads show heterogeneous performance on different hardware accelerators due to architectural differences. Not being aware of this result in suboptimal allocations.
2. **Generality across policies**: Recent scheudlers like Allox and Gandiva optimize for a single scheduling objective, while in reality, clusters may take differnet (and sophisticated) scheduling policies.
3. **Colocation and optimization optimizations**: The performance benefits of these optimizations should be considered explicitly while optimizing for global scheduling objectives, since these optimizations are more effective when deployed in a heterogeneity-aware way.

![](/files/-M_X0wzuinpZkvdCD6PZ)

## Design and Implementation

![](/files/-M_X37Mq5zzj042cBSGy)

### Allocations as time fractions

Gavel expresses scheduling policies as optimization problems and produces an allocation matrix that indicates the fraction of time each job should spend on the different accelerators between allocation recomputations (new jobs arriving, old jobs finishing, periodic recomputations, etc).&#x20;

![In this case, job 0 spends 60% of the time on a V100 and 40% of the time on a P100.](/files/-M_X3Ibbp1bk2EIG_M_Y)

![](/files/-M_X3Vv7spactJlJ7OGD)

### Effective throughput

![m: model, X: allocation matrix, T: throughput matrix](/files/-M_XzZAYNdv_p3NRwDuY)

### Policies as optimization problems

Gavel converted the following policies into optimization problems and included them in the code base:

* LAS/Max-min fairness: Least Attained Service policy, used by [Tiresias](/machine-learning-systems/machine-learning-systems-index/tiresias-a-gpu-cluster-manager-for-distributed-deep-learning)
  * LAS w/ weights
* Minimize makespan
* Finish time fairness: Used by [Themis](/machine-learning-systems/machine-learning-systems-index/themis-fair-and-efficient-gpu-cluster-scheduling)
* FIFO
* Shortest job first (SJF)
* Minimize cost (in public cloud instances)
  * Minimize cost w/ SLOs
* Hierarchical (multi-level policy: FIFO, fairness, etc.)

![](/files/-M_Y-wSWF8ZLho6Qtaxv)

### Realizing the optimal allocation: round-based scheduling + priorities

With the optimal allocation computed, Gavel tries to dispatch jobs while matching the optimal allocation as close as possible by using two techniques:

1. **Round-based scheduling**: Gavel allows users to set the round length (optimal length that allows for the effective approximation is 6 minutes). This mechanism ensures that jobs receive time on accelerator types according to the optimal allocation.&#x20;
2. **Per-job priority score**: In each round, the scheduler runs jobs in decreasing priority order. The priority for each job is computed as the target allocation divided by the number of rounds received.

![Round-based scheduling](/files/-M_X5D7MmXBrPiQte5on)

![This priority mechanism helps the approximation of the optimal allocation as shown in X^example.](/files/-M_X3oMWMxU7TCsK1zdk)

## Evaluation

![Heterogeneity-aware policies reduces avg. JCT by 1.5x](/files/-M_Y1pd0x_7ODT_wJoj0)

![How well the Gavel scheduling/dispatching mechanism works](/files/-M_Y3NT9AIS_obhAFJF-)

![](/files/-M_Y3lxRlE0LlDICLMpK)

![](/files/-M_Y43bnVDBCiHGfXo0X)

The paper also covers extensive evaluations on:

* How well the heterogeneity-aware policies improve objective metrics, both in simulation and physical experiments
* How do Gavel's policies scale
* Whether Gavel can accurately estimate the throughputs of co-located jobs when using space sharing

## Links

* [Paper PDF](https://cs.stanford.edu/~matei/papers/2020/osdi_gavel.pdf)
* [Presentation video at OSDI '20](https://www.youtube.com/watch?v=I2PsnUo6WPk)
* [Presentation slides at OSDI '20](https://www.usenix.org/sites/default/files/conference/protected-files/osdi20_slides_narayanan.pdf)
* [Gavel on GitHub](https://github.com/stanford-futuredata/gavel)


# \[2020 OSDI] AntMan: Dynamic Scaling on GPU Clusters for Deep Learning

## Summary

AntMan is a cluster scheduler for GPU sharing. It introduces two techniques, dynamic memory scaling and opportunistic computation management, to accommodate multiple jobs and avoid interference.

![System architecture/workflow](/files/HSi1eNBYQii7jDYZNP0W)

## Background & Motivation

* GPUs in a shared cluster are not properly utilized (both SM and GRAM are under-utilized). One of the reasons is multi-GPU jobs require gang scheduling, which creates GPU idleness. Moreover, DL training jobs have dynamic resource demand over time.&#x20;
* Training jobs in the Alibaba cluster have the following characterstics:
  * Small model size: Most GPU memory can be shared
  * Short mini-batch: Fast resource coordination
  * Similar mini-batch: Mini-batch time can be used to quantify inter-job interference

![](/files/hQXug93N47a01OPjlvhd)

![](/files/RZnu4VI6GJGfpb1pPTi9)

## Design & Implementation

### Dynamic Memory Scaling

AntMan dynamically co-locates jobs on shared GPUs. The goal is for resource-guarantee jobs to maintain the same performance as dedicated execution while co-locating opportunistic jobs to best utilize the resources.

AntMan monitors the memory usage of DL jobs and sets the corresponding memory upper bounds, allowing other jobs to utilize the spare memory. However, since DL jobs have dynamic resource demand, jobs may require more memory than before, which creates OOM and fails all jobs. In this case (Fig. 7a), these memory bursts are cached on the host (CPU) memory, and are moved back to GRAM after re-allocation. The same technique is applied to jobs that need to shrink their memory requirements to make way for other jobs (Fig. 7b).&#x20;

![](/files/H7i9Pg8J4CZp8pHoDqk1)

### Computation management for minimizing interference

The GpuOpManager is introduced in DL frameworks to opportunistically launch computation kernels during idle time slots to reduce interference.

![](/files/l7KK7JdGsdxH7r9tJ8Tb)

## Evaluation

![Micro benchmark 1: Memory scaling](/files/Adly3uubTOAhapgXsHE9)

![Micro benchmark 2: Computation management. Here, ESPnet is a resource-guaranteed job, while ResNet50 is an opportunistic job.](/files/KW0Y6oxAWyHDmKsC0ymB)

![End-to-end evaluation](/files/1iTTmUez5dzdCb5G45r3)

## Links & References

* [Paper PDF](https://www.usenix.org/system/files/osdi20-xiao.pdf)
* [Presentation video at OSDI '20](https://www.youtube.com/watch?v=8PSzcqL0eUA)
* [Presentation slides at OSDI '20](https://www.usenix.org/sites/default/files/conference/protected-files/osdi20_slides_xiao.pdf)
* [GPU-cluster-for-deep-learning on GitHub](https://github.com/alibaba/GPU-scheduler-for-deep-learning)


# \[2020 OSDI] BytePS: A High Performance and Generic Framework for Distributed DNN Training

## One-line Summary

In this paper, the authors introduced BytePS, a unified architecture for accelerating distributed DNN training in heterogeneous GPU/CPU clusters. Yes, this is the OG title of the paper, but it's so long that it does not fit in the GitBook title, so...

## Paper Structure Outline

1. Introduction
2. Background
   1. Distributed DNN Training
   2. All-reduce
   3. Parameter Server (PS)
3. Motivation and BytePS Architecture
   1. Motivation
   2. Architecture Overview
4. BytePS Communication Design
   1. Inter-machine Communication
      1. Communication Efficiency Analysis
   2. Intra-machine Communication
      1. PCIe-only Topology
      2. NVLink-based Topology
      3. Discussion
5. Summation Service
6. Implementation
   1. Multi-Stage Pipeline
   2. Address RDMA Performance Issues
   3. BytePS Usage
7. Evaluation
   1. Inter-machine Microbenchmarks
   2. Leverage CPU Machines
   3. Adapt to Intra-machine Topology
   4. Scalability
8. Observations and Discussion
9. Related Work
10. Conclusion

## Background & Motivation

Existing architectures (all\_reduce and parameter server) for distributed DNN training are insufficient.

![The two architectures for distributed training based on data parallelism: All-reduce and PS](/files/-MNPAWZr-UZFytUrQOkm)

![Even with ByteScheduler, we are still 30% away from the optimal performance](/files/-MNP98gq7-84zdrrdkDa)

The paper analyzed three problems that led to this slowdown, and then presented a solution to each problem:

1. Sub-optimal Inter-machine Communication
2. Sub-optimal Intra-machine Communication
3. The CPU Bottleneck

### Sub-optimal Inter-machine Communication

![](/files/-MNPBZDxN-l2tuXpXDlm)

For allreduce, the CPUs are not leveraged properly as the communication is between GPUs only. For PS, if there are insufficient CPUs for servers, bottlenecks may be created. **The solution is an optimal communication strategy that unifies allreduce and PS.**

### Sub-optimal Intra-machine Communication

![](/files/-MNPC1ieHg3DXzomTT5w)

There are often multiple GPUs in a GPU machine IRL. The internal topology is also a network, which results in bottlenecks. **The solution is intra-machine optimizations that accommodate diverse intra-machine topologies.**

### The CPU Bottleneck

In the parameter server setup, the GPU workers send the gradients to CPU servers. The CPU servers first aggregate the gradients received, and then update the parameters using the optimizer function. The problem is that CPUs might not match network rates, thus creating bottlenecks. **The solution is a summation service that moves parameter updates from CPUs to GPUs.**

## Design and Implementation

### Sub-optimal Inter-machine Communication

PS only uses links between CPUs and GPUs and does not utilize the bandwidth between GPU machines. In the allreduce setup, the communication solely relies on inter-GPU communications, not utilizing the CPU at all. BytePS takes the best of both worlds and combines these two strategies. In the paper, the authors presented an optimal partition strategy that adjusts the proportion of CPU-GPU and GPU-GPU communications.

![](/files/-MNPIGEmcHjppfz-ptoW)

### Sub-optimal Intra-machine Communication

![I might come back to this later, this is a bit complicated to understand for now :P](/files/-MNPJ3czQTbBzNZjRdcT)

### The CPU Bottleneck

The PS server's role can be divided into two parts: Gradient Summation & Parameter Update. Typically, forward propagation and backward propagation get placed on GPUs, while gradient summation and parameter update are placed on GPUs. The authors found that the gradient summation step is CPU-friendly, while the parameter update step is heavy. To resolve this issue, the authors presented the Summation Service, which moves the parameter update to GPUs to resolve the aforementioned bottleneck.

### System Architecture Overview

![](/files/-MNPKFu40VQ6t4Z9tPND)

## Evaluation

![BytePS achieves near-optimal communication performance](/files/-MNPML8xTZeYrUMLd7_7)

![10% - 20% gains](/files/-MNPN-z9w8EJuzSqpuuj)

![BytePS outperforms allreduce & PS by up to 84% and 245%, respectively](/files/-MNPNK1Sj4x-Yj0oOb3l)

![Breakdown of Performance Gains](/files/-MNPNX9wy5pD1kpe0KCX)

## New Vocabulary

* Network Interface Controller (NIC):&#x20;
* PCI Express (PCIe):&#x20;
* Goodput: Throughput that's good. Goodput is the rate at which **useful** data passes through a link, while throughput measures the rate at which **all** data passes through a link. For example, in a local area network (LAN), goodput only measures the throughput of the original data, while throughput also measures all the protocol overhead information (packet headers, etc.).

## Links

* [Paper PDF](https://www.usenix.org/system/files/osdi20-jiang.pdf)
* [Presentation Video at OSDI '20](https://www.youtube.com/watch?v=j8PHNglSZX8\&feature=emb_logo\&ab_channel=USENIX)
* [Presentation Slides PDF](https://www.usenix.org/sites/default/files/conference/protected-files/osdi20_slides_jiang.pdf)
* [BytePS on GitHub](https://github.com/bytedance/byteps)
* [The rationale for BytePS](https://github.com/bytedance/byteps/blob/master/docs/rationale.md)


# \[2020 SIGCOMM] Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video Analytics

## Summary

Frame differencing is an existing technique that improves the efficiency of video analytics pipelines. Reducto makes this approach more efficient by performing on-camera frame filtering using server-guided decisions to dynamically adjust the filtering threshold and smartly select the best differencing feature.

![](/files/Tl8IGYz4HMPl8am0yarQ)

## Background & Motivation

In real-time video analytics pipelines, cameras send video streams to cloud servers, which then immediately run object detection models to answer user queries. Such pipelines aim for high accuracy and low latency.

Video analytics is resource-intensive (compute & network)! A technique to improve efficiency is frame filtering, which removes (i.e., do not send) frames that wouldn't change the query results. There are three existing approaches to this end:

1. Approximate model: Only send frames to DNN if the on-camera compressed model is not confident
   1. Con: Still too slow on camera
2. Specialized binary classifiers: Only send frames that contain certain objects
   1. Con: Misses filtering opportunities (objects are present but query results do not change across frames)
3. Pixel-level frame differencing: Only send frames if the low-level features like pixel values have changed drastically as we would expect a different result from the previous frames
   1. Cons: Existing approaches use single, static thresholds while the video content can be highly dynamic, so they cannot reliably meet accuracy targets. Moreover, they rely solely on pixel comparison, whereas there might be other low-level frame differences that are potentially more effective.

![](/files/BIuE3De6ID227LPyQ1sd)

This work focuses on improving (3) by addressing two questions: (1) How do we dynamically determine the filtering threshold? (2) Which differencing feature should we use?

## Design & Implementation

### Dynamic threshold

![](/files/jeM5wtBFCjg54jvtrl72)

* What's the threshold that filters the most frames while meeting the target accuracy?
* Using a small, unfiltered video, split the video into several segments and construct a hash table that maps the diff values to thresholds
* Server: builds the hash table (expensive)
* Camera: looks up in the hash table (cheap)

### Choosing the differencing feature

![](/files/jC0vn6wAaSkmr8cipSlt)

* Server: calculates the best feature (expensive)
* Run once per query type (e.g. counting)

### Overall pipeline

![](/files/tKFtWJyvXa5xgho8Hl2G)

![](/files/ETe4TypMzF1xPZpTvQNn)

## Evaluation

![](/files/lbOicqKfSN9M0s0kUpGI)

![](/files/SnIjR2bmj4ifzJGRX9i1)

![](/files/YqEqgHYSr1AHKwqyl0Md)

## Links & References

* [Paper PDF](https://www.cs.princeton.edu/~ravian/publications/reducto_sigcomm20.pdf)
* [Presentation video at SIGCOMM '20](https://www.youtube.com/watch?v=IllEKLVUiYM)


# \[2020 MLSys] Salus: Fine-Grained GPU Sharing Primitives for Deep Learning Applications

## One-line Summary

Salus presents two GPU sharing primitives for fine-grained GPU sharing among multiple DL applications: fast job switching (for time-sharing and preemption) and GPU lane abstraction (for dynamic memory sharing). Together, these primitives improve all aspects of training performance and open the gate to implementing novel policies.

## Paper Structure Outline

1. Introduction
2. Background and Motivation
   1. DL Workloads Characteristics
   2. Existing Techniques for Sharing GPUs
3. Salus
   1. Architectural Overview
   2. Efficient Job Switching
      1. Characterizing DL Memory Allocations
      2. Scheduling Granularity
   3. Spatial Sharing via GPU Lane
      1. Lane Auto Defragmentation
      2. Lane Assignment
4. Scheduling Policies in Salus
   1. PACK to Maximize Efficiency
   2. SRTF to Enable Prioritization
   3. FAIR to Equalize Job Progress
5. Evaluation
   1. Long-Running Training
      1. Overall Comparison
      2. Impact of Fast Job Switching
   2. Hyper-Parameter Exploration
   3. Inference
   4. Overhead
6. Concluding Remarks

## Background & Motivation

The modern GPU allocation model (coarse-grained, one-at-a-time) is not good enough, in that it creates a high overhead for flexible scheduling (time-sharing, preemption, migration). Also, as having multiple coexisting processes in a GPU is not efficient, the GPU becomes underutilized.

## Design and Implementation

Existing GPU sharing techniques are not efficient. Salus tackles this by exposing two GPU sharing primitives: fast job switching and memory sharing.&#x20;

![The adaptor transfers computation graph, while the execution service consolidates all GPU accesses](/files/-MRuQTW7Jo7JvZen9IOF)

### Fast Job Switching

Modern frameworks use checkpointing for job switching to achieve second-scale suspend-resume. Nevertheless, checkpointing results in large data transfer from/to the GPU memory, making the communication cost non-negligible. In Salus, memory allocations are classified into three types:

1. Model: Model parameters, persistent, no temporal variations.
2. Ephemeral: Intermediate layer's outputs/temporal data generated by the algorithm itself. Only needed during computations and are released between iterations.
3. Framework-internal: Used by the DL framework for book-keeping/data preparation pipeline. Persistent across iterations.

![](/files/-MRuRqKPoT8COp9o3xR3)

As persistent memory is significantly less than ephemeral memory, more than one job's persistent memory can be kept in GPU while still having space for either job's ephemeral memory. Thus, fast job switching is enabled by not removing persistent memory from GPU at all.

### Spatial Sharing via GPU Lane

![](/files/-MRvEv_C-MYBBvwn5Bzr)

![](/files/-MRvEouN_kerZPLB4HCD)

In Salus, the GPU memory space is divided into ephemeral and persistent regions, and the ephemeral region is furthur divided into lanes, which are continuous memory spaces that can contain ephemeral memory allocation for iterations. This allows time-slicing within lanes and parallelism across lanes. Salus also supports automatic in-lane defragmentation and dynamic lane assignment (re-partitioning).

### Scheduling Policies

1. PACK (Maximize Efficiency): Multiple jobs can be packed in separated GPU lanes to achieve higher utilization and minimize makespan. This allows training many jobs in parallel and enables efficient inference serving.
2. SRTF (Enable Prioritization): Salus supports job priorities by doing preemption of larger jobs.
3. FAIR (Equalize Job Progress): This equalizes total service over time for jobs in each lane.
4. FIFO: The de facto mechanism nowadays.

## Evaluation

![Salus introduces little overhead (observed from similar makespan). This also confirms that SRTF improves the avg JCT compared to FIFO.x](/files/-MRvJaiAjm3ECFwZ8uyS)

![Sub-second Level Switching](/files/-MRvJ2PsJ_f85WYoDSXI)

!["TF": FIFO in TensorFlow.](/files/-MRvKRkUu6h6R6mI7pm3)

![Packing inference applications allows Salus to reduce the number of GPUs needed while maintaining reasonable latency overhead.](/files/-MRvKeiRnGPHpm87htcQ)

## New Vocabulary

*

## Links

* [Paper PDF](https://www.mosharaf.com/wp-content/uploads/salus-mlsys20.pdf)
* [Presentation slides at MLSys '20](https://mlsys.org/media/Slides/mlsys/2020/balla\(03-10-30\)-03-10-30-1426-fine-grained_gp.pdf)
* [Poster](https://unlimitedcodeworks.xyz/assets/pub/yu20mlsys/yu20mlsys-poster.pdf)
* [Salus on GitHub](https://github.com/SymbioticLab/Salus)


# \[2020 EuroSys] AlloX: Compute Allocation in Hybrid Clusters

## One-line Summary

Allocate interchangeable resources in a hybrid cluster (CPU, GPU, FPGA, problem-specific accelerator);

Scheduling -> Min-cost bipartite matching problem, provides dynamic fair allocation

reduce avg jct while providing fairness and preventing starvation

## Paper Structure Outline

1. Introduction
2. Background & Motivation
   1. Interchangeable Resources
3. Algorithm Design
   1. Optimal Approach for Queued Up Jobs
      1. Generate input for the matching problem
      2. Solve the matching problem
      3. Convert the matching solution to job scheduling
   2. Handling Online Arrivals
   3. Incorporating Fairness
      1. Existing Fair Allocation Algorithms are Insufficient
      2. Our Idea
      3. Incorporating Fairness into AlloX
4. AlloX Implementation
   1. Estimator
   2. Scheduler
   3. Placer
   4. Operational Issues
5. Evaluation
   1. Experimental Methodology
      1. Baselines
   2. AlloX Performance
      1. Experiments on a Cluster
      2. Simulation Results
      3. Starvation
      4. Performance and Fairness Trade-offs
   3. Sensitivity Analysis
      1. Estimation Errors
      2. Profiling Overhead
6. Related Work
7. Concluding Remarks
8. Acknowledgments

## Background & Motivation

### Background X: Heterogeneous/interchangeable resources

CPU, GPU, distinct speedup rates for different applications

### Motivation X: Current schedulers do not consider&#x20;

best fit: pick the optimal config for each job. This creates load imbalance (heavy load on GPUs while CPUs idle)

join the shortest queue: choose resource with the shortest completion time. This is short-sighted (each job optimizes for itself w/o considering later jobs)

shortest job first: maintain ordered queue of all jobs by increasing processing time. Whenever a resource becomes avaialable, scheulde job wiht the shortest run time

## Design and Implementation

## Evaluation

## Links

* [Paper PDF](https://www.mosharaf.com/wp-content/uploads/allox-eurosys20.pdf)


# \[2020 VLDB] PyTorch Distributed: Experiences on Accelerating Data Parallel Training

## One-line Summary

This paper presents the design, implementation, and evaluation of the PyTorch distributed data parallel module.

## Paper Structure Outline

1. INTRODUCTION
2. BACKGROUND
   1. PyTorch
   2. Data Parallelism
   3. AllReduce
3. SYSTEM DESIGN
   1. API
   2. Gradient Reduction
      1. A Naive Solution
      2. Gradient Bucketing
      3. Overlap Computation with Communication
      4. Gradient Accumulation
   3. Collective Communication
4. IMPLEMENTATION
   1. Python Front-end
   2. Core Gradient Reduction
5. EVALUATION
   1. Latency Breakdown
   2. Bucket Size
   3. Scalability
   4. Round-Robin Process Group
6. DISCUSSION
   1. Lessons Learned
   2. Future Improvements
      1. Gradient Order Prediction
      2. Layer Dropping
      3. Gradient Compression
7. RELATED WORK
8. CONCLUSION

## Background & Motivation

There are three steps in training a DNN model:

1. Forward pass: Computes loss
2. Backward pass: Computes gradients
3. Optimizer step: Updates parameters

To train large models on large datasets, data parallelism is applied so that multiple workers work together to do the training. Each worker holds a replica of a model, trains the model (forward & backward pass) using a partition of the dataset, and averages the gradients/parameters among the workers.

## System Design

### API

There are two design goals when designing the API:

1. Non-instrusive: Converting local training scripts to distributed scripts should require minimal code modifications.
2. Interceptive: For as many optimizations as possible to work, the API needs to allow the implementation to intercept various signals and trigger appropriate algorithms correctly.

### Gradient Reduction

1. **Naive solution**: DDP controls all training processes to (1) start from the same model state and (2) consume the same gradients in every iteration. (2) can be implemented by inserting a gradient synchronization phase after the local backward pass, or by adding a hook to trigger computation after every backward pass. There are two performance concerns:
   1. Collective communication performs poorly on small tensors
   2. By separating the gradient computation and synchronization, we lose the chance to overlap the two phases
2. **Gradient bucketing**: We can observe that collective communications are more effective on large tensors than on smaller tensors. As a result, we can use gradient reduction to bucket multiple gradients into one allreduce operation. However, DDP should not compact all gradients in one single allreduce, otherwise the communication cannot overlap with computation.
3. **Overlap computation with communication**: With bucketing, DDP only needs to wait until all contents in the same bucket is ready before launching communications. There are two things that requires attention, though. The first thing is that the reducing order must be the same across all processes, otherwise mismatches might occur. The other thing is that the backward pass could hang due to some gradients being skipped and never saying "I'm ready" to their corresponding buckets. See the graph below for two examples.
4. **Gradient accumulation**: Instead of doing allreduce for every iteration, do allreduce every n interations.

The paper covered some level of technical details for each of the above four subsections.

![Two cases of failures during gradient synchronization due to overlapping](/files/-MOt5G9Obg0i9Zl7U5WG)

![Overall architecture](/files/-MOtAfbMpGZy1dWe3dd8)

### Collective Communication

DDP is built on top of communication libraries like NCCL, Gloo, and MPI. The APIs from all three libraries are wrapped into the same ProcessGroup API. In DDP, workers are expected to join a process group for commuication primitives to work on.

### Implementation Details

1. Python Front-end
   1. Configurable knobs
   2. Model device affinity
   3. Model buffers
2. Core Gradient Reduction
   1. Parameter-to-bucket mapping
   2. Autograd hook
   3. Bucket allreduce
   4. Globally unused parameters

## Evaluation

![The effectiveness of overlapping computation with communication](/files/-MOxTpLSjpwbeLwyxRIb)

![Bucket size vs. latency](/files/-MOyUFjvpNdtyBNqPyil)

![Latency vs. number of GPUs](/files/-MOyUsfb-5LwXzTKfF_o)

![Latency vs. doing gradient reduction every n iterations (n = 1, 2, 4, 8)](/files/-MOyVEq_Iw44Bu_SVsYX)

![Besides the per iteration latency, it’s also crucial to measure the convergence speed to ver- ify if the acceleration might be erased by convergence slow- down.](/files/-MOyYWV6xcFZSKIebTLe)

![Using multiple process groups to bypass the intrinsic concurrency limitations in process group backend implementations.](/files/-MOyZ4V2jMfOGB2MU6YP)

## Discussion

There is no single configuration that would work for all use cases, but there are some rules that can be summarized and can help us narrow down the range of the optimal configuration:

1. **Communication backend**: In most cases, NCCL is considerablly faster than Gloo.
2. **Bucket size**: The optimal bucket size lies in between the small extreme and the large extreme. The optimal bucket sizes are likely to increase with the size of the model in a sub-linear manner.
3. **Resource allocation**: In NCCL, it is recommended to keep the workers in a DDP group within the same machine, otherwise there will be significant slowdown due to the bandwidth across the machines being lower than that between same-machine GPUs.

Some future directions for optimizations:

1. **Gradient order prediction**: trace the backward order using autograd hooks and update parameter to bucket mapping accordingly
2. **Layer dropping**: drop layers randomly during the forward pass
3. **Gradient compression**: only communicates gradients with the necessary precision

## New Vocabulary

* Hooks in PyTorch: A hook can be registered on a Tensor or a nn.Module. A hook is basically a function that is executed when either forward or backward is called.
* [Computation graph in PyTorch](https://jdhao.github.io/2017/11/12/pytorch-computation-graph/)

## Links

* [Paper PDF](https://arxiv.org/pdf/2006.15704.pdf)
* [CS 744 Slides & Notes](http://pages.cs.wisc.edu/~shivaram/cs744-fa20-slides/cs744-pytorch-notes.pdf)
* [Debugging and Visualization in PyTorch using Hooks](https://blog.paperspace.com/pytorch-hooks-gradient-clipping-debugging/)


# \[2020 NetAI] Is Network the Bottleneck of Distributed Training?

## Summary

Distributed training suffers from sub-linear scale-out. The authors argue that this is due to the network not being fully saturated as a result of the poor implementation of network transport. If the network can be fully utilized, distributed training can achieve an almost-linear scale-out. Also, in a highly-utilized network, the extent of gradient compression does not need to be that high.

## Background & Motivation

Current distributed DNN training using data parallelism suffers from a sub-linear scaling when scaled out. People have been optimizing the communication phase of distributed training, as the computation phase (the other phase) is embarrassingly parallel and should scale almost linearly. An example of those optimizations is gradient compression, which lies at the application level.

![](/files/ENGgJ8X2DWql5i5fs0Ci)

The authors of this paper instead look at the network layer (network-level optimizations do not require changes at the application level).

## Evaluation

The authors first argue that the computation phase is not the bottleneck. They found that due to (1) distributed backward pass overlaps with all-reduce and (2) Horovod injects per-layer hooks in distributed training, the computation time has a slight increase in distributed training. However, this inevitable side effect (at most 15%) does not offset the extent (merely 56% - 75%) of the sub-linear scaling.

![](/files/QWvb4c6r9V0wQSvwdGAh)

Then, the only possibility is that the communication phase is the bottleneck, so the authors tried different network bandwidths. Surprisingly, the scaling factor line plateaus after 25Gbps, meaning that a faster network does not necessarily benefit the system.

![](/files/vOYVIK75fIYOMGBElCal)

![](/files/YX5F03n066wIRIPExzlh)The low network utilization at a high bandwidth explains the issue, as only a small fraction of the bandwidth is properly utilized. This might be explained by TCP being CPU-intensive at high speed (100 Gbps), but modern GPU instances have sufficient CPUs, and the authors found the actual CPU utilization is low. **The conclusion is that the poor implementation of network transport cannot fully saturate the available bandwidth during communication.**

![](/files/kUa2jh97n02b5GQFH1mB)A what-if simulation analysis shows that under a fully-utilized network, almost-linear scale-out is possible.

![](/files/AFOT2vAteI0KZLDCYZsQ)Also, gradient compression is useful in low-speed networks, but a large compression ratio is not necessary.

## Links & References

* [Paper PDF](https://dl.acm.org/doi/pdf/10.1145/3405671.3405810)
* [training-bottleneck on GitHub](https://github.com/netx-repo/training-bottleneck)


# \[2020 NSDI] Themis: Fair and Efficient GPU Cluster Scheduling

## One-line Summary

The authors present a new fairness metric, finish-time fairness, and a newer scheduler architecture and API that supports resource division according to the metric.

## Paper Structure Outline

1. Introduction
2. Motivation
   1. Preliminaries
   2. Characterizing Production ML Apps
   3. Our Goal
3. Finish-Time Fair Allocation
   1. Fair Sharing Concerns for ML Apps
      1. ML Task Durations
      2. Placement Preferences
   2. Metric: Finish-Time Fairness
   3. Mechanism: Partial Allocation Auctions
      1. One-Shot Auction
      2. Multi-round auctions
4. System Design
   1. Design Requirements
   2. THEMIS Scheduler Architecture
      1. Need for a new scheduling architecture
      2. Two-Level Semi-Optimistic Scheduling
   3. AGENT and AppScheduler Interaction
      1. Single-Job ML Apps
      2. Generalizing to Multiple-Job ML Apps
      3. End-to-end Example
5. Implementation
6. Evaluation
   1. Experimental Setup
   2. Macrobenchmarks
      1. Sources of Improvement
      2. Effect of Contention
      3. Systems Overheads
   3. Microbenchmarks
   4. Sensitivity Analysis
7. Related Work
8. Conclusion

## Background & Motivation

![An analysis of existing workloads](/files/-MOSflM_29LmJQSNphml)

In large GPU clusters, existing scheduling disciplines do a poor job in fair sharing.&#x20;

The authors presented the Sharing Incentive (SI): "If N deep learning apps are sharing a cluster, then no application should run slower than on a private cluster with 1/N resources". Similar fairness metrics include Pareto Efficiency (PE) and Envy-Freeness (EF).

Existing cluster scheduling frameworks are not adequate:

1. Dominant Resource Fairness (DRF): The metric is "application resource share". It only uses instantaneous resource fairness (whenever resources are available, they are allocated to the task with the least current share). This is fine for big data analytics workloads, where task durations are short. For ML apps, though, running long, resource-intensive, gang-scheduled tasks might lead to newly-arrived jobs waiting, which is a violation of SI. Also, DRF does not take into account the placement preferences of ML apps. Modern ML apps have vastly different model architectures and placement preferences. For example, VGG16 is affected greatly by the hardware placement, while Inception-v3 is not. This is due to the difference in the amount of communication and synchronization for different workloads.
2. Least Attained Service (LAS, Tiresias uses this): The metric is "app attained service, #GPUs \* time". In Tiresias, GPUs are leased for a certain duration, and when leases expire, available GPUs are given to the job that received the least GPU time thus far. While this resolves the starvation issue mentioned above, it fails to address the placement issue: For two (sparse vs. dense) placements, Tiresias considers them to be the same as the attained service is equal, while actually, a poor placement leads to a slower execution time.

![Different placement preferences of jobs](/files/-MOSf9AtYl6iGLGWTSbn)

![A violation of the SI due to the placement preferences being ignored](/files/-MOSegGSO2hYEo2NwEAX)

## Design and Implementation

### Metric

The authors presented the new metric finish-time fairness, $$\rho$$:$$\rho = T\_{sh} / T\_{id}$$

* $$T\_{sh}$$: finish-time of app in shared cluster
* $$T\_{id}$$: finish-time of app in exclusive 1/N share of cluster
* $$N$$: average contention during app lifetime

The SI requires that for every application, $$\rho \leq 1$$.

![Calculating \rho for single-job ML apps](/files/-MOSnLMr_3jBDlmB_ST8)

### Interface

Hyperparameter Optimizers (Hyperparam-Opt, like Google Vizier) manage deep learning applications. The Hyperparam-Opt tracks per-job progress and determines which jobs to terminate early. Applications calculate $$\rho$$, and the scheduler pulls updated values of rho from the Agent co-located with the app's Hyperparam-Opt. For the CS 736 final, this is as deep as it will cover. In the future, I'll do a second read to try to dig deeper.

### Mechanism

SI's focus is minimizing the max rho: min(max(rho))

![Strawman mechanism](/files/-MOSq62lKDi7zp9vUfjf)

Strawman Mechanism: When resources are available, the interface gets rho estimates from all apps, and then allocates the resources to the app with the highest rho for *lease* duration. There are two drawbacks (compared with Themis):

1. May not find the most efficient allocation: It indeed allocates resources to the job that needs resources the most, but the job might not be the one that can utilize the resource the best.
2. Cannot handle applications that lie about high rho values: Applications have the incentive to lie with high rho values to hoard GPU resources, which leads to starvations of honest jobs.

![The Themis approach](/files/-MOSrxhwFNaYXojMvoIl)

Themis introduces a knob, f, that manages a tradeoff between SI and efficiency.

1. Solving inefficient allocation: When f = 0, more applications are allocated resources and, as a result, get better opportunities to match placement preferences of apps to resources. Analysis suggests f = 0.8 gives a good tradeoff.
2. Incentivizing truth-telling of rho: Partial Allocation Auction within 1-f apps.

![Partial Allocation Auction: Too detailed for CS 736, will cover in 2nd pass](/files/-MOStVGrRKXUoUs49NN1)

## Evaluation

Baseline frameworks for comparison:

* Tiresias: Least Attained Service Job First
* Optimus: Best Throughput Scaling First
* Gandiva: Best Packing Job First
* SLAQ: Best Loss Gradient Job First

![](/files/-MOSuPHMOzPt640p--0O)

From Figure 9, we can observe a few things:

1. Some apps perform very poorly with Tiresias due to placement inefficiency
2. Themis is fair (rho <= 1.2)
3. Themis is minimizing the max rho compared to other algorithms

![](/files/-MOSuti162U8cnUPkZbE)

From Figure 19, we can observe:

1. In general, increasing f (filtering out more jobs) increases fairness and decreases max rho
2. Increasing f gives us fewer scheduling choices, so it decreases GPU performance
3. Short lease -> switch more rapidly between different jobs -> fairer

## New Vocabulary

* Hyperparameters: Data that govern the training process itself. Hyperparameters are set before the learning process begins. For example, the number of hidden layers, the number of nodes each layer should use, etc.&#x20;
* Hyperparameter tuning: The process of finding the best values of the hyperparameters.
* [Preemptive and Non-Preemptive Scheduling](https://www.geeksforgeeks.org/preemptive-and-non-preemptive-scheduling/)

## Links

* [Paper PDF](https://www.usenix.org/system/files/nsdi20-paper-mahajan.pdf)
* [Presentation Video at NSDI '20](https://www.youtube.com/watch?v=K2a7DRcZdIU\&ab_channel=USENIX)
* [Presentation Slides](https://www.usenix.org/sites/default/files/conference/protected-files/nsdi20_slides_mahajan.pdf)


# \[2021 MLSys] Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification

## One-line Summary

Accordion dynamically adjusts the gradient compression rate and batch size during critical regimes in training to do better compression, reduce communication, and achieve an end-to-end speedup w/o losing accuracy.

## Paper Structure Outline

1. Introduction
2. Related work
3. Distributed SGD
4. ACCORDION
   1. Adaptive communication using critical regimes
   2. ACCORDION's Design
   3. Relationship between gradient compression and adaptive batch-size
5. Experimental evaluation
   1. Experimental setup
   2. Results
   3. ACCORDION with PowerSGD
   4. ACCORDION with TopK
   5. ACCORDION with Large Batch size
   6. Comparison with Prior Work
6. Future Work and Limitations
7. Conclusion
8. Appendix
   1. Detailed Experimental Settings
   2. Connection Between Gradient Compression and Batch Size
   3. ACCORDION on Extremely Large Batch Size
   4. Results and Detailed Analysis
      1. Language Model
      2. Computer Vision Models
   5. Detailed Analysis of Batch Size Results
   6. Compression Ratio Selection of Adasparse
   7. Model Descriptions

## Background & Motivation

Current methods to alleviate gradient communication bottlenecks include:

* Lossy gradient compression (reduce the size of data communicated)
  * Choosing the compression ratio is a tradeoff between final accuracy & communication overhead)
  * Can be generalized into three groups: quantization, sparsification, and low rank approximation
* Increase batch size (reduce the frequency of per-epoch communication)
  * This leads to degradation in final accuracy

In this work, the authors relax the "fixed communication" scheme and use adaptive schemes. The authors build on the idea of critical regimes so that avoiding gradient compression (lowering the compression rate) during critical regimes mitigates accuracy loss. Accordion is also able to adjust the batch size.

## Design and Implementation

![](/files/-MT9OT8iPx_1eQVuJoBE)

In the example above, if low compression is used for the first 20 epochs and the 10 epochs after epoch 150 and high compression is used in other places, the overall communication will be close to high compression, and the accuracy will be the same as using low compression throughout (communication is also reduced significantly).

![Whatever that symbol is, it is the threshold for declaring critical regimes.](/files/-MT9PWi6pzqmqqaJmXpB)

Critical regimes are identified by measuring the rate of change in gradient norms. This technique has a low computational and memory overhead.

For batch sizes, small batches are used only in critical regimes, and this results in performance similar to using small batches everywhere.

## Evaluation

![Accordion with PowerSGD](/files/-MT9QQ-6NAsB4iWb43ev)

![Accordion with TopK](/files/-MT9QWKH7TVt6SFkkMv2)

![Accordion with large batch size](/files/-MT9QZwV3FRNXMljfkpm)

More evaluations are available in the paper appendix. This paper has the longest appendix I've ever seen :)

## Links

* [Paper PDF](https://proceedings.mlsys.org/paper/2021/file/1d7f7abc18fcb43975065399b0d1e48e-Paper.pdf)
* [Presentation video (Poster) at MLSys '21](https://slideslive.com/38952703/accordion-adaptive-gradient-communication-via-critical-learning-regime-identification?ref=account-folder-80869-folders)
* [Presentation video (Oral) at MLSys '21](https://slideslive.com/38952732/oral-accordion-adaptive-gradient-communication-via-critical-learning-regime-identification?ref=account-folder-80870-folders)
* Presentation slides at MLSys '21
* [Accordion on GitHub](https://github.com/uw-mad-dash/Accordion)


# \[2021 VLDB] Analyzing and Mitigating Data Stalls in DNN Training

## One-line Summary

This work investigates the data pipeline aspect of DNN training. The authors found that DNN training is dominated by prefetching and preprocessing of data (data stalls). A tool for measuring data stalls is built, and techniques are presented to mitigate data stalls.

## Paper Structure Outline

1. Introduction
   1. Contributions
2. Background
   1. The DNN ETL Requirements
   2. DALI: Fast Data Pipelining
3. Data Stalls in DNN Training
4. Analyzing Data Stalls
   1. Methodology
   2. Measuring Data Stalls using DS-Analyzer
   3. Data Stalls in DNN Training
      1. When dataset resides on remote storage
      2. When datasets cannot be fully cached
      3. When datasets fit in memory
      4. Data stalls exist across training frameworks
      5. Analysis of NLP models
5. DS-Analyzer: Predictive Analysis
   1. Example: Predicting Optimal Cache Size
6. Mitigating Data Stalls
   1. The MinIO Cache
   2. Partitioned MinIO Caching
   3. Coordinated Prep
   4. Tying it all together with CoorDL
7. Evaluation
   1. Single-Server Multi-GPU Training
   2. Multi-Server Distributed Training
   3. Hyperparameter Search
   4. Training to Accuracy with CoorDL
   5. Resource Utilization
   6. CoorDL on DGX-2
8. Discussion

## Background & Motivation

There are 5 stages in each iteration of an epoch:

1. A minibatch of data items is fetched from storage.&#x20;
2. The data items are pre-processed, for e.g., for image classification, data items are decompressed, and then randomly cropped, resized, and flipped.
3. The minibatch is then processed at the GPU to obtain the model’s prediction.
4. A loss function is used to determine how much the prediction deviates from the right answer.&#x20;
5. Model weights are updated using computed gradients.

Steps 3-5 constitute the actual computation, while steps 1-2 are data preparation. If the computation rate is bigger than the minimum of data prefetching rate and data preprocessing rate, a GPU waits for steps 1-2 to happen (a data stall occurs). More specifically, step 1 is termed a fetch stall, (I/O bound during loading minibatch from storage), while step 2 is termed a prep stall (CPU bound waiting for the data items to be processed). The following conclusions are drawn from the analysis:

![](/files/-MaKVFRPeh19IoOLsHHY)

* When dataset resides on remote storage (distributed fs/object stores): Large datasets usually fit entirely on local storage. Thus, a one-time download cost of the dataset is paid, and the benefits of local SSD is taken advantage of afterwards.
* When datasets cannot be fully cached
  * Fetch stalls are common if the dataset is not fully cached in memory
  * OS page cache is inefficient for DNN training
  * Lack of coordination among caches leads to redundant I/O in distributed training
  * Lack of coordination in HP search results in redundant I/O
* When dataset fits in memory
  * DNNs need 3-24 CPU cores per GPU for pre-processing
  * DALI is able to reduce, but not eliminate prep stalls
  * Decoding is very expensive, offloading prep to the GPU trades GPU memory usage for a speedup
  * Larger batch sizes utilize the GPU parallelism better. However, as compute gets faster, data stalls  become the bottleneck
  * Redundant pre-processing in HP search results in high prep stalls
* Data stalls exist across training frameworks (TensorFlow, MxNet)

## Design and Implementation

### DS-Analyzer: Perform predictive what-if analysis of data stalls

DS-Analyzer analyzes the implication of CPU, memory, and storage on the performance of a DNN and does what-if analyses. DS-Analyzer uses a differential approach and runs in three phases to measure prep stall and fetch stall:

1. Measure vanilla ingestion rate (with no fetch or prep stalls).
2. Measure prep stalls by training with a subset of the dataset entirely cached in memory. With this, any throughput decrease compared to (1) is due to prep stalls.
3. Measure fetch stalls by clearing all caches and setting max cache size to a user-given limit. The difference between (2) and (3) is due to fetch stalls.

### CoorDL: Mitigating data stalls

CoorDL is built on top of DALI and it incorporates the three following techniques:

#### MinIO: DNN-aware software caching to reduce cache misses per epoch (benefits single-server training)

Currently, the caching of the training dataset relies on the OS page cache. DNN training has the data access pattern of "repetitive across epochs and random within an epoch". This means that all data items in the dataset have equal probabilities of access in an epoch, so it is not important which data item is cached, but instead, it is crucial that cached items are not replaced before they are used in order to minimize I/O per epoch.

In MinIO, items, once cached, are never replaced in the DNN cache. This technique results in only capacity misses while using the page cache + LRU leads to more misses because of thrashing. MinIO is implemented in user space instead of a policy in the kernel.

![Toy example: dataset size = 4, cache size = 2. MinIO incurs 2 capacity misses, while page cache result in 2-4 misses because of thrashing.](/files/-MaKQ3El4nck3eyq9Zvq)

#### Partitioned caching to coordinate remote MinIO caches (benefits distributed training)

In distributed training, the dataset is partitioned across all servers with a random partition every epoch. The authors found that on a cache miss, data transfer over commodity TCP stack is much faster than fetching from local storage. As a result, the authors present Partitioned MinIO. At the end of the first epoch, Partitioned MinIO collectively caches a part of the dataset of size equal to the sum of capacities of individual MinIO caches. Metadata about data items present in each server's cache is maintained. In the case of a local cache miss, the item is looked up in the metadata. If present, it is fetched from the respective server over TCP (otherwise from local storage).

#### Coordinated prep to eliminate redundant fetch & prep across jobs (benefits hyperparameter search)

When colocating hyperparameter search jobs, pre-processed minibatches created by one job can be reused by all other jobs. However, there is currently no coordination in data fetch & prep among these jobs, leading to stalls. In coordinated prep, each job receives a random shard of the dataset and processes it. After pre-processing the minibatches, they are exposed to all other jobs in a staging area in the memory region (with minimal memory overhead). Coordinated prep ensures that a minibatch is deleted once it is used exactly once by all jobs to ensure that it's not used across epochs (leads to lower accuracy & OOM).

## Evaluation

![](/files/-MaLrGXfi0Dhw_C8cLUQ)

![](/files/-MaLrImlLWPEqQw4ZKnj)

## Links

* [Paper PDF](https://www.cs.utexas.edu/~vijay/papers/vldb21-datastalls.pdf)
* [msr-fiddle/DS-Analyzer on GitHub](https://github.com/msr-fiddle/DS-Analyzer)
* [msr-fiddle/CoorDL on GitHub](https://github.com/msr-fiddle/CoorDL)


# \[2021 FAST] CheckFreq: Frequent, Fine-Grained DNN Checkpointing

## One-line Summary

CheckFreq pipelines checkpointing with computation for automated, frequent, fine-grained checkpointing in DNN training.

## Paper Structure Outline

1. Introduction
2. Background
3. The Current State of Checkpointing
   1. Checkpointing is Incorrect
   2. Checkpointing is Inefficient
   3. Summary
4. CheckFreq: Design and Implementation
   1. Goals
   2. CheckFreq Recovery Guarantees
   3. Design
      1. Checkpointing Mechanism
      2. Checkpointing Policy
   4. Implementation
5. Evaluation
   1. Experimental Setup
   2. Accuracy Implications
   3. Performance of Checkpointing Mechanism
      1. Checkpoint Stalls
      2. Breakdown of Benefits
   4. Checkpointing Policy
   5. Recovery Time
   6. End-to-End Training
6. Discussion
7. Related Work
8. Conclusion

## Background & Motivation

During DNN training, checkpointing is performed to ensure fault tolerance. Current checkpointing schemes are synchronous, thus leading to large checkpoint stalls. Furthermore, due to bigger models and larger datasets, epoch times are increasing. Typically, checkpointing is performed at epoch boundaries and the checkpointing frequency needs to be set manually. → We need fine-grained, iteration-level checkpointing.

## Design and Implementation

CheckFreq is an automated checkpointing framework for DNN training.

### Mechanism: Low-cost, pipelined checkpointing

#### Low checkpoint stalls: 2-phase DNN-aware checkpointing

CheckFreq decouples the traditional checkpointing into two phases: `snapshot()` and `persist()`. `snapshot()` serializes the training state and copies it from the GPU memory to a user-space buffer in CPU memory. `persist()` writes the serialized content to disk. These two phases are pipelined with DNN training computation.

![](/files/-MaFSKqRwPTEa7JcwOWE)

In the optimal case, as the model weights are synchronized in the last phase of an iteration, we can pipeline the `snapshot()` with the forward & backward pass of the next iteration, minimizing the checkpoint stall.

The authors also found that doing the snapshot on the GPU has an orders-of-magnitude lower cost than that on the CPU, as the latter involves a memory copy from GPU to CPU. Therefore, if spare GPU memories are available, the snapshot is done on the GPU memory.

#### Maintain data invariant: Resumable data iterator

Current data iterators do not guarantee the order of data items after resuming. CheckFreq resolves this by using a seed that is a function of the epoch number to reconstruct the shuffle order after resuming.

![](/files/-MaFXT5tlg7UUDIsFhOu)

### Policy: When to checkpoint?

#### Initial frequency: Systematic online profiling

The key idea is to come up with a frequency of checkpointing every k iterations such that:

1. The cost of 1 checkpoint can be amortized over k iterations
2. The runtime overhead of checkpointing is within a user-defined threshold of the actual compute time (say 5%)

To accomplish this, CheckFreq profiles: the iteration time (Ti), time to perform weight update (Tw), time to create an in-memory GPU copy (Tg), time to create an in-memory CPU copy (Tc), time to write to storage (Ts), size of checkpoint (m), peak GPU memory utilization (M), and total GPU memory (Mmax). Then, the frequency is determined as follows:

![](/files/-MaFV4-qFjRchclulqKE)

#### Adaptive rate tuning: Manages interference from other jobs

Consider the following example

![](/files/-MaFZoLanIFKPJKbJjSq)

* Isolated: When a job runs alone, the checkpointing overhead is kept at 5% as specified by the user
* Static: When another job space-shares the same GPU, checkpointing at the previous frequency results in a 35% overhead
* Adaptive: CheckFreq's adaptive policy reduced the checkpoint frequency and keeps the overhead at 5%

## Evaluation

![](/files/-MaF_KefZddKuh_haFmC)

![](/files/-MaF_XdtscncM_AVmNBb)

## Links

* [Paper PDF](https://www.usenix.org/system/files/fast21-mohan.pdf)
* [Presentation video at FAST '21](https://www.youtube.com/watch?v=E3uaeaqfjcY)
* [Presentation slides at FAST '21](https://www.usenix.org/sites/default/files/conference/protected-files/fast21_slides_mohan.pdf)
* [msr-fiddle/CheckFreq on GitHub](https://github.com/msr-fiddle/CheckFreq)


# \[2021 EuroMLSys] Interference-Aware Scheduling for Inference Serving

## Summary

This work proposes a scheduler for inference workloads on heterogeneous hardware. The scheduler is aware of and proactive to interferences between co-located jobs, therefore outperforming baseline policies like lease-loaded.

## Background & Motivation

Inference serving schedulers co-locate models to improve resource utilization. However, the least-loaded scheduling policy, popular in the context of VM task scheduling, is agnostic to the interference/latency degradation created by co-location, thus yielding sub-optimal scheduling result.&#x20;

## Design & Implementation

![](/files/qZm0sHtOTwm2koAlUhak)

![](/files/gEG7V5ZkDDY5uX5vPnVz)

By using a unified predictor instead of maintaining separate predictors for different co-location degrees and machine types, we are able to (1) reduce the efforts needed to train multiple predictors and (2) exploit the similarity across co-location configurations (e.g., the same models on an 8vCPU VM vs. a 32vCPU VM).

## Evaluation

![](/files/sanQ09lzrhVJzvw5jFKQ)

## Links & References

* [Paper PDF](https://dl.acm.org/doi/pdf/10.1145/3437984.3458837)


# \[2021 OSDI] Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning

## One-line Summary

Pollux co-adaptively optimizes DL job execution at both per-job level and cluster-wide level. At the per-job level, Pollux uses GNS to dynamically tune the batch size to optimize goodput, a new metric that considers both system throughput and statistical efficiency of DL training. At the cluster-wide level, Pollux optimizes the generalized mean goodput of all jobs, alongside cluster-level goals including fairness and JCT.

## Paper Structure Outline

1. Introduction
2. Background: Distributed DL Training
   1. System Throughput
   2. Statistical Efficiency
   3. Existing DL Schedulers
3. The Goodput of DL Training and Pollux
   1. Modeling Statistical Efficiency
   2. Modeling System Throughput
4. Pollux Design and Architecture
   1. PolluxAgent: Job-Level Optimization
   2. PolluxSched: Cluster-wise Optimization
   3. Implementation
5. Evaluation
   1. Experimental Setup
   2. Testbed Macrobenchmark Experiments
   3. Simulator Experiments
      1. Scheduling Fairness
      2. Other Effects on Scheduling
   4. More Applications of Pollux
      1. Cloud Auto-scaling
      2. Hyper-parameter Optimization (HPO)
   5. Artifact
6. Additional Related Work
7. Conclusion
8. Acknowledgments

## Background & Motivation

### Motivation: System Throughput & Statistical Efficiency, Dynamicity in DL Training Jobs

* Training using larger batch sizes increases the system throughput
* However, as batch sizes increase, potential issues may occur:
  * If the learning rate is not tuned accordingly, the final model quality may be suboptimal
  * Increasing the batch size decreases the statistical efficiency of DL training
    * Statistical efficiency: training progress per unit of training data processed
  * Even further increasing the bs results in worse model generalization (in terms of validation performance, due to unknown reasons)
* Gradient Noise Scale (GNS) measures the noise-to-signal ratio of the stochastic gradient. It allows training jobs to increase their batch sizes later on during training without hurting the statistical efficiency.

![Suboptimal statistical efficiency after increasing batch size](/files/-MexDXwlI8suXxmC-oNs)

![To achieve the best training performance, we should aim for the point where statistical efficiency \* system throughput reaches its maximum](/files/-MexIBjaEVNUHSbQ0XLz)

![How jobs can check GNS and increase their batch sizes without hurting the statistical efficiency. This graph is excerpted from Kungfu (OSDI '20)](/files/-MexCyEKW4Y5kziABZNr)

### Background: Existing DL Schedulers

* Non-scale-adaptive schedulers are not aware of jobs' performance scalabilities w\.r.t. the amount of allocated resources
  * [Tiresias](/machine-learning-systems/machine-learning-systems-index/tiresias-a-gpu-cluster-manager-for-distributed-deep-learning) requires users to specify #GPUs during job submission
  * [Gandiva](/machine-learning-systems/machine-learning-systems-index/gandiva-introspective-cluster-scheduling-for-deep-learning) does the same, and although it may dynamically change the number of GPUs used by a job, it does so opportunistically
* Scale-adaptive schedulers automatically decide the amount of resources allocated to each job to speed up jobs
  * Optimus: Learns a predictive model for the system throughput given different amounts of resources, and optimizes the avg JCT
  * SLAQ: Minimize the avg loss values for training general ML models
  * [Gavel](/machine-learning-systems/machine-learning-systems-index/gavel-heterogeneity-aware-cluster-scheduling-policies-for-deep-learning-workloads): Takes into account the performance heterogeneity of underlying accelerators
  * AntMan: Uses dynamic scaling & fine-grained GPU sharing to improve cluster utilization, resource fairness, and JCTs
  * [Themis](/machine-learning-systems/machine-learning-systems-index/themis-fair-and-efficient-gpu-cluster-scheduling): Introduces the notion of finish time fairness
* Importantly, existing schedulers/policies are agnostic to the statistical efficiency of DL training and the inter-dependence of resource decisions and training parameters.

## Design

### Goodput = Throughput \* Statistical Efficiency

* For each job, Pollux optimizes for a new metric called the goodput

![Goodput definition](/files/-MexHzJyP3HhmPpvz-GK)

* When a user submits the job, he/she submits an initial batch size and learning rate, and Pollux will run the job with these initial hyperparameters (with s=0). As the job progresses, Pollux learns and refines predictive models for both throughput and efficiency through profiling. Then, Pollux periodically re-tunes (a, m, s).
* Learning rate scaling: Pollux allows users to implement/select scaling rules like AdaScale, square-root scaling, linear scaling, etc.

### Modeling Statistical Efficiency

* The statistical efficiency E is measured relative to the initial batch size & learning rate -> 0 < E <= 1, and training using batch size M will need to process 1/E times as many training examples to make the same progress as using batch size M0.

![Definition of pre-conditioned GNS](/files/-MexSwp90vYhliJ_hf13)

![Expression for efficiency](/files/-MexLqQAWDwnGLIU6mkk)

* During training, Pollux estimates the value of ϕt, then uses the efficiency expression to predict the efficiency at different batch sizes. Note that ϕt varies according to the training progress at iteration t.

![Takeaways: (1) For larger batach sizes, the initial efficiency is low, but this improves later in training. (2) The model can accurately predict the efficiency of using a different batch size without training at that batch size](/files/-MexU3opzyCrWshsbfyj)

### Modeling System Throughput

* Pollux separately models T\_grad, time for local gradient computations, and T\_sync, the avg time spent in each iteration for gradient averaging/model synchronizations:
  * Tgrad: The run time scales linearly with the per-GPU batch size m. Thus, Tgrad is modeled as `Tgrad(m) = α_grad + β_grad * m`, where α and β are fittable parameters.
  * Tsync: For single-GPU jobs, Tsync = 0. Otherwise, Pollux models Tsync by using a linear function of #GPUs in data parallelism and taking into account the performance retrogressions when using 3+ GPUs (due to stragglers/network bottlenecks)
    * N: Number of physical nodes occupied by at least one replica
    * α and β (local, sync): Constant and retrogression parameters for when all processes are co-located onto the same node
    * α and β (node, sync): Analogous parameters for when at least two processes are located on different nodes
    * This model can also account for rack-level locality by adding a third pair of parameters

![Modeling Tsync](/files/-MexXxnzhRVySElj_0Ow)

* Overlapping computation and synchronization: Then, Pollux combines Tgrad and Tsync. If there is no overlapping between gradient computation and synchronization, then `Titer = Tgrad + Tsync`. With perfect overlapping, `Titer = max(Tgrad, Tsync)`. Realistically, Pollux sets a learnable parameter γ to express the level of overlapping
  * γ >= 1. When γ == 1, there is no overlapping. As γ→∞, it transitions towards perfect overlapping.
* Gradient accumulation is a technique that allows for larger batch sizes, bypassing the GPU memory constraints ([quick tutorial](https://kozodoi.me/python/deep%20learning/pytorch/tutorial/2021/02/19/gradient-accumulation.html#2.-What-is-graident-accumulation): instead of updating the network weights on every batch, we save gradient values, proceed to the next batch, and add up the new gradients). With s = 0, there is no gradient accumulation. Otherwise, one iteration of SGD spans s accumulation steps and one synchronization step.
* Combining all the above, we have:

![](/files/-MexLdH5kSysedApjgvA)

* ... and the predictions are pretty good!

![](/files/-Mex_wlZdRpL2oMLx-DT)

* Note that there are many other factors (except number and co-locality of GPUs, bs, gradient accumulation steps in Pollux) that may affect the data-parallel throughput, including specialized hardware, sophisticated synchronization algorithms, different parallelization strategies, larger scales, or hidden resource contention. Pollux does not cover those, but it designs the goodput metric so that different equations for throughput may be easily plugged in. Dayum!

## Implementation

![](/files/-MexaVI9ZXaf593kpLCK)

### Job-level optimization: PolluxAgent

* An instance of PolluxAgent is started with each training job
* PolluxAgent then measures GNS & throughput, fits the EFFICIENCY and THROUGHPUT functions for that job, and tunes its batch size and learning rate for efficient utilization of its current allocated resources
* Finally, PolluxAgnet periodically reports to PolluxSched

### Cluster-wide optimization: PolluxSched

* PolluxSched periodically optimizes the resources allocations for jobs in the cluster to maximize FITNESS, taking into account the goodput function for each job and cluster-wide resource contention. The decisions also account for {re-allocation overhead, slowdowns due to network interference between jobs, resource fairness}

![Note that af is a fair resource allocation for the job, defined to be an exclusive 1/J share of the cluster. This is similar to finish time fairness, but SPEEDUP is related to training performance at a moment in time, whereas FTF is related to end-to-end JCT.](/files/-Mez5jaQDbjZyhhf8AsN)

* Re-allocation penalty: The per-job SPEEDUP is applied a penalty, which results in jobs with historically higher rates of re-allocations being penalized more for future re-allocations.
* Interference avoidance: Pollux simply disallows different distributed nodes from sharing the same node.
* Non-adaptive jobs: EFFICIENCY is fixed to be 1, and Pollux can continue to adapt its resource allocations based on system throughput.

## Evaluation

* [Tiresias](/machine-learning-systems/machine-learning-systems-index/tiresias-a-gpu-cluster-manager-for-distributed-deep-learning) and Optimus are used as the baseline schedulers. Optimus only adapts the number of GPUs and Tiresias adapts neither.
* For baseline schedulers, manually-tuned jobs are used as follows. A "#GPUs" is considered valid if using the optimal batch size for that "#GPUs" achieves 50% - 80% of the ideal/linear scalability vs. using the optimal batch size on a single GPU. For each job, the #GPUs and batch size are selected randomly from its set of valid configurations.
  * These configurations assume that users are highly rational and knowledgeable about the scalability of the models (...) in favor of the baseline schedulers
  * Less than 50% -> undertilization of resources
  * More than 80% -> more GPUs can still be utilized efficiently
* Pollux does really well in trading-off throughput with efficiency. For example, during periods of low cluster contention, Pollux can allocate more GPUs (& larger batch sizes) to boost training throughput & goodput, even if it decreases the statistical efficiency, and vice versa

![](/files/-MezD4mgnI8_8q7SbFHe)

![](/files/-MezDjq4U6Um3yy_DlG_)

![](/files/-MezDmYHO8cGOGagkR4h)

## Links

* [Paper PDF](https://www.pdl.cmu.edu/PDL-FTP/CloudComputing/osdi21-pollux.pdf)
* [Presentation video at OSDI '21](https://www.usenix.org/conference/osdi21/presentation/qiao)
* [Presentation slides at OSDI '21](https://www.usenix.org/system/files/osdi21_slides_qiao.pdf)
* [Pollux on GitHub](https://github.com/petuum/adaptdl/tree/osdi21-artifact)
* My [presentation video](https://youtu.be/lJ3_iM13A5k) on Pollux (in Mandarin Chinese) at the Systems & Networking Reading Group hosted by [Xiangfeng Zhu](https://xzhu27.me/)


# \[2021 MLSys] Wavelet: Efficient DNN Training with Tick-Tock Scheduling

## One-line Summary

Both data and model parallelism suffer from system under-utilization. Wavelet exploits the under-utilized memory & compute by scaling up the number of training tasks and launching the additional tasks with a delay to fully utilize the on-chip memory and improve the compute utilization, speeding up individual jobs.

That was an unnecessarily long sentence... GRE took its toll on me!

![Example of how Wavelet is applied to data parallel training](/files/-M_T5sjvVPUFScO9_KXZ)

## Paper Structure Outline

1. Introduction
2. Background and Motivation
   1. Distributed DNN Training Schemes
   2. Jobs Characteristics of Distributed DNN Training
      1. Zoom-in analysis on data parallel training
      2. Sub-iteration analysis on model parallel training
3. Wavelet Design
   1. System Overview
   2. Wavelet in Data Parallelism
      1. Memory overlapping
      2. Computation overlapping
      3. Model synchronization between waves
   3. Wavelet in Model Parallelism
      1. Launching multiple tock-wave tasks
      2. Model partition switching
      3. Inter-batch synchronization
4. Evaluation
   1. Data parallelism
      1. Single machine multi-GPU
      2. Multi-machine multi-GPU
   2. Model parallelism
      1. Single machine multi-GPU
      2. Multi-machine multi-GPU
      3. Overhead analysis
5. Related Work
   1. Resource allocation for distributed DNN training
   2. GPU sharing
6. Conclusion

## Background & Motivation

Bigger models & datasets call for large-scale distributed machine learning training. The current scheduling policy, gang scheduling, where all training tasks on all workers need to be launched at the same time, contributes to the under-utilization of system resources (compute & memory).

In Fig. 1 (see above), computation is memory-bounded during the forward propagation. Between time 0.4 and 0.6, memory is underutilized in the backward propagation. Moreover, \~60% of on-chip compute cores are underutilized.

![The same thing happens for model parallelism where the memory valley is longer. With pipelining = using GPipe.](/files/-M_T6rQ-z_gNU6mUSQt5)

There are existing job multiplexing schemes that boost system utilization. [Gandiva](/machine-learning-systems/machine-learning-systems-index/gandiva-introspective-cluster-scheduling-for-deep-learning) lets 2 low-utilization jobs space-share a GPU. [Salus](/machine-learning-systems/machine-learning-systems-index/2020-sigcomm-reducto-on-camera-filtering-for-resource-efficient-real-time-video-analytics/salus-fine-grained-gpu-sharing-primitives-for-deep-learning-applications) provides fine-grained memory sharing via the GPU Lane abstraction. However, neither scheme contributes to the training progress of a single job. In this work, Wavelet relaxes the gang scheduling scheme and accelerates a single job while improving the system utilization.

## Design and Implementation

### Data parallelism

#### Model synchronization

![How Wavelet handles the extra tock-wave during model synchronization](/files/-M_T8uUJX0l50JtjOT0A)

In vanilla allreduce, there are only Main & Tick waves. With this Tock wave added, Wavelet doubles the number of model synchronizations. This is the same as synchronizing over 2\*N data parallel tasks on 2\*N GPUs, thus guaranteeing convergence. &#x20;

#### Overlapping memory

![](/files/-M_T9ieUQYSnwvF_HZFV)

In gang scheduling, the memory of all GPUs is underutilized during backprop. Tick-tock scheduling injects tock-wave tasks right after the tick-wave tasks finish the forward pass. To concurrently run 2 tasks (tick & tock), 2 model replicas are maintained on the GPU since the two waves train on different data. In the memory, the size of the model is way smaller than the size of the intermediate results, so no need to worry about the extra memory.

#### Overlapping computation

CUDA computation kernels are launched in separate CUDA streams to ensure ordered execution within a stream and non-blocking across different streams. The empty bubbles between kernels is due to the latency of CPUs sending instructions to GPUs.

### Model parallelism

![](/files/-M_TAtcPPNyfS_KLQqbc)

In the vanilla pipelined process (white blocks), only 1 batch is active in the system and at most 1 GPU is active at a time. Each GPU also holds the same model partition during the whole training process.

In Wavelet, we inject 3 (N-1 w/ N GPUs) tock waves on 1 tick wave. The model partition is swapped on each GPU using a round-robin fashion. There exists an extra model synchronization for each model partition, and the context switching also brings overhead.

## Evaluation

### Data parallelism

![](https://lh4.googleusercontent.com/VsZF6a3W2RzY3BYXFP7kHY7GKpLN-5VyPr01iWp9PiBj0c6iOypabT8VcZ0GCe6P9KqBqUqkhzxi_ssKLUIvzYhONHTu-fGPuilYAt_MKiiqag0o-ffGhyAYYOKh6eM4vydfTl_1O8E)

* Single machine: Up to 1.88x speedup (avg: 1.4x, theoretically 2x) over DP baseline
* Multiple machine: Up to 1.18x speedup over baseline. The worse throughput than baseline is the overhead kicking in: The cross-machine low-bandwidth network becomes the bottleneck during the extra allreduce

### Model parallelism

![](https://lh3.googleusercontent.com/cXWkPGS1OhaFhnrv2vvFVuNrTm1ySzDU9JuHMIBS2_x9rLsul1nIJrzYAm8nPEDQa47hHXG2mUmw_U3gTRAoBV1Fnwy_c9LrqDRDkpxOjhqFD_l9E38gbwHVyBJiVfhz99gzTLbJ79k)

* Only \~2.5x speedup in 4x/8x parallelism
  * Gpipe/PipeDream breaks a mini-batch into smaller micro-batches -> High-frequency but small-size data chunks
  * Number of CUDA kernel calls **↑**
  * Intermediate result transfer between GPUs that hold different model partitions **↑**
  * Linear scalability in theory

### Overhead analysis

![Note that this is in log-scale](https://lh3.googleusercontent.com/RTdgfup_rPnHzhgPbMq_1eflWzzhpgvO7W7aSB0pousFIRUQ6cokC8gVYZ5K_wZ_e1mXOxGm7FBnHpGInHYltnTsbtW9aKmZNl6v6k9dYw0_iQ3rwulgMpqn-gXGeqlNKGYe6pmdNRs)

* Context switch: Switching model partition, \~4% of total training time
* Communication: Transferring intermediate results across GPUs (\~15% of total training time)
* All reduce: Model synchronization during backprop (\~4% of total training time)

## Links

* [Paper PDF](https://proceedings.mlsys.org/paper/2021/file/c81e728d9d4c2f636f067f89cc14862c-Paper.pdf)
* [Presentation video at MLSys '21](https://mlsys.org/virtual/2021/oral/1586)
* [Presentation slides at MLSys '21](https://mlsys.org/virtual/2021/oral/1586)


# \[2021 NSDI] SwitchML: Scaling Distributed Machine Learning with In-Network Aggregation

## Summary

Modern distributed ML training is communication-intensive. Thanks to the corporate overlords, emerging hardware shows up for help. Programmable switches can aggregate model updates in-network, making the network itself an accelerator for ML.

## Background & Motivation

In recent years, we have seen orders of magnitude faster capability improvements in compute than networks. Furthermore, the ratio of communication to computation in the workload itself has shifted. As a result, in distributed training, the network is becoming the bottleneck.

A new approach for model updates is in-network aggregation. In this approach, workers send their model updates over the network, where an aggregation primitive in the network sums the updates and distributes only the resulting value. This offers a fundamental advantage over all-reduce and PS since it avoids end-host processing required to perform aggregation and therefore provides "sub-RTT" latency.

## Design & Implementation

The idea sounds amazing but it comes with challenges. First, switches' packet processing capabilities are limited, and ML uses floating-point values, while integer computing is the norm in programmable switches. Second, on-chip memory is also small (tens of MBs while model updates might have hundreds of megabytes of gradients). Finally, the system must be resilient to packet loss without impact on efficiency or correctness. To this end, the authors propose SwitchML which co-designs in-switch processing with an end-host transport layer and ML frameworks.

### SwitchML overview

* Combined switch-host architecture: The switch performs integer aggregation, while end hosts are responsible for managing reliability and performing more complex computations.
* Pool-based streaming aggregation: SwitchML streams aggregation through the switch. End hosts handle the management of aggregators in a pool, leaving the switch dataplane with a simple design.
* Fault-tolerant protocols: Recover from packet loss with minimal overheads & handles worker/network failures
* Quantized integer-based aggregation: Floating-point values are converted to 32-bit integers to satisfy the computing power of switches. This process is done at end hosts without impacting training accuracy.

![](/files/yqQY10oju87ShDyxJF1g)

### Aggregation protocol

* Switch-side: A pool-based design addresses two limitations. First, it removes the need to store an entire model update on a switch at once. Second, it allows processing to be done at the packet level by performing the aggregation in small pieces, at most k integers at a time.
* Worker-side: After the initial batch of packets is sent, each worker only sends a new packet with the next piece of update once it has received the aggregated packets returned from the switch. This simple communication scheme does not require any explicit coordination among workers yet still achieves agreement on which slots to use.

### Packet loss

The natural way to deal with packet losses is retransmissions after timeouts. However, this naive approach has two main challenges: (1) differentiating packets that are lost on the upward paths vs. the downward ones, and (2) being able to retransmit an aggregated response that is lost on the way back to a worker. The solutions are (1) explicitly maintaining information as to which workers have already contributed updates to a given slot to ignore duplicate transmissions, and (2) maintaining a shadow copy of the previous result for each slot, which allows the switch to retransmit a dropped result packet for a slot even when the switch has started reusing the slot for the next chunk. This ensures that no worker node can ever lag more than one chunk behind any of the others for a particular slot.

### Quantizing floating-point values

SwitchML uses a numeric representation, inspired by block floating-point, that combines 32-bit fixed-point addition in the switch with adaptive scaling on the workers. This representation is used only when aggregating gradients. Empirically, this does not hurt convergence.

## Evaluation

![](/files/Mu1IsfQBFtYOOeYUUM2N)

## Links & References

* [Paper PDF](https://www.usenix.org/system/files/nsdi21-sapio.pdf)
* [Presentation video at NSDI '21](https://www.youtube.com/watch?v=FIZsXfeZrvE)
* [Presentation slides at NSDI '21](https://www.usenix.org/system/files/nsdi21_slides_sapio.pdf)
* [SwitchML on GitHub](https://github.com/p4lang/p4app-switchML)


# Big Data Systems - Index

## Table of Contents

### Infrastructure, Frameworks, and Paradigms

* [NFS: Sun's Network File System](/earlier-readings-and-notes/index/nfs-suns-network-file-system)
* [\[SOSP '03\] The Google File System](/machine-learning-systems/index/the-google-file-system)
* [\[OSDI '04\] MapReduce: Simplified Data Processing on Large Clusters](/machine-learning-systems/index/mapreduce-simplified-data-processing-on-large-clusters)
* \[SOSP '09] FAWN: A Fast Array of Wimpy Nodes
* [\[NSDI '11\] Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center](/machine-learning-systems/index/mesos-a-platform-for-fine-grained-resource-sharing-in-the-data-center)
* [\[NSDI '12\] Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing](/machine-learning-systems/index/resilient-distributed-datasets-a-fault-tolerant-abstraction-for-in-memory-cluster-computing)
* \[HotOS '15] Scalability! But at what COST? ([pdf](https://www.usenix.org/system/files/conference/hotos15/hotos15-paper-mcsherry.pdf))
* [\[HotOS '21\] From Cloud Computing to Sky Computing](/machine-learning-systems/index/from-cloud-computing-to-sky-computing)

### Scheduling & Resource Allocation

* \[NSDI '11] Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center
* \[EuroSys '13] Omega: flexible, scalable schedulers for large compute clusters
* \[SoCC '13] Apache Hadoop YARN: Yet Another Resource Negotiator
* \[SoCC '14] Wrangler: Predictable and Faster Jobs using Fewer Resources
* \[OSDI '14] Apollo: Scalable and Coordinated Scheduling for Cloud-Scale Computing ([pdf](https://www.usenix.org/system/files/conference/osdi14/osdi14-paper-boutin_0.pdf))
* \[SIGCOMM '14] Tetris: Multi-Resource Packing for Cluster Schedulers ([pdf](https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/tetris_sigcomm14.pdf))
* \[ASPLOS '14] Quasar: Resource-Efficient and QoS-Aware Cluster Management
* \[SIGCOMM '15] Network-Aware Scheduling for Data-Parallel Jobs: Plan When You Can
* \[OSDI '16] CARBYNE: Altruistic Scheduling in Multi-Resource Clusters ([pdf](https://www.usenix.org/system/files/conference/osdi16/osdi16-grandl-altruistic.pdf))
* \[OSDI '16] Packing and Dependency-aware Scheduling for Data-Parallel Clusters
* \[NSDI '16] HUG: Multi-Resource Fairness for Correlated and Elastic Demands
* \[EuroSys '16] TetriSched: global rescheduling with adaptive plan-ahead in dynamic heterogeneous clusters
* \[SoCC '17] Selecting the best vm across multiple public clouds: A data-driven performance modeling approach
* \[ATC '18] On the diversity of cluster workloads and its impact on research results

### Cloud/Serverless Computing

* \[SoCC '17] Occupy the Cloud: Distributed Computing for the 99%
* \[arXiv '19] Cloud Programming Simplified: A Berkeley View on Serverless Computing
* \[SoCC '19] Centralized Core-granular Scheduling for Serverless Functions
* \[SoCC '19] Cirrus: a Serverless Framework for End-to-end ML Workflows
* \[NSDI '19] Shuffling, Fast and Slow: Scalable Analytics on Serverless Infrastructure
* \[SIGMOD '20] Le Taureau: Deconstructing the Serverless Landscape & A Look Forward
* \[SoCC '20] Serverless linear algebra
* \[ATC '20] Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud Provider
* \[SIGMOD '21] Towards Demystifying Serverless Machine Learning Training
* \[OSDI '21] Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads
* \[SoCC '21] Atoll: A Scalable Low-Latency Serverless Platform
* \[NSDI '21] Caerus: Nimble Task Scheduling for Serverless Analytics
* \[ASPLOS '22] Serverless computing on heterogeneous computers
* \[arXiv '22] Groundhog: Efficient Request Isolation in FaaS ([pdf](https://arxiv.org/pdf/2205.11458.pdf))

### Network Flow Scheduling

* \[SIGCOMM '11] Managing Data Transfers in Computer Clusters with Orchestra
* [\[HotNets '12\] Coflow: A Networking Abstraction for Cluster Applications](/machine-learning-systems/index/big-data-systems-papers-short-notes#2012-hotnets-coflow-a-networking-abstraction-for-cluster-applications)
* [\[SIGCOMM '14\] Efficient coflow scheduling with Varys](/machine-learning-systems/index/big-data-systems-papers-short-notes#2014-sigcomm-efficient-coflow-scheduling-with-varys)
* \[SIGCOMM '14] Barrat: Decentralized task-aware scheduling for data center networks
* [\[SIGCOMM '15\] Aalo: Efficient coflow scheduling without prior knowledge](/machine-learning-systems/index/big-data-systems-papers-short-notes#2015-sigcomm-aalo-efficient-coflow-scheduling-without-prior-knowledge)
* \[SIGCOMM '16] CODA: Toward Automatically Identifying and Scheduling COflows in the DArk
* \[SIGCOMM '16] Scheduling Mix-flows in Commodity Datacenters with Karuna ([pdf](https://people.csail.mit.edu/alizadeh/papers/karuna-sigcomm16.pdf))
* \[SIGCOMM '18] Sincronia: Near-Optimal Network Design for Coflows
* \[SPAA '19] Near Optimal Coflow Scheduling in Networks

### Graphs

* [MIT's 6.886 Graph Analytics reading list](https://people.csail.mit.edu/jshun/6886-s18/) by [Prof. Julian Shun](https://people.csail.mit.edu/jshun/)
* [\[SIGMOD '10\] Pregel: A System for Large-Scale Graph Processing](/machine-learning-systems/index/pregel-a-system-for-large-scale-graph-processing)
* [\[OSDI '12\] PowerGraph: Distributed Graph-Parallel Computation on Neural Graphs](/machine-learning-systems/index/powergraph-distributed-graph-parallel-computation-on-natural-graphs)
* \[PPoPP '13] Ligra: A Lightweight Graph Processing Framework for Shared Memory ([pdf](https://people.csail.mit.edu/jshun/ligra.pdf))
* \[OSDI '14] GraphX: Graph Processing in a Distributed Dataflow Framework
* \[ATC '17] Garaph: Efficient GPU-accelerated Graph Processing on a Single Machine with Balanced Replication
* \[EuroSys '17] MOSAIC: Processing a Trillion-Edge Graph on a Single Machine ([pdf](https://dl.acm.org/doi/pdf/10.1145/3064176.3064191))
* \[VLDB '18] A Distributed Multi-GPU System for Fast Graph Processing ([pdf](https://dl.acm.org/doi/pdf/10.14778/3157794.3157799))
* \[SoCC '20] PaGraph: Scaling GNN Training on Large Graphs via Computation-aware Caching ([pdf](https://dl.acm.org/doi/pdf/10.1145/3419111.3421281))
* [\[EuroSys '21\] NextDoor: Accelerating graph sampling for graph machine learning using GPUs](/machine-learning-systems/index/accelerating-graph-sampling-for-graph-machine-learning-using-gpus)
* \[OSDI '21] Marius: Learning Massive Graph Embeddings on a Single Machine
* \[arXiv '22] Marius++: Large-Scale Training of Graph Neural Networks on a Single Machine
* \[MLSys '22] Graphiler: Optimizing Graph Neural Networks with Message Passing Data Flow Graph

### Distributed Tracing

* \[Textbook] Distributed Tracing in Practice
* \[SOSP '15] Pivot tracing: dynamic causal monitoring for distributed systems ([pdf](https://dl.acm.org/doi/pdf/10.1145/2815400.2815415))
* \[SoCC '16] Principled Workflow-Centric Tracing of Distributed Systems ([pdf](https://dl.acm.org/doi/pdf/10.1145/2987550.2987568))
* \[SOSP '17] Canopy: An End-to-End Performance Tracing And Analysis System ([pdf](https://dl.acm.org/doi/pdf/10.1145/3132747.3132749))
* \[SoCC '18] Weighted Sampling of Execution Traces: Capturing More Needles and Less Hay ([pdf](https://dl.acm.org/doi/pdf/10.1145/3267809.3267841))
* \[SoCC '19] Sifter: Scalable Sampling for Distributed Traces, without Feature Engineering ([pdf](https://dl.acm.org/doi/pdf/10.1145/3357223.3362736))
* \[HotNets '21] Snicket: Query-Driven Distributed Tracing ([pdf](https://www.cs.princeton.edu/~ravian/publications/snicket_hotnets21.pdf))
* \[NSDI '23] The Benefit of Hindsight: Tracing Edge-Cases in Distributed Systems ([pdf](https://arxiv.org/pdf/2202.05769.pdf))

### Caching

* \[SoCC '11] Small Cache, Big Effect: Provable Load Balancing for Randomly Partitioned Cluster Services
* \[NSDI '16] Be Fast, Cheap and in Control with SwitchKV
* \[SOSP '17] NetCache: Balancing Key-Value Stores with Fast In-Network Caching
* [\[FAST '19\] DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed Caching](/machine-learning-systems/index/2019-fast-distcache-provable-load-balancing-for-large-scale-storage-systems-with-distributed...)

### New Data, Hardware Models

* \[ISCA '17] In-Datacenter Performance Analysis of a Tensor Processing Unit

### Databases

* \[SIGMOD '12] Towards a Unified Architecture for in-RDBMS Analytics
* \[arXiv '13] Bayesian Optimization in a Billion Dimensions via Random Embeddings
* \[SIGMOD '17] Automatic Database Management System Tuning Through Large-scale Machine Learning
* \[HotStorage '20] Too Many Knobs to Tune? Towards Faster Database Tuning by Pre-selecting Important Knobs
* \[arXiv '21] Facilitating Database Tuning with Hyper-Parameter Optimization: A Comprehensive Experimental Evaluation
* \[VLDB '21] An Inquiry into Machine Learning-based Automatic Configuration Tuning Services on Real-World Database Management Systems ([pdf](https://db.cs.cmu.edu/papers/2021/p1241-aken.pdf))
* \[VLDB '22] LlamaTune: Sample-Efficient DBMS Configuration Tuning

## Meta stuff

* Reading lists
  * [CS 294 @ Berkeley: Machine Learning Systems](https://ucbrise.github.io/cs294-ai-sys-fa19/)
  * [CS 744 @ UW-Madison: Big Data Systems](http://pages.cs.wisc.edu/~shivaram/cs744-fa20/)
  * [CS 6787 @ Cornell: Advanced Machine Learning Systems](https://www.cs.cornell.edu/courses/cs6787/2020fa/), with a focus on the ML side
  * [Awesome-System-for-Machine-Learning](https://github.com/HuaizhengZhang/Awesome-System-for-Machine-Learning): An open-sourced reading list
  * [The MLSys conference](https://mlsys.org/)
  * [SOSP AI Systems workshop](http://learningsys.org/sosp19/acceptedpapers.html)
* Some other stuff
  * Meta papers
    * [A Berkeley View of Systems Challenges for AI](https://thodrek.github.io/CS839_spring18/papers/EECS-2017-159.pdf)
    * [MLSys: The New Frontier of Machine Learning Systems](https://arxiv.org/pdf/1904.03257.pdf)
  * [Systems Benchmarking Crimes](https://www.cse.unsw.edu.au/~gernot/benchmarking-crimes.html)
  * [CSE 559W @ U Washington Slides](http://dlsys.cs.washington.edu/schedule): Not a paper reading class, more of an end-to-end comprehensive introduction of foundations of DL Systems
  * [CS 759 @ UW-Madison (HPC) Course Notes](/earlier-readings-and-notes/cs759-hpc-course-notes): A great overview of HPC, CUDA, OpenMP, MPI


# Big Data Systems Papers - Short Notes

## \[2012 HotNets] Coflow: A Networking Abstraction for Cluster Applications

Coflows make it easier for applications to convey their communicatino semantics to the network, which in turn enables the network to better optimize common communication patterns.

![](/files/ynCqmlJFh62kxAiWA7Er)

![](/files/xsjDlOg1T78eWvYMT8tQ)

## \[2014 SIGCOMM] Efficient coflow scheduling with Varys

* Smallest-Effective-Bottleneck-First (SEBF) heuristic: greedily schedules a coflow based on its bottleneck’s completion time
* Minimum-Allocation-for-Desired-Duration (MADD) algorithm: slows down all the flows in a coflow to match the completion time of the flow that will take the longest to finish, so that other coexisting coflows can make progress

<img src="/files/4n9TDbjBjI1BIxfbrkgc" alt="" data-size="original">

![](/files/NR0UkK8el352cbaEBGoz)

![](/files/RX1oC61XvKOHOhn4iV7R)

## \[2015 SIGCOMM] Aalo: Efficient coflow scheduling without prior knowledge

* Removed the clairvoyance requirements of the coflow information
* Aalo employs Discretized Coflow-Aware Least-Attained Service (D-CLAS) to separate coflows into a small number of priority queues based on how much they have already sent across the cluster. By performing prioritization across queues and by scheduling coflows in the FIFO order within each queue, Aalo’s non-clairvoyant scheduler reduces coflow completion times while guaranteeing starvation freedom.

![](/files/SUExmu2r41V7xkELNbYS)

![](/files/sArPobPgghcqXzoVEYsz)


# \[2003 SOSP] The Google File System

## One-line Summary

GFS is a system for distributed file storage. The design of GFS is motivated by Google's cluster architecture paradigm and its workload characterizations. It provides fault tolerance while running on inexpensive commodity hardware, and it delivers high aggregate performance to a large number of clients. Hadoop (HDFS) is an open-source implementation of GFS.

## Paper Structure Outline

1. Introduction
2. Design Overview
   1. Assumptions
   2. Interface
   3. Architecture
   4. Single Master
   5. Chunk Size
   6. Metadata
      1. In-Memory Data Structures
      2. Chunk Locations
      3. Operation Log
   7. Consistency Model
      1. Guarantees by GFS
      2. Implications for Applications
3. System Interactions
   1. Leases and Mutation Order
   2. Data Flow
   3. Atomic Record Appends
   4. Snapshot
4. Master Operation
   1. Namespace Management and Locking
   2. Replica Placement
   3. Creation, Re-replication, Rebalancing
   4. Garbage Collection
      1. Mechanism
      2. Discussion
   5. Stale Replica Detection
5. Fault Tolerance and Diagnosis
   1. High Availability
      1. Fast Recovery
      2. Chunk Replication
      3. Master Replication
   2. Data Integrity
   3. Diagnostic Tools
6. Measurements
   1. Micro-benchmarks
      1. Reads
      2. Writes
      3. Record Appends
   2. Real World Clusters
      1. Storage
      2. Metadata
      3. Read and Write Rates
      4. Master Load
      5. Recovery Time
   3. Workload Breakdown
      1. Methodology and Caveats
      2. Chunkserver Workload
      3. Appends vs. Writes
      4. Master Workload
7. Experiences
8. Related Work
9. Conclusions

## Background & Motivation

* A little bit of history:
  * NFS
    * File stored on central file server
    * Implements POSIX -> transparent to users
    * Client-side caching + STAT request to validate cache entries (data blocks)
    * Bad scalability: Thousands of nodes ask the server for cache validation
  * Andrew File System
    * Designed for scalee
    * Whole-file caching: Wheen a client opens a file, the entire file is read from the server & cached at the client
  * Oceanstore/Past (late 90s to early 2000s)
    * Wide area storage systems
    * Fully decentralized
    * Built on Distributed Hash Tables (DHT)

![\~80% of the files are < 4K](/files/-Mjj37zRsNaNLIhbkVdy)

* Google was only five years old in 2003: It was a relatively young company. Instead of scaling up (buying expensive servers), they chose to scale out (using commodity, inexpensive hardware), due to the huge amount of data.
* Certain aspects of the workload drive the design choices. The following observations are different from the assumptions made by previous distributed storage systems:
  * In a cluster of thousands of commodity servers, component failures are inevitable (and more frequent than that on expensive servers?), so **fault tolerance** is an important design consideration.
  * GFS is optimized for the reading and writing of a modest number (a few million) of **large files** (hundreds of MBs to multiple GBs).
  * Writes to files are mostly **append-only**: There are hardly random writes (consider a web crawler that keeps adding crawled content to a single file).
  * Reads are either large sequential reads (consider a batch processing system that reads the large file and creates a search index) or small random reads.
  * Latency is not a big concern

## Design and Implementation

![](/files/-MjeiLkWomesqOGqxVbV)

* Files are split into chunks. Each 64-MB chunk (this is much larger than traditional file system block sizes) of a file can be identified by a 64-bit ID, and the chunks are distributed on multiple machines (GFS chunkservers). Moreover, multiple (3 by default) replicas of each chunk are stored for fault tolerance. If the replication factor of a file falls below a goal (due to machine failures/corrupted replicas), chunks are re-replicated.
  * Chunk size trade-offs:
    * Client -> Master: If chunk size is small, clients need to issue more calls to master
    * Client -> Chunkserver: Too small -> need to open connections to many chunkservers; Too large -> chunkserver may become a hotspot
    * Metadata: Chunk size too small -> Metadata size grows, master's memory will suffer
  * Replication
    * Primary replica for each chunk
    * Chain replication
* GFS master: A single master server stores in memory the metadata of the cluster: The file and chunk namespaces, the mapping from files to chunks, and the locations of each chunk's replicas. Having a single master makes things vastly easier.
* GFS clients only communicate with the GFS master about the metadata of the file, and the actual I/O is done between the client and the chunkservers. GFS also caches (chunk handle, chunk location) to reduce the GFS master's workload.
* The master keeps track of an operation log, the only persistent record of metadata. In case of a master failure, it can recover the file system state by replaying the operation log.

![What happens during a write (this graph is used to describe a lease mechanism for maintain a global mutation order)](/files/-Mjf-aB0ZbZe1nCPJrDj)

* Data flow: The data is pushed linearly along a chain of chunkservers to fully utilize each machine's outgoing bandwidth. Aside from that, each machine forwards the data to the closest (the distance can be estimated from IP addresses) machine to avoid network bottleneck.
* GFS also provides an atomic append operation, record append, that allows multiple clients to concurrently append to the same file. If this is done using traditional writes, clients would need to do complicated & expensive synchronization.
  * Consistency
    * At-lease once: The append record will appear at least once in the chunk
    * Atomic: The entire record will appear

## Evaluation

![](/files/-Mjf9xhPXCrxcEwVh-wU)

![Cluster A is for development and cluster B is for production. Read rate > write rate](/files/-MjfA-HXhK2m4-nSPnu5)

## Links & References

* [Paper PDF](http://pages.cs.wisc.edu/~shivaram/cs744-readings/GFS.pdf)
* [Presentation video by Defog Tech on YouTube](https://youtu.be/eRgFNW4QFDc)


# \[2004 OSDI] MapReduce: Simplified Data Processing on Large Clusters

## One-line Summary

MapReduce is a simple programming model on large clusters with frequent failures. It provides a set of limited but general functional API (Map, Reduce, Sort), fault tolerance, and straggler mitigation through retries.

## Paper Structure Outline

1. Introduction
2. Programming Model
   1. Example
   2. Types
   3. More Examples
3. Implementation
   1. Execution Overview
   2. Master Data Structures
   3. Fault Tolerance
   4. Locality
   5. Task Granularity
   6. Backup Tasks
4. Refinements
   1. Partitioning Function
   2. Ordering Guarantees
   3. Combiner Function
   4. Input and Output Types
   5. Side-effects
   6. Skipping Bad Records
   7. Local Execution
   8. Status Information
   9. Counters
5. Performance
   1. Cluster Configuration
   2. Grep
   3. Sort
   4. Effect of Backup Tasks
   5. Machine Failures
6. Experience
   1. Large-Scale Indexing
7. Related Work
8. Conclusions

## Programming Model

The data type for each record is of the form (key, value).&#x20;

The terms "map" and "reduce" are borrowed from functional languages like Lisp. The Map function (parallelly) processes (a large number of) individual records to generate intermediate (key, value) pairs. The Reduce function (parallelly) processes and merges all intermediate values associated per key by partitioning keys (e.g., hash partitioning).

![Map function & Reduce function](/files/-Md38aCm64zBbIw8vvNJ)

## Example Workloads

### Word Count

![](/files/-Md3A8R8tYTBsgqAjTAN)

### Distributed grep

![](/files/-Md3AEqbFPknEpL6S-Kq)

### Reversed Web-Link Graph

![](/files/-Md3AKO_7_fujyC2cEzt)

### Count of URL Access Frequency

![](/files/-Md3AR_IZlXQSG6Kuso3)

### Sort

![](/files/-Md3AY3Q1zkji43h4_oh)

## MapReduce Scheduling

### Inside MapReduce

For the users, they only need to write the map & reduce programs, then submit the job and wait for the results, without the need to know about the internal parallel/distributed computing. For the paradigm and the scheduler, the following need to be handled:

1. Parallelize Map
2. Transfer data from Map to Reduce: Use partitioning function, ensuring all map output records with the same key are assigned to the same Reduce task
3. Parallelize Reduce
4. Implement storage for Map input, Map output, Reduce input, Reduce output
   1. Map input: From distributed FS (GFS, HDFS, etc.)
   2. Map output: To local FS/disk at Map node
   3. Reduce input: From (multiple) remote disks; Uses local FS
   4. Reduce output: To distributed FS
5. Ensure the barrier between the Map phase and Reduce phase

![](/files/-Md3DGvseulot7GQZdss)

### The YARN Scheduler

Yet Another Resource Negotiator (YARN) is offered in Hadoop 2.x+. It treats each server as a collection of containers (some CPU + some memory). It has three main components:

1. Global Resource Manager (RM): Scheduling
2. Per-Server Node Manager (NM): Daemon and Server-specific functions
3. Per-Application/job Application Master (AM): Handles container negotiations with RMs and NMs, detect task failures of that job

![](/files/-Md3EhS0O0DUmtgnbcjO)

## Other Designs

### Fault Tolerance: Failures

* Server Failure
  * NM heartbeats to RM: If server fails, RM lets all affected AMs know, and AMs take action
  * NM keeps track of each task running at its server: If a task fails while in progress, mark the task as idle and restart it. If the same task fails repeatedly, end the job
  * AM heartbeats to RM: On failure, RM restarts AM, which then syncs up with its running tasks
* RM Failure
  * Use old checkpoints and bring up secondary RM&#x20;

### Fault Tolerance: Stragglers

* The slowest machine slows the entire job down
* Possible reasons: bad disk, network bandwidth, CPU, or memory
* Keep track of the progress of each task (% done). When a straggler appears, launch a second copy of a task on another node and take the output of whichever finishes first (this is called Speculative Execution).

### Locality

* Cloud has hierarchical topology (e.g., racks)
* GFS/HDFS stores 3 replicas of each chunk (e.g., 64 MB in size), possibly on different racks
* MapReduce attempts to schedule a Map task on (preference from high to low):
  * A machine that contains a replica of corresponding input data
  * On the same rack as a machine containing the input
  * Anywhere

## Links

* [Paper PDF](https://static.googleusercontent.com/media/research.google.com/en//archive/mapreduce-osdi04.pdf)
* [Course notes from CS 744 @ UW-Madison](http://pages.cs.wisc.edu/~shivaram/cs744-fa20-slides/cs744-mapred-notes.pdf)
* [Official MapReduce Tutorial from Apache Hadoop](https://hadoop.apache.org/docs/r1.2.1/mapred_tutorial.html)
* [Cloud Computing by Prof. Indranil Gupta from UIUC, offered on Coursera](https://www.coursera.org/specializations/cloud-computing)


# \[2010 SIGMOD] Pregel: A System for Large-Scale Graph Processing

## One-line Summary

Pregel is a computational model/message passing abstraction that allows users to express many graph algorithms with ease. In Pregel, vertex programs run in sequences of super-steps, and the programming model is essentially "thinking like a vertex".&#x20;

## Paper Structure Outline

1. Introduction
2. Model of Computation
3. The C++ API
   1. Message Passing
   2. Combiners
   3. Aggregators
   4. Topology Mutations
   5. Input and Output
4. Implementation
   1. Basic Architecture
   2. Fault Tolerance
   3. Worker Implementation
   4. Master Implementation
   5. Aggregators
5. Applications
   1. PageRank
   2. Shortest Paths
   3. Bipartite Matching
   4. Semi-Clustering
6. Experiments
7. Related Work
8. Conclusions and Future Work

## Background & Motivation

Graphs are getting bigger (in terms of #vertices/#edges). A list of challenges in implementing large-scale graph processing algorithms is as follows.

![](/files/-McXTCJ7Mi_rTzLVVCQb)

Pregel is a system for processing large-scale graphs in a distributed fashion. Its API allows for arbitrary graph algorithms to be expressed with ease.

## Design

### The Pregel Programming Model

> Pregel computations consist of a sequence of iterations, called supersteps. During a superstep the framework invokes a userdefined function for each vertex, conceptually in parallel. The function specifies behavior at a single vertex V and a single superstep S. It can read messages sent to V in superstep S − 1, send messages to other vertices that will be received at superstep S + 1, and modify the state of V and its outgoing edges. Messages are typically sent along outgoing edges, but a message may be sent to any vertex whose identifier is known.\
> \
> Algorithm termination is based on every vertex voting to halt. In superstep 0, every vertex is in the active state; all active vertices participate in the computation of any given superstep. A vertex deactivates itself by voting to halt. This means that the vertex has no further work to do unless triggered externally, and the Pregel framework will not execute that vertex in subsequent supersteps unless it receives a message. If reactivated by a message, a vertex must explicitly deactivate itself again. The algorithm as a whole terminates when all vertices are simultaneously inactive and there are no messages in transit.

![A simple example where the largest value is propageted to every vertex.](/files/-McWrg-_FbVCaqzbYIIN)

![The Pregel API](/files/-McXNhwvjHkgbbdwRqCi)

### Combiners

To reduce the overhead when sending a message, Pregel provides combiners, which are user-defined functions that allow multiple messages to be coalesced, reducing the message traffic. Note that combiners should only be enabled for commutative and associative operations, as there is no guarantee about {which messages are combined, the order of combining, the groupings presented to the combiner, etc}.

### Aggregators

Aggregators are used for global information exchange:

> Each vertex can provide a value to an aggregator in superstep S, the system combines those values using a reduction operator, and the resulting value is made available to all vertices in superstep S + 1.

### Topology Mutations

An example use case is clustering algorithms, in which each cluster might be replaced with a single vertex. In the Compute() function, requests can be issued to add/remove vertices/edges.&#x20;

## Implementation

### Pregel Architecture

A graph is partitioned into partitions, each containing a set of vertices and their outgoing edges. The default partitioning function is a hash function, while users can also use custom assignment functions to better exploit locality (e.g., colocating vertices representing pages of the same site).

The execution stages are as follows:

> FIXME: This is the part where I get a bit confused -- if the user input is loaded in step 3, then how come we can already determine the partitions in step 2?

1. Many copies of the user program begin executing on a cluster of machines. One of the copies acts as the master to coordinate worker activities.
2. The master determines the number of partitions and assigns partitions to machines. It is possible to have multiple partitions per worker for parallelism & load balancing. Each worker maintains the state of its partition, executes Compute(), and manages messages between workers.
3. Each worker gets assigned a portion of the user's input by the master. If a worker loads a vertex that belongs to that worker's section of the graph, then "Aal Izz Well". Otherwise, the worker enqueues a message to the remote peer that owns the vertex.
4. The master instructs each worker to do a superstep. Compute() is called for each active vertex. Messages are sent asynchronously to overlap computation and communication. When a worker finishes, it reports to the master, telling it how many vertices will be active in the next superstep. This step iterates until all vertices are inactive.
5. The computation halts and the master may instruct each worker to save its portion of the graph.

### Fault Tolerance

At the beginning of a superstep, the master instructs the workers to save the partition states (vertex/edge values, incoming messages) to persistent storage. The checkpoint frequency is determined using a mean time to failure model. A heartbeat mechanism is used to detect failures, either for a worker to terminate or for the master to mark a worker as failed. After a failure, the master reassigns partitions to the currently available workers, each of which loads from the checkpoint.

## Example Workloads

### PageRank

![](/files/-McXeaauu9HA5Eanroyj)

### Shortest Paths

![In this example, as the receiving vertex of a message only needs to know the shortest distance (the minimum of all distances sent by neighboring vertices), a combiner can be utilized.](/files/-McXei_G8-oDBHXEUwVR)

See more examples (bipartite matching, semi-clustering) in the paper :\_D

## Evaluation

![SSSP: Single-source shortest paths. Scalability! Let's go!](/files/-McXfyA1_YKzhOzthO4T)

![In a realistic setting (log-normal random graphs), the performance also scales.](/files/-McXgEhKz3Kr1D1eWG_9)

## Links

* [Paper PDF](https://www.dcs.bbk.ac.uk/~dell/teaching/cc/paper/sigmod10/p135-malewicz.pdf)
* [Reading notes by the morning paper](https://blog.acolyer.org/2015/05/26/pregel-a-system-for-large-scale-graph-processing/)


# \[2011 NSDI] Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center

## One-line Summary

Mesos is a scheduler for sharing cluster resources between multiple services/applications. Using two-level scheduling (application-specific schedulers), its orchestration mechanism can provide scalable, decentralized scheduling.

## Paper Structure Outline

1.

## Background & Motivation

* Motivation: Share resources among multiple frameworks/services
* Background: OS scheduling: time sharing
  * Mechanism: Pre-empt processes
  * Policy: Which process is chosen to run
* Background: Cluster scheduling
  * Time sharing: Context switches are more expensive
  * Space sharing: Partition resources at the same time across jobs
  * Policies: Should be aware of locality
  * Scale
  * Fault tolerance
* Background: Target environment
  * Multiple MapReduce versions
  * Mix of frameworks: MPI, Spark, MapReduce
  * Data sharing across frameworks
  * Avoid per-framework clusters (not good for resource utilization)

## Design and Implementation

### Architecture

![Mesos architetcture](/files/-Mkgs62LyB_pvNzcLalE)

* Centralized master (& its backups)
* Agent on every machine
* Two-level scheduling: Framework scheduler attached to Mesos
  * "Simplicity across frameworks"

### Resource offers

![Resource offers](/files/-MkgtAk8GO8oCQoxSTGe)

* Slave send heartbeats with information on available resources
* Mesos master sends resource offers to frameworks -> Frameworks replies tasks & their granularity
* Constraints
  * Example of constraints
    * Soft constraints: Prefer to run the task at a particular location (e.g. for data locality)
    * Hard constraints: Task needs GPUs
  * Constraints in Mesos
    * Applications can reject offers
    * Optimization: Filters (reduces the number of rejected offers)
* Allocation
  * Assumption: Tasks are short -> allocate when they finish
  * Long tasks: Revocation beyond guaranteed allocation
    * E.g., MPI has guaranteed allocation of 8 machines; currently assigned 12 machines; can take away 4 machines
* Isolation
  * Containers (Docker/Linux cgroups)

### Fault tolerance

* Node fails -> forward the failure to Hadoop, let them decide what to do
* Master fails -> recover (soft) state by communicating with framework schedulers/workers
* Also has a standby master
* Framework scheduler fails -> Mesos doesn't handle that

### Placement preferences

* Problem: More frameworks have preferred nodes than available. Who gets the offers?
* Lottery scheduling: Offers weighted by num allocations

### Design choices & implications

* Centralized vs. Distributed
  * Framework complexity: Every framework developer needs to implement a scheduler
  * Fragmentation, starvation: Especially if a job has large resource requirements. Partial workaround: min offer size
  * Inter-dependent framework: 2 frameworks cannot be colocated (e.g. due to security, privacy, ...)

## Evaluations

![Excerpted from Shivaram's CS 744 course notes](/files/-Mki2NIz47ybIGGVkN_H)

## Links & References

* [Paper PDF](http://pages.cs.wisc.edu/~shivaram/cs744-readings/mesos.pdf)
* [CS 744 course notes](https://pages.cs.wisc.edu/~shivaram/cs744-fa21-slides/cs744-mesos-notes.pdf)


# \[2012 NSDI] Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster ...

...Computing

## One-line Summary

Spark is a generalized MapReduce model (in-memory Hadoop :P). It uses RDDs to support in-memory computations.

## Paper Structure Outline

1. Introduction
2. Resilient Distributed Datasets (RDDs)
   1. RDD Abstraction
   2. Spark Programming Interface
      1. Example: Console Log Mining
   3. Advantages of the RDD Model
   4. Applications Not Suitable for RDDs
3. Spark Programming Interface
   1. RDD Operations in Spark
   2. Example Applications
      1. Logistic Regression
      2. PageRank
4. Representing RDDs
5. Implementation
   1. Job Scheduling
   2. Interpreter Integration
   3. Memory Management
   4. Support for Checkpointing
6. Evaluation
   1. Iterative Machine Learning Applications
   2. PageRank
   3. Fault Recovery
   4. Behavior with Insufficient Memory
   5. User Applications Built with Spark (In-mem analytics, traffic modeling, twitter spam classification)
   6. Interactive Data Mining
7. Discussion
   1. Expressing Existing Programming Models
   2. Leveraging RDDs for Debugging
8. Related Work
9. Conclusion

## Background & Motivation

* Programmability
  * Most real applications require multiple MapReduce stages:
    * Google indexing pipeline: 21 steps
    * Analytics queries: 2-5 steps
    * Iterative algorithms: 10s of steps
  * Multi-step jobs create spaghetti code
    * 21 MapReduce steps -> 21 mapper & reducer classes
* Performance
  * MapReduce only provides one pass of computation (must write data to file system in between)
  * Expensive for apps that need to reuse data (e.g., multi-step algorithms like PageRank, interactive data mining)

## Design and Implementation

### Apache Spark

* Simple, functional API
  * 5x - 10x less code than MapReduce
  * Parallel transformations on collections
  * Available in Scala, Python, Java, and R
  * Can trace how operations are chained -> free type checking!
  * Mimics local programs
* Performance
  * In-memory computing primitives
    * In-mem caching: LRU
    * Caches also get cleared when workers go away
  * Optimization across operators
* Lazy evaluation/execution
* Fault tolerance
  * Lineage graphs: Records of transformations that created this RDD
    * Assumption: Input file is still available
    * Assumption: Storage for lineage graphs is stable (similar to the MapReduce master)
* Job scheduling
  * Captures RDD dependency graph
  * Pipelines function into "stages"
  * Cache-aware for data reuse, locality (move computation to where data is cached)
  * Partition-aware to avoid shuffles

### RDDs

* Immutable, partitioned collection of objects
* Can be cached in memory for faster reuse
* Operations on RDDs: Transformations (build RDDs) & Actions (compute results)

## Links & References

* [Paper PDF](https://www.usenix.org/system/files/conference/nsdi12/nsdi12-final138.pdf)
* [CS 744 course notes](https://pages.cs.wisc.edu/~shivaram/cs744-fa21-slides/cs744-spark-notes.pdf)


# \[2012 OSDI] PowerGraph: Distributed Graph-Parallel Computation on Natural Graphs

## One-line Summary

The key contributions are:

* GAS (gather, apply, scatter) programming model
* Using vertex cut instead of edge cut to layout data for power-law graphs
* Balancing computation & minimizing communication

## Paper Structure Outline

1. Introduction
2. Graph-Parallel Abstractions
   1. Pregel
   2. GraphLab
   3. Characterization
3. Challenges of Natural Graphs
4. PowerGraph Abstraction
   1. GAS Vertex-Programs
   2. Delta Caching
   3. Initiating Future Computation
      1. Bulk Synchronous Execution
      2. Asynchronous Execution
   4. Comparison with GraphLab/Pregel
5. Distributed Graph Placement
   1. Balanced p-way Vertex-Cut
   2. Greedy Vertex-Cuts
6. Abstraction Comparison
   1. Computation Imbalance
   2. Communication Imbalance
   3. Runtime Comparison
7. Implementation and Evaluation
   1. Graph Loading and Placement
   2. Synchronous Engine (Sync)
   3. Asynchronous Engine (Async)
   4. Async. Serializable Engine (Async+S)
   5. Fault Tolerance
   6. MLDM Applications
8. Related Work
9. Conclusions and Future Work

## Background & Motivation

* **Background 1: Natural Graphs**
  * Graphs IRL (e.g., social networks/the Internet) follow a power-law degree distribution
    * A small subset of the vertices have very high degrees, while most vertices have a small degree
  * Existing graph-parallel frameworks depend on a balanced degree distribution for performance

![](/files/-MfC1aHOOPjceGHZpUkt)

* **Background 2: Existing frameworks (**[**Pregel**](/machine-learning-systems/index/pregel-a-system-for-large-scale-graph-processing)**, GraphLab) cannot handle natural graphs well**
  * Work balancing: Existing graph-parallel frameworks treat vertices symmetrically and have storage/communication/computation costs linear in degree
  * Partitioning: Pregel/GraphLab depends on partitioning the graph, which is hard to do in natural graphs. Their solution, random partitioning, is bad.
  * Communication/storage: Major bottlenecks at high-degree vertices due to the skewed distribution
  * Computation: Existing frameworks do not parallelize individual vertex programs, limiting their scalability in skewed graphs

## **Design and Implementation**

### GAS Model

* Gather: Information from adjacent vertices/edges is reduced by a generalized "sum" operation (commutative and associative)
* Apply: The gathered sum is used with the current value to update the current vertex value
* Scatter: The new value is used to update data on adjacent edges

### Vertex-Cuts instead of Edge-Cuts

* Edge-Cuts
  * Every vertex is placed on a machine, and edges span across machines
    * If adjacent vertices are on different machines, they use "ghost" vertices -> changes need to be synchronized to ghosts
  * In natural graphs, there are lots of edges spanned across machines; Balanced edge-cut algorithms perform poorly, so GraphLab and Pregel uses randomized placement (bad)
* Vertex-Cuts
  * Every edge is placed on a machine, and vertices may be across machines
    * Intuition: The distribution of vertex degree is highly skewed, but the number of vertices adjacent to a given edge is constant (always 2)
    * Each vertex is replicated ("mirrors") across the machines where its adjacent edges lie
  * This results in a better balance for natural graphs

![](/files/-MfC5F7SzjYGA5qkLej7)

![](/files/-MfC7q3GUHD-LNQ5uJlU)

### Delta Caching, Execution Model

* Delta caching
  * At each vertex, the accumulator values are cached, and the scatter function can return a delta value to directly apply to the neighboring cached accumulator.
  * If this value is not returned, the neighboring cache is cleared
* Execution model: Sync vs. Async
  * Sync (bulk synchronous)
    * 3 "minor-steps": Gather for all active vertices -> Apply -> Scatter
    * Barrier after each minor-step; Changes are committed at the end of each minor-step and visible on the next
  * Async (asynchronous)
    * Changes are immediately available to other vertices
    * Execute active vertices as cores become available

## Evaluation

### Reduced vertex replication/communication costs

![](/files/-MfC6nRKYHynIIq5RpGW)

![](/files/-MfCClRTN6Od7VAcQOI1)

![](/files/-MfCDI7jxtr-2_R7HG8U)

## Links

* [Paper PDF](https://www.usenix.org/system/files/conference/osdi12/osdi12-final-167.pdf)
* [Presentation slides by 6.886 @ MIT](https://people.csail.mit.edu/jshun/6886-s20/lectures/lecture11-2.pdf)
* [Presentation slides by CS 744 @ UW-Madison](http://pages.cs.wisc.edu/~shivaram/cs744-fa20-slides/cs744-powergraph-notes.pdf)
* [PowerGraph on GitHub](https://github.com/jegonzal/PowerGraph)


# \[2019 FAST] DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed...

...Caching

## Summary

This paper presents a new distributed caching mechanism for addressing load imbalance in large-scale storage systems.

![](/files/S2bJKoQjJTHuLPnaQZ6B)

## Background & Motivation

Cloud service providers use large clusters to store data. The data access workload is skewed (power law distribution), which creates load imbalance, resulting in low throughput and long tail latencies. The objective is to achieve load balancing in distributed storage systems.

![](/files/HQxKs4t3gkgKuJFriwVw)A common approach is to add a front-end cache node as a load balancer.&#x20;

![](/files/MV6xGMUdhvAwZzhaeMOl)The problem is that nowadays, cloud-scale distributed storage spans across many clusters, which exposes scalability issues. Given that the cache throughput is 10-100 times of the server throughput, one caching node (e.g., a switch) can only guarantee load balancing for 10-100 servers (a few racks of servers within a cluster). In other words, a single cache node only guarantees intra-cluster load balancing, not inter-cluster load balancing.

![](/files/yBzai3w2ZWFmU5xP0A69)Adding one cache node as the load balancer within each cluster also doesn't work: between clusters, load imbalance still exists. Adding another cache node atop all the per-cluster cache nodes does not work due to the throughput constraint.

Thus, we need a layer of distributed caching as the load balancer.

## Design & Implementation

Some key design choices include:

* Cache allocation with independent hash functions: The intuition is that if one cache node in a layer is overloaded by receiving too many queries to its cached objects, because the hash functions of the two layers are independent, the set of hot objects would be distributed to multiple cache nodes in another layer with high probability.
* Query routing with the power-of-two-choices: The sender of a query looks at the loads of the cache nodes that cache the queried object and sends the query to the less-loaded node.

![](/files/Q2PchK1SHFilkSiyR6rO)

These mechanisms can be applied recursively for multi-layer hierarchical caching.

## Evaluation

![](/files/XvpjAJqwjyMTTX3f4SI8)

## Links & References

* [Paper PDF](https://www.usenix.org/system/files/fast19-liu.pdf)
* [Presentation video at FAST '19](https://www.youtube.com/watch?v=iLsBC1yjH40)
* [Presentation slides at FAST '19](https://www.usenix.org/sites/default/files/conference/protected-files/fast19_slides_liu.pdf)


# \[2021 HotOS] From Cloud Computing to Sky Computing

## One-line Summary

This paper envisions sky computing, the possible future, and a more commoditized version of cloud computing, by drawing lessons from the history of the Internet. It then introduces the technical/economical barriers of fulfilling this vision of utility computing.

![](/files/-Maoh0AYZS6vP1-vCKTA)

## Paper Structure Outline

1. Introduction
2. Historical Context
3. Lessons from the Internet
4. Compatibility Layer
5. Intercloud Layer
6. Peering Between Clouds
7. Speculations about the Future
8. Conclusion

## Background & Motivation

> Computation may someday be organized as a public utility, just as the telephone system is a public utility. We can envisage computer service companies whose subscribers are connected to them \[...]. Each subscriber needs to pay only for the capacity that he actually uses, but he has access to all programming languages characteristic of a very large system.”    -- John McCarthy on the future of computing, 1961

Currently, from the user's point of view, many of the cloud computing services (AWS, Microsoft, Google, etc.) are proprietary/differentiated (e.g., APIs for cluster management, object store, data warehouse, serverless offering), and thus applications developed on one cloud cannot be easily migrated to another.

From the provider's point of view, business models are built around "attracting and retaining customers", which goes against the idea of offering a purely commoditized service.

The benefits of sky computing are:

* New capabilities: If one cloud in the sky provides access to new hardware (e.g., TPU), any app in the sky can use it
* Better security: Eliminate a single point of attack by distributing trust across multiple clouds
* Better reliability: Avoids major cloud outages
* Better performance: Aggregates all resources to use the best resources for a job
* Lower cost: Use most cost-effective cloud for a job

## Design and Implementation

To fulfill the sky computing vision, three design issues (the Internet also faced them) must be addressed:

* Compatibility layer: Mask low-level technical differences/heterogeneity
* Intercloud/Routing layer: Route jobs to the right cloud
* Peering layer: Allow clouds to have agreements with each other about how to exchange services

![](/files/-Maikf2xoXza4YQBMEz-)

### Compatibility Layer

Similar to the IP layer, a compatibility layer abstracts away the services provided by a cloud and allows an application developed on top of this layer to run on different clouds without change. The authors conclude that this is not technically difficult, as the high-level management and service interfaces users interact with are now more than ever supported by open source software (OSS). The compatibility layer could be constructed out of some set of the OSS solutions. One glaring gap is the storage layer (AWS has S3, Azure has Blob storage, etc.), but there are currently efforts underway to provide more compatibility and fill this gap.

OSS projects for different levels of the software stack include:

* OS: Linux
* Cluster resource managers: Kubernetes, Apache Mesos
* Application packaging: Docker
* Databases: MySQL, Postgres
* Big data execution engines: Apache Spark, Apache Hadoop
* Streaming engines: Apache Flink, Apache Spark, Apache Kafka
* Distributed query engines and db: Cassandra, MongoDB, Presto, SparkSQL, Redis
* ML libraries: PyTorch, Tensorflow, MXNet, MLFlow, Horovod, Ray RLlib
* General distributed frameworks: Ray, Erlang, Akka

![Status quo: multi-cloud, porting from one cloud to another is expensive](/files/-MaoiJLouu3HDysnpE2L)

![What sky computing will bring](/files/-MaoiZKuuWLhhjOJ3_zv)

### Intercloud Layer

The intercloud layer should allow users to specify policies describing the tradeoff between performance, availability, and cost (e.g., a user might specify that this is a Tensorflow job, it involves data that cannot leave Germany, and must be finished within the next two hours for under a certain cost), but not require users to make low-level decisions. There should be few technical limitations as this (moving jobs across clouds) is similar to moving jobs within the same cloud across datacenters.

### Peering Layer

Under certain scenarios, instead of processing all data in the same cloud, moving data between clouds can be cost-effective. For example, although moving a 150 GB ImageNet dataset out of AWS costs $13, training ResNet50 on ImageNet on AWS costs \~$40, while training on Azure costs $20. If clouds adopt reciprocal data peering arrangements, it allows data to be moved freely between peering clouds and enables greater freedom in job movement.

### Speculations about the Future

The authors' vision is as follows. While large providers may not be incentivized to build a compatibility layer, smaller cloud providers will embrace such a layer and form a sky. Within the sky, providers may specialize in supporting one or more services. E.g., Oracle can provide a database-optimized cloud, NVIDIA can provide GPU-optimized, hardware-assisted ML services, Samsung can provide a storage-optimized cloud. In the long term, both the standalone providers and in-sky providers will exist: the standalone providers compete with each other and the sky, and the in-sky providers both compete within the sky and collectively compete with the standalone providers.

## Links

* [Paper PDF](https://sigops.org/s/conferences/hotos/2021/papers/hotos21-s02-stoica.pdf)
* [Presentation video at HotOS '21](https://www.youtube.com/watch?v=Q6MsEucsmGM\&list=PLl-7Fg11LUZe_6cCrz6sVvTbE_8SEobNB)


# \[2021 EuroSys] NextDoor: Accelerating graph sampling for graph machine learning using GPUs

## One-line Summary

In Graph Neural Network (GNN) training, existing approaches use CPUs to sample the graph before using GPUs to train the GNN, but sampling is a major overhead (up to 62% of training time). Nextdoor uses GPUs to accelerate graph sampling by up to 4x, and its main contributions are:

1. Simple abstractions & API to express diverse graph sampling algorithms
2. A new "transit parallel" approach to increase the parallelism of graph sampling
3. Optimizations (load balancing & caching) to improve GPU utilization

![Nextdoor structure](/files/-MhYVzDfNa4x41KSZDeG)

Takeaways from Shivaram's group meeting after discussing this paper include:

* In evaluations, except from the relative numbers, post absolute values as well
* Parallelization works very well on GPUs, more CPU-based tasks may be identified and transformed into GPU-based tasks with high levels of parallelization
* Nextdoor introduces this abstraction/API that "bounds" graph sampling algorithms so that they can be properly parallelized

## Paper Structure Outline

1. Introduction
2. Background and Motivation
   1. Representation Learning on Graphs
   2. Requirements for GPU Performance
3. An Abstraction for Graph Sampling
4. Graph Sampling using NEXTDOOR
   1. Programming API
   2. Use Cases
5. Paradigms for Graph Sampling on GPUs
   1. Sample-Parallelism
   2. Transit-Parallelism
6. Efficient Transit Parallelism on GPUs
   1. Sampling in Individual Transit Sampling
   2. Transit-Parallel Collective Transit Sampling
   3. Unique Neighbors
   4. Graph Sampling using Multiple GPUs
   5. Integration in GNNs using Python API
   6. Advantages of NEXTDOOR's API
7. Alternative Graph Processing Systems
8. Evaluation
   1. Execution Time Breakdown
   2. Graph Sampling Performance
   3. Alternative GPU-Based Abstractions
   4. Sampling Large Graphs
   5. Sampling on Multiple GPUs
   6. End-to-End Integration in GNN Systems
9. Related Work
10. Conclusion

## Background & Motivation

### Background 1: How GNN training works

* GNNs maps vertices of (input) graphs to an embedding in an N-dimensional space so that the similarity of embeddings between nodes indicate their similarity in the network
  * The embeddings are then used for many downstream tasks (e.g., product recommendation, clustering)
* There are two types of GNNs, and this work focuses on the first one (they're more common):
  * Sampling-based GNNs samples the input graph and train using these samples
  * Whole-Graph-based GNNs train on the whole input graph directly
* Workflow of Sampling-based GNNs: First, a graph sampling algorithm is used to sample the input graph, and the samples are then used for data parallel training
* Currently, most implementations use CPUs for sampling, because the implementation is easier

![Different graph sampling algorithms](/files/-MhYPY4xj5B0kwuvYcdv)

### Background 2: How to best utilize GPUs

* Level of parallelism should be high (in GPU computing, # threads == # samples)
* Accesses to the global memory should be coalesced and aligned
* Shared memory and registers for each SM can be used as software-managed cache
* Avoid warp divergence

![](/files/-MhYPjwoBvA1a1zb5u14)

![](/files/-MhYSZs6W_amLI-iZbmt)

### Motivation: Graph sampling on CPUs is a major overhead

![Existing implementations spend as much as 62% of the training time on graph sampling](/files/-Mgy1o-q2Fnu6hwrJnYy)

Currently, graph sampling is done on CPUs because of the ease of implementation. Nextdoor attempts to provide both easy-to-implement and fast graph sampling.

## Design and Implementation

### Powerful abstraction/API to express sampling algorithms

* Input to Nextdoor
  * A graph
  * An initial set of samples, each with >=1 root vertices
  * User-defined functions to describe the algorithm
* Output of Nextdoor: An expanded set of examples
* Nextdoor abstractions:
  * A graph sampling appplication runs for k steps
  * At each step i,
    * A transit vertex for i is a vertex whose neighbors may be added to the sample
    * Sample mi of those neighbors
  * There are two types of sampling:
    * Individual transit sampling: Sample mi neighbors per-transit-node
    * Collective transit sampling: Sample mi neighbors per-sample

![Example algorithms expressed using this abstraction](/files/-MhYRNy9qbjxgQ5d0wJ9)

![Required user-defined functions](/files/-MhYRZiQ69OHlO63DIVJ)

![Use cases of Nextdoor](/files/-MhYRin3Zz16YU5nEQvz)

### Transit parallel to increase parallelism

* Status quo: One thread for each sample -> poor parallelism. How can we increase the parallelism?
* Sample parallel: In each thread, one neighbor of a transit vertex is sampled, and samples are assigned to consecutive threads
  * The parallelism is better
  * However, sample parallel suffers from irregularity: The access to the global memory is random, and shared memory/registers cannot be used as caches
* Transit parallel: Assign samples with common transits to consecutive threads
  * A GroupBy operation is needed to invert the sample-transit mapping to a transit-sample mapping
  * Here, consecutive threads access edges of the same transit vertices. Therefore, the global memory accesses are coalesced, and shared memory/registers can be used for caches
* The Nextdoor API exposes three levels of parallelism
  1. Each transit is mapped to a threadblock
  2. Each sample is assigned to a group of mi threads at step i
  3. Each thread samples one neighbor

![Sample parallel](/files/-MhYU947A7ZC9Rp1BDzS)

![Transit parallel](/files/-MhYUE9uSJmGWGRipCSr)

### Optimization techniques for GPUs (load balancing, caching)

Nextdoor uses different {types of kernels, caching strategies, neighbor access strategy, transit scheduling strategy} to process transit vertices based on the number of neighbors to sample for the transit vertex, which helps to best utilize the memory/compute resources

![Sub warp: A set of contiguous threads of the same warp assigned to the same sample. All sub warps have the same size, which is determined using sampleSize function for the current step.](/files/-MhYUga2_S40n1tXl3pa)

## Evaluation

![End-to-end speedups for GNN training](/files/-MhYVKfBP_0XlKgU-243)

![Nextdoor against existing graph sampling implementations](/files/-MhYVTIES0UGsBTdM4VN)

The original paper also included some microbenchmarks of speedups of SP, TP, and the overhead of the GroupBy operation, etc.

## Links & References

* [Paper PDF](https://marcoserafini.github.io/projects/nextdoor/nextdoor.pdf)
* Presentation video at EuroSys '21 ([Long](https://www.youtube.com/watch?v=GsffY0j6tVE\&list=PLzDuHU-z7gNjuSbEYCFXZtWAl3nAdNF2f\&index=19) & [Short](https://www.youtube.com/watch?v=lwB7KcMIpkQ\&list=PLzDuHU-z7gNghxOWGcdLK_xWtqHjxaYTm\&index=19))
* [Presentation slides at EuroSys '21](https://2021.eurosys.org/docs/presentations/6-Jangda%20-%20Abhinav%20Jangda.pdf)
* Graph sampling algorithms referenced in Nextdoor
  * [DeepWalk](https://arxiv.org/pdf/1403.6652.pdf)
  * [node2vec](https://arxiv.org/pdf/1607.00653.pdf)
  * [GraphSAGE](https://arxiv.org/pdf/1706.02216.pdf)
  * [FastGCN](https://arxiv.org/pdf/1801.10247.pdf)
  * [ClusterGCN](https://arxiv.org/pdf/1905.07953.pdf)
  * [LADIES](https://arxiv.org/pdf/1911.07323.pdf)


# High Performance Computing Course Notes

ECE/ME/EMA/CS 759: High Performance Computing for Engineering Applications, Spring 2021 by Prof. Dan Negrut

## Acknowledgments

* All slides/files linked are accessible on Box using a UW-Madison account
* Almost every figure and piece of code in these notes is excerpted from Prof. Dan Negrut's course slides. Some of the slides are taken from other places by Prof. Negrut -- he cited those in his slides.
* [Slides for ME759 (of the whole semester)](https://uwmadison.app.box.com/s/oboe3t95di8rne0g002ydj8tpd0pwwkt)
* [Slides from ME459 (Computing Concepts for Applications in Engineering)](https://uwmadison.app.box.com/s/943jyv29y4u145uajfedgxamhn4ru9qx)

## Table of Contents

| Date | Title                                                                                                                                                                                                                                                         | Recommended Readings                                                                                                                                                                                                                                     |
| ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1/25 | [Lecture 1: Course Overview](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-1-course-overview)                                                                                                                                                    | [Basic Linux Command Line Usage](https://www.lynda.com/Linux-tutorials/Learning-Linux-Command-Line/753913-2.html); Slurm usage (ME459 p95-97)                                                                                                            |
| 1/27 | [Lecture 2: From Code to Instructions. The FDX Cycle. Instruction Level Parallelism.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-2-from-code-to-instructions.-the-fdx-cycle.-instruction-level-parallelism.)                                  | C recap (ME459 p114-); [Euler usage](https://uwmadison.app.box.com/s/eu45vz9uc1a913i831b1saiu554ueb4z)                                                                                                                                                   |
| 1/29 | [Lecture 3: Superscalar architectures. Measuring Computer Performance. Memory Aspects.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-3-superscalar-architectures.-measuring-computer-performance.-memory-aspects.)                              | gdb recap (ME459 p649-); Ch.5 of the [C book](https://www.amazon.com/Programming-Language-2nd-Brian-Kernighan/dp/0131103628)                                                                                                                             |
| 2/1  | [Lecture 4: The memory hierarchy. Caches.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-4-the-memory-hierarchy.-caches.)                                                                                                                        | Build mgmt & cmake (ME459 p354-)                                                                                                                                                                                                                         |
| 2/3  | [Lecture 5: Caches, wrap up. Virtual Memory.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-5-caches-wrap-up.-virtual-memory.)                                                                                                                   | Git (ME459 p449-); [How to Write a Git Commit](https://chris.beams.io/posts/git-commit/)                                                                                                                                                                 |
| 2/5  | [Lecture 6: The Walls to Sequential Computing. Moore’s Law.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-6-the-walls-to-sequential-computing.-moores-law.)                                                                                     | [Validity of the single processor approach to achieving large scale computing capabilities (Amdahl, '67)](https://uwmadison.app.box.com/s/z21zx63u3n3swxk23luex6mv0n9az96a)                                                                              |
| 2/8  | [Lecture 7: Parallel Computing. Flynn’s Taxonomy. Amdahl’s Law.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-8-parallel-computing.-flynns-taxonomy.-amdahls-law.)                                                                              | [Structured Programming w/ go to Statements (Knuth, '74)](https://uwmadison.app.box.com/s/40oh4cw2j0tlouf6ip6fqv2guegrn6rz)                                                                                                                              |
| 2/10 | [Lecture 8: GPU Computing Intro. The CUDA Programming Model. CUDA Execution Configuration](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-8-gpu-computing-intro.-the-cuda-programming-model.-cuda-execution-configuration)                        | [Modern Microprocessors: A 90-Minute Guide (Patterson, '01)](http://www.lighterra.com/papers/modernmicroprocessors/)                                                                                                                                     |
| 2/12 | [Lecture 9: GPU Memory Spaces.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-9)                                                                                                                                                                 | [Optimizations in C++ Compilers (Godbolt, 2019)](https://queue.acm.org/detail.cfm?id=3372264)                                                                                                                                                            |
| 2/15 | [Lecture 10: GPU Scheduling Issues.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-10-gpu-scheduling-issues.)                                                                                                                                    | [NVIDIA Tesla Architecture](https://uwmadison.app.box.com/s/c3j9jiy6feq31qh4nuuli9ce40xebvlb)                                                                                                                                                            |
| 2/17 | [Lecture 11: Execution Divergence. Control Flow in CUDA. CUDA Shared Memory Issues.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-11-execution-divergence.-control-flow-in-cuda.-global-memory-access-patterns-and)                             | [CUDA C++ Programming Guide](https://docs.nvidia.com/pdf/CUDA_C_Programming_Guide.pdf)                                                                                                                                                                   |
| 2/19 | [Lecture 12: Global Memory Access Patterns and Implications.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-12-cuda-shared-memory-issues.)                                                                                                       | [The GPU Computing Era (Nickolls & Dally, '10)](https://uwmadison.app.box.com/s/o63jve7gq6kn9f473btedx2k4tp33m34)                                                                                                                                        |
| 2/22 | [Lecture 13: Atomic operations in CUDA. GPU ode optimization rules of thumb.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-12-cuda-shared-memory-issues.-atomic-operations-in-cuda.)                                                            | [Unified Memory in CUDA 6: A Brief Overview](https://www.drdobbs.com/parallel/unified-memory-in-cuda-6-a-brief-overvie/240169095)                                                                                                                        |
| 2/24 | [Lecture 14: CUDA Case Studies. (1) 1D Stencil Operation. (2) Vector Reduction in CUDA](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-14-tiling-as-a-programing-pattern-in-cuda.-example-vector-reduction-in-cuda.)                              | [Maximizing Unified Memory Performance in CUDA (Sakharnykh, '17)](https://developer.nvidia.com/blog/maximizing-unified-memory-performance-cuda/)                                                                                                         |
| 2/26 | [Lecture 15: CUDA Case Studies. (3) Parallel Prefix Scans on the GPU. Using Multiple Streams in CUDA.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-15-cuda-optimization-issues.-resource-utilization-issues.-parallel-prefix-scan-on-the-gpu.) | [Titles of GTC '21 Talks](https://www.nvidia.com/en-us/gtc/on-demand/)                                                                                                                                                                                   |
| 3/1  | [Lecture 16: Streams, and overlapping data copy with execution.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-16-streams-and-overlapping-data-copy-with-execution.)                                                                             | [Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking (Citadel, '18)](https://uwmadison.app.box.com/s/qdmt5f9qxpnbx431t2oo7neh6a7ri6zs); [CUDA C++ Best Practices Guide](https://docs.nvidia.com/cuda/pdf/CUDA_C_Best_Practices_Guide.pdf) |
| 3/3  | [Lecture 17: GPU Computing: Advanced Features.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-17-gpu-computing-advanced-features.-unified-memory-usage.)                                                                                         | [GTC '18 Talk on Unified Memory](https://on-demand.gputechconf.com/gtc/2018/video/S8430/)                                                                                                                                                                |
| 3/5  | [Lecture 18: GPU Computing with thrust and cub.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-18-gpu-computing-with-thrust-and-cub.)                                                                                                            | [Thrust: A Productivity-Oriented Library for CUDA (Bell & Hoberock, '11)](https://uwmadison.app.box.com/s/5gdq2gaqf15xjl1cd550ttsko782fbaz)                                                                                                              |
| 3/8  | [Lecture 19: Hardware aspects relevant in multi-core, shared memory parallel computing.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-19-hardware-aspects-relevant-in-multi-core-shared-memory-parallel-computing.)                             | [Unified Memory in CUDA 6: A Brief Overview and Related Data Access/Transfer Issues (by Dan and some other guys! '14)](https://sbel.wisc.edu/wp-content/uploads/sites/569/2018/05/TR-2014-09.pdf)                                                        |
| 3/10 | [Lecture 20: Multi-core Parallel Computing with OpenMP. Parallel Regions.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-20-multi-core-parallel-computing-with-openmp.-parallel-regions.)                                                        | Cache Coherence on Power 9 - Volta systems w/ NVLINK2                                                                                                                                                                                                    |
| 3/12 | [Lecture 21: OpenMP Work Sharing.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-21-openmp-work-sharing.)                                                                                                                                        | [Node-Level Performance Engineering (SC '19)](https://uwmadison.app.box.com/s/cvva3ybaq0867e160hqf6l92yp9gbthx)                                                                                                                                          |
| 3/15 | [Lecture 22: OpenMP Work Sharing.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-22-openmp-work-sharing)                                                                                                                                         | [Advanced OpenMP: Performance and 5.0 Features (SC '19)](https://uwmadison.app.box.com/s/dftxq7z83u6bc1e33lbihhcuvb5ek2uc)                                                                                                                               |
| 3/17 | [Lecture 23: OpenMP NUMA Aspects. Caching and OpenMP.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-23-openmp-numa-aspects.-caching-and-openmp.)                                                                                                | [Mastering Tasking with OpenMP (SC '19)](https://uwmadison.app.box.com/s/40yxdvu41prgvfuzvv668lhx78eafos0)                                                                                                                                               |
| 3/19 | [Lecture 24: Critical Thinking. Code Optimization Aspects.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-24-critical-thinking.-code-optimizatino-aspects.)                                                                                      | [Ch. 12 of Optimizing Software in C++](https://www.agner.org/optimize/optimizing_cpp.pdf)                                                                                                                                                                |
| 3/22 | [Lecture 25: Computing with Supercomputers.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-25-computing-with-supercomputers.)                                                                                                                    |                                                                                                                                                                                                                                                          |
| 3/24 | [Lecture 26: MPI Parallel Programming General Introduction. Point-to-Point Communication.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-26-mpi-parallel-programming-general-introduction.-point-to-point-communication.)                        | [HPC Perspectives (Dongarra et. al., '05)](https://uwmadison.app.box.com/s/fi2h0s0d4rgvviepc1dd92m9jqoemq86)                                                                                                                                             |
| 3/26 | [Lecture 27: MPI Parallel Programming Point-to-Point communication: Blocking vs. Non-blocking sends.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-27-mpi-parallel-programming-point-to-point-communication-blocking-vs.-non-blocking-sends.)   | [Advanced MPI Programming (SC '19)](https://uwmadison.app.box.com/s/ymhcrw7xc49u3cvs86jeva9bnek99sfe)                                                                                                                                                    |
| 3/29 | [Lecture 28: MPI Parallel Programming: MPI Collectives. Overview of topics covered in the class.](/earlier-readings-and-notes/cs759-hpc-course-notes/lecture-28-mpi-parallel-programming-mpi-collectives.-overview-of-topics-covered-in-the-class.)           |                                                                                                                                                                                                                                                          |


# Lecture 1: Course Overview

## Course Description

This grad-level course seeks to:&#x20;

1. Provide an overview of various advanced computing software and hardware solutions
2. Introduce **CUDA** for parallel computing on the Graphics Processing Unit (GPU)
3. Introduce the **OpenMP** solution to enabling parallelism across multiple CPU cores
4. Introduce the Message Passing Interface (**MPI**) standard for leveraging parallelism on a CPU cluster
5. Promote an understanding instrumental in deciding what parallel computing model is suitable for which problems.

## Linux "module" utility

[Linux man page](https://linux.die.net/man/1/module)

{% code title="Linux module usage" %}

```bash
[dan@euler ~]$ gcc --version
gcc (GCC) 4.8.5 20150623 (Red Hat 4.8.5-16)
...
[dan@euler ~]$ module load gcc/6.4.0
[dan@euler ~]$ gcc --version
gcc (GCC) 6.4.0
...

[dan@euler ~]$ nvcc main.cu -o cudaprogram
bash: nvcc: command not found
[dan@euler ~]$ module avail cuda

--------------/usr/local/share/modulefiles ---------------------
cuda/0_user/cuda  cuda/7.5  cuda/8-rc  cuda/9    cuda/9.1  
cuda/7            cuda/8    cuda/8.0   cuda/9.0

[dan@euler ~]$ module load cuda/9
[dan@euler ~]$ nvcc main.cu -o cudaprogram

[dan@euler ~]$ module list
Currently Loaded Modulefiles:
1) gcc/6.4.0   2) gcc/0_cuda/6.4.0   3) cuda/9
[dan@euler ~]$ module unload cuda gcc
[dan@euler ~]$ nvcc
bash: nvcc: command not found
```

{% endcode %}

## The Euler cluster

* Files on the Euler remote cluster can be easily edited using the [Remote-SSH plugin for VS Code](https://marketplace.visualstudio.com/items?itemName=ms-vscode-remote.remote-ssh)

## Slurm (Simple Linux Utility for Resource Management)

Slurm is used on Euler for job management and scheduling.

Slurm usage (SBATCH flags documentation) can be found [here](https://slurm.schedmd.com/sbatch.html).

{% code title="Example of a Slurm-specific batch script" %}

```bash
#!/usr/bin/env bash                     # intepret file as bash script
#SBATCH --job-name=HelloScript
#SBATCH-p wacc                          # a partition is a logical chunk of cluster
#SBATCH --time=0-00:00:10
#SBATCH --output=“hello_output-%j.txt”

#SBATCH --ntasks=1 --cpus-per-task=1    # simple jobs: one core suffices
#SBATCH --nodes=2
#SBATCH --ntasks-per-node=8             # for mpi
#SBATCH --cpus-per-task=4               # multithreaded jobs

#SBATCH --gres=gpu:1                    # gres: Generic RESource
                                        # --gres=type[:model]:N
                                        # e.g., gpu:gtx1080:3or infiniband:1                                        
#SBATCH --constraint=haswell


# regular bash script
cd $SLURM_SUBMIT_DIR                    # directory where the script is submitted from

name_str=“World”
echo “Hello, $name_str!”
```

{% endcode %}

{% code title="sbatch (Slurm batch) usage" %}

```bash
[dan@euler~]$ sbatch hello_slurm.sh     # submit a scheduling script to Slurm
Submitted batch job 1975385
[dan@euler~]$ cat hello_output-1975385.txt
Hello, World!
[dan@euler~]$
```

{% endcode %}


# Lecture 2: From Code to Instructions. The FDX Cycle. Instruction Level Parallelism.

## Lecture Summary

This class is basically a recap of an "Intro to Machine Organization" class. The topics covered are instruction, assembly code, registers, CISC vs. RISC, CPU organization (CU/ALU), FDX cyle.

[CS 61C @ Berkeley (Su19)](https://inst.eecs.berkeley.edu/~cs61c/su19/) is a great class with many more in-depth slides on these topics.

## CPU Organization

![A (somewhat simplified) schematic architecture](/files/-MS3rk-1osn1681N2bsJ)

The Control Unit (CU) controls the "datapath" (i.e., the hardware collection of functional units + registers + data buses), while the Arithmetic Logic Unit (ALU) executes arithmetic and load/store operations.

## From C to Machine Code

![C code -> intermediate representation -> assembly code -> machine code/instructions](/files/-MS3txmaILdvPUr-9H6b)

The same C code leads to different assembly code using different ISAs and even using different flags during compilation. An ISA (Instruction Set Architecture) is a set of commands (e.g., sw, addiu, lw) that the CU understands. The two paradigms for ISAs are RISC (Reduced Instruction Set Computing Architecture) and CISC (Complex). The major difference is that in RISC, an instruction is encoded into a fixed set of bits (64), while CISC (e.g., Intel/AMD x86) instructions have various lengths.

### The FDX (Fetch-Decode-Execute) Cycle

* Fetch: An instruction is fetched from memory
* Decode: The string of 1s and 0s are decoded by the CU. Example: [RISC-V Green Card](https://www.cl.cam.ac.uk/teaching/1617/ECAD+Arch/files/docs/RISCVGreenCardv8-20151013.pdf)
* Execute: Once all data (operands) available, instruction is executed

![Integrated Circuits: From Transistors to Chip Microarchitecture](/files/-MS41c10THufbe9DYvPX)


# Lecture 3: Superscalar architectures. Measuring Computer Performance. Memory Aspects.

## Lecture Summary

* Registers
* ILP (Instruction Level Parallelism), focus on pipelining, also mentions OOOE and multiple-issue&#x20;
* TLP (Thread Level Parallelism), and HTT (Intel Hyper-Threading Technology) discussion
* Execution times

## Registers

* A register is a hardware asset whose role is to store information (data value/instruction)
* It's the storage type with the shortest latency (closest to CU & ALU)
* The number & size of registers used are specific to an ISA
* Take-home message: When processing one instruction, the chip needs all instruction-related operands in registers

Types of registers include:

* Instruction Register (IR): Holds the instruction that is executed
* Program Counter (PC): Holds the address of the instruction executed next
* Memory Data Register (MDR): Holds data read in from memory/produced by the ALU and waiting to be stored in memory
* Memory Address Register (MAR): Holds the address of RAM memory location where I/O data is supposed to be read in/written out
* Return Address (RA): The address where upon finishing a sequence of instructions, the execution should jump and commence with the execution of subsequent instruction
* Others include registers for:
  * Subroutine Arguments
  * Temporary Variables
  * Saved Temporary Variables
* Several other registers for handling function calls are:
  * Stack Pointer (SP): Holds an address to the top of the stack
  * Global Pointer (GP): Holds on to a global pointer that points into the middle of a 64KB block of memory in the heap that holds constants and global variables
  * Frame Pointer (FP): Holds an address that points to the beginning of the procedure frame (e.g., the previous SP before this function changed its value)

## ILP: Pipelining

* The concept of the "clock cycle": In a factory assembly line, it's the time from the moment a station takes an input to the moment the output leaves the station. In a processor, the clock cycle is the time between two ticks of the internal clock of the microprocessor/chip. The clock speed is typically measured in Hz (pulses/s).&#x20;
* The FDX cycle can be expanded to a five-stage process:
  * Fetch instruction
  * Decode instruction
  * Data access
  * Execute the operation
  * Write-back into register file
* Pipelining idea: Different stages of different instructions can be worked upon simultaneously

Consider these instructions:

```
sw $t0,  0($s2) //store what is in register $t0 at mem location  0 bytes from address in register $s2
sw $t1, 32($s2) //store what is in register $t1 at mem location 32 bytes from address in register $s2
sw $t2, 64($s2) //store what is in register $t0 at mem location 64 bytes from address in register $s2
```

![Case 1: No pipelining](/files/-MSFA5RsteyK4OsBV_TB)

![Case 2: With pipelining](/files/-MSFABnleBAtuYIPF7ar)

Ideally, in a balanced pipeline, each component/stage takes the same amount of time for completion to prevent a "bottleneck". Today, a typical pipeline depth is \~12-15 stages. Using pipelining gives a speed up (duh), and it also requires no changes on the user-code level.

Things can go south, though...

### Structural Hazards

An example is resource contention (e.g., two pipelined instructions have stages that need to use the same special register at the same time). The solutions are:

1. Commandeer a register for temporary use: Fortuitous
2. Serialize the access (introduce a bubble in the pipeline): Guaranteed to work, but introduces slowdown
3. OOOE (Out of Order Execution) performed statically at compile time or dynamically at run time: Good compromise

### Data Hazards

```
add  $t0, $t2, $t4   // $t0 = $t2 + $t4
addi $t3, $t0, 16    // $t3 = $t0 + 16 “add immediate instruction”
```

In the example above, you might think that `t0` is unavailable until the first instruction completely finishes. This is partially true: actually, `t0` becomes available after stage 3 of the pipeline (after it goes through the ALU). A solution, Intermediate Result Forwarding, makes the result in the ALU available to other stages of the pipeline right away. This is not a panacea, and occasionally we still need to do bubbling for a couple of cycles. OOOE also definitely helps here.

### Control Hazards

For instance, if there's an if statement in the C code (`if (sin(x)/x > 0.5`), we don't know the next instruction until the computation completes, which takes a few cycles. Bubbling again works, but it introduces slowdown. An alternative is to do branch prediction. There are two versions:

1. Static Branch Prediction: Always predict that the branch will not be taken and schedule accordingly (always the then branch, never the else branch). In some other cases (e.g., a do-while loop), it makes more sense.
2. Dynamic Branch Prediction: Make the branching decision based on recent history. In some cases, the accuracy rate can reach 90%.

## ILP: Multiple-Issue

```
int a, b;
float c, d;
//some code setting up a, b, c, d
a += b;
c += d;
```

In sequential computing, a multiple-issue processor core has the hardware chops to issue more than one instruction per cycle. This is another way to speed up execution. For example, in the code above, there is no dependency between updating a and c. Multiple-Issue can be done statically (predefined) or dynamically (determined at run time).

![](/files/-MSFO-Gu9k55ZFzx7Hdw)

A chip that is capable of doing multiple-issue is also called a superscalar architecture. Title card!

## ILP to TLP

![Various ILP techniques](/files/-MSFQkqQ239pBj6ZqgVn)

To wrap up, pipelining, OOOE, and multiple-issue are techniques for Instruction-Level Parallelism (ILP). These techniques work within one thread, and we can get more optimizations by going up to the thread level, TLP, where a chip executes simultaneously from different processes or different threads. Note that at this point, we are still talking about parallelism within one core, not multicore.

## HTT

* HTT is an example implementation of TLP.
* The scheduler tries to issue instructions from both processes at the same time.
* HTT allows the OS to see one physical chip as two virtual chips.
* HTT is particularly useful when running simultaneous modestly demanding processes. In HPC, if one stream of instruction saturates the memory bandwidth, then it's less useful.

![](/files/-MSFVsb7VpDtAi8_NKZV)

![When one thread stalls (due to cache miss, branch mispredict, pipeline bubbles, etc.), the other thread chimes in at the same rate as a single thread running on the core](/files/-MSFVvIVO_95PsAYUvZr)

A taxonomy for multi-threading is (Hennessey & Patterson):

* Coarse-grain multi-threading
* Fine-grain multi-threading
* Simultaneous multi-threading

To wrap up superscalar vs. TLP:

* Superscalar: Instructions associated with one PC
  * HW allows more than one instruction per cycle
  * One thread
* TLP: Instructions associated with two PCs
  * Processor handles instructions from different threads/processes


# Lecture 4: The memory hierarchy. Caches.

## Lecture Summary

* Execution times
* Memory related issues
* The memory hierarchy
* Caches

## Execution Times - Nomenclature

* Wall Clock Time: Amount of time from the beginning to the end of a program
* CPU Execution Time: Amount of time on the CPU that's dedicated to your program, requires a profiling tool to access
  * User Time: Time spent processing instructions compiled out of code generated by the user or in libraries that are directly called by user code
  * System Time: Time spent in support of the user’s program but in instructions that were not generated out of code written by the user (e.g., OS support for opening/reading a file, throwing an exception, etc.)
* Clock cycle: The length of the period for the processor clock (e.g., a 1GHz processor has a clock cycle of 1 nanosecond)
* The CPU Performance Equation: CPU Execution Time = Instruction Count \* Clock-Cycles per Instructions (CPI) \* Clock Cycle Time = Instruction Count \* Clock-Cycles per Instructions (CPI) / Clock Rate

![The SPEC CPU benchmark. CPI<1: Multiple-issue is in play. For combinational optimization, there are probably a lot of pipeline stalls](/files/-MSTe-MQNFNzSDIsyTCN)

## Memory & Cache

* SRAM (Static Random Access Memory): Expensive but fast (short access time), bulky, transistor hog, needs no refresh
* DRAM (Dynamic \~): Cheap but slow, information stored as a charge in a capacitor, higher capacity per unit area, needs refresh every 10-100ms, sensitive to disturbances

![](/files/-MSU96h6oBJjJDyh3113)

![](/files/-MSU9jbjU9RjdSz3xbQa)

The memory hierarchy (the pyramid of tradeoffs):

* A dedicated hardware asset called MMU (Memory Management Unit) is used to manage the hierarchy
* Tradeoff:
  * DRAM off-chip: Main memory
  * SRAM on-chip: Cache
    * Caches have a deeper hierarchy: L1+L2+L3. L1 is faster and smaller than L2 & L3.
    * Different types of caches
      * Data caches: Feeds processor with data manipulated during execution
      * Instruction caches: Stores instructions
    * The ratio between cache size & main memory size: \~1:1000

![](/files/-MSUABEmaiBF07xEoJNU)

![](/files/-MSUAFE_Akzl9ZaP5-5u)

The reason why cache works is the principle of locality: Programs tend to use data and instructions with addresses near or equal to those they have used recently.

* Temporal locality: Recently referenced items are likely to be referenced again in the near future
  * Data references: For example, in the code snippet below, the variable sum gets referenced at each iteration
  * Instruction references: The loop is cycled through repeatedly
* Spatial locality: Items with nearby addresses tend to come into use together
  * Data references: The elements in the array abc are accessed in succession (stride-1 reference pattern)
  * Instruction references: The instructions are referenced in sequence

```
sum = 0;
for (i = 0; i < n; i++)
    sum += abc[i];
return sum;
```

### Case study: Adding the entries in an N-dimensional matrix (not covered in class)

Take-home message: Well-written programs leverage data/instruction locality (which brings cache into the play) for better performance


# Lecture 5: Caches, wrap up. Virtual Memory.

## Lecture Summary

* Wrap up Cache
* Virtual Memory

## Caches

Types of cache misses, ranged by the amount of delay caused:

* Cache read miss from instruction cache
* Cache read miss from data cache
* Cache write miss to data cache

Reasons for cache misses:

* Cold (compulsory) miss: Cache is empty
* Capacity miss: Not enough space
* Conflict miss: Enough space, a lot of conflicts (and thus replacements)

Common placement policies are:

* Fully associative (M-way associative, if M blocks in total)
* K-way associative: Each set fits K blocks
* Direct mapped (1-way associative)

[Here are some comic illustrations for understanding cache basics, cache misses, and cache associativity. Source: CS Illustrated from Berkeley.](http://csillustrated.berkeley.edu/illustrations.php)

![](/files/-MSd5Io4uRwVpa6CTcgS)

General Cache Organization:

* B: Number of bytes in a cache line (typically 64)
* E: Number of cache lines/blocks that combine to make up a set (typically 2^{0,1,2,3,4})
* S: Number of sets that make up the cache
* T: Total cache size
* B \* E \* S = T

### Case study: Adding the entries in an N-dimensional matrix

Accessing data with locality gives a huge speedup:

```
int sum_array_rows(int** a){
    int i, j, sum = 0;
    
    // Option 1: Accessed with locality
    for (i = 0; i < M; i++)
        for(j = 0; j < N; j++)
            sum += a[i][j];
    
    // Option 2: Accessed w/o locality
    for (j = 0; j < N; j++)
        for (i = 0; i < M; i++)
            sum += a[i][j];

    return sum;
}
```


# Lecture 6: The Walls to Sequential Computing. Moore’s Law.

## Lecture Summary

* Wrap up Caches
* Virtual Memory

## Caches

![](/files/-MT4Gz0D9R0lC7yJDhbe)

* Handling a write-hit
  * Write-through
  * Write-back
* Handling a write-miss
  * Write-allocate
  * No-write-allocate
* Typical combos in practice
  * Write-back + Write-allocate (more common)
  * Write-through + No-write-allocate

Miss rate is more important than the hit rate: 97% hit rate is \~2 times worse than 99% hit rate

![Cache Capacity Effects from Memory Mountain](/files/-MT4NLNO8skPv67NM0ir)

## Case Study: Rearranging Loops to Improve Spatial Locality

## Virtual Memory

Why memory virtualization?

* Ease of use (running programs that require more memory than physically available)&#x20;

* Isolation (running multiple programs simultaneously)

* Protection

* A page of virtual memory corresponds to a frame of physical memory

* Page table enables the translation of virtual address into physical addresses

* The page table is stored in main memory
  * If the page table is accessed for each address translation, this would be very costly

* Translation Lookaside Buffer (TLB): "Cache" for the addr translation process

![](/files/-MT4TqtQ-6GaZW2HvDkn)

![](/files/-MT4TyJfcTfwIFBPdeA3)


# Lecture 7: Parallel Computing. Flynn's Taxonomy. Amdahl's Law.

## Lecture Summary

* Wrap up Virtual Memory
* Intuitions for Parallel Computing
* Flynn's Taxonomy
* Amdahl's Law

## Why Parallel Computing?

Sequential computing is facing these steep hills to climb:

* Memory Wall: Speed difference between CPU & memory outside the chip
* ILP Wall
* Power Wall: Latency & limited communication bandwidth beyond chip boundaries

### Memory Wall

![](/files/-MT4a2iWQ7l99EZ7X4TY)

Take-home message: Try to stay away from long and winding conversations with the main memory

### ILP Wall

![ILP elicits very complex microarchitecture](/files/-MT4axpClWcbfFSwnde4)

Instruction pipelining; Superscalar execution; Out-of-order execution; Register renaming; Speculative execution; Branch prediction

Predicting the future comes at the cost of microarchitecture complexity and power cost

### Power Wall

Power, and not manufacturing, limits traditional general-purpose microarchitecture improvements

### Recap

![](/files/-MT4cZeoePMCw5T0unpc)

## Now What?

![](/files/-MT4dSLA9YNz8Hrcy0tl)


# Lecture 8: GPU Computing Intro. The CUDA Programming Model. CUDA Execution Configuration.

## Lecture Summary

* Flynn's Taxonomy
* Amdahl's Law
* Start GPU computing

## Flynn's Taxonomy

* SISD: Single Instruction/Single Data
* SIMD: Single Instruction/Multiple Data
* MISD: Multiple Instruction/Single Data
* MIMD: Multiple Instruction/Multiple Data

![](/files/-MTCW1eYJShRco27pP_X)

![](/files/-MTCW5CQ3ggl5t6KLqDG)

![](/files/-MTCWDPALIAWy9R5ur0E)

![](/files/-MTCWGj8vHZqL7dTzuRu)

## Amdahl's Law (Law of Diminishing Returns)

The overall speedup relies on the worst-performing sections the most (see more explanations [here](/earlier-readings-and-notes/index/raid-a-case-for-redundant-arrays-of-inexpensive-disks#background-and-motivation)).

## GPU Computing with CUDA

![](/files/-MTCY1N7G585nZMvP7TZ)

![](/files/-MTCY6bbnqKgosahC27h)

![Fermi](/files/-MTCZ2upP99L7gFQCBg_)

![Volta](/files/-MTCZBoe6MnKuow3-sXs)

![Ampere](/files/-MTCZGQ_nLoQKiSyANlj)

![](/files/-MTCZR6i2dI4UQGRRF0E)


# Lecture 9: GPU Memory Spaces

## Lecture Summary

* GPU computing: generalities
* GPU computing: execution configuration
* ~~GPU computing: scheduling execution~~

## Prerequisite: Parallelism

![Coarse Grain vs. Fine Grain Parallelism](/files/-MUVifdFA-2vQaaQGP28)

* Coarse grain parallelism: Good for CPUs
  * Few tasks
  * Tasks are heterogeneous
  * Tasks are in general complex, lots of control flow
  * Example: {Bake a cake, make coffee, watch lectures} at the same time
* Fine grain parallelism: Very good for GPUs, ok for CPUs
  * A lot, a lot of tasks
  * Tasks are basically identical
  * Tasks are in general pretty straightforward, lots of math, not much control flow
  * Example: Image processing (lots of pixels to deal with)

## GPU Computing

* GPGPU: General Purpose GPU Computing
  * Started in the early 2000s using graphics libraries
  * GPUs had high bandwidths
  * Data need to be moved into the GPU to process it (this may be a bottleneck!)
    * PCIe: 16-32 GB/s
    * NVLink: 5-12 times faster than PCIe 3
    * The tradeoff is worth it if the data transfer overhead is smaller than our gain
  * Idea: Use the GPU as a co-processor to handle big, parallel jobs
    * In the meanwhile, the CPU handles control of execution & corner tasks
* CUDA: Compute Unified Device Architecture, distributed by NVIDIA
  * Eliminated the graphics-constraints associated with GPGPU
  * Enables a general-purpose programming model
* GPUs:
  * Is a co-processor to the CPU/host
  * Has its own memory (device memory)
  * Runs many threads in parallel
  * The data parallel portion of an application runs on the devices as kernels executed in parallel by many threads
  * As compared to CPU threads:
    * GPUs threads are extremely lightweight
    * A GPU needs 1000s of threads for full efficiency
* Compute capability vs. CUDA version:
  * Compute capability: Refers to hardware
  * CUDA version: Refers to software that manages the hardware
* Compatibility issues
  * The CUDA driver API is backward, but not forward compatible
    * Code that works for CUDA 8.0 should work for 11.0, but not the other way around

![The CUDA execution model](/files/-MUVmJE-Ry7ItrisZL0n)

* CUDA host stream
  * The CUDA runtime places all calls that invoke the GPU in a stream (i.e., ordered collection) of calls
    * The stream is FIFO: In the picture above, Kernel1 is only called after Kernel0 finishes
  * Asynchronicity between host and device: The host continues execution right after launching a kernel
    * Synchronization can be forced
* Three opportunities for asynchronous:
  * The GPU and CPU work in async mode
  * The GPU has three engines that can work at the same time (copy-in, copy-out, execution)
  * Multiple GPUs can work at the same time on one host
* Language supported by CUDA
  * C/C++: [Check out this introduction by NVIDIA](https://developer.nvidia.com/blog/even-easier-introduction-cuda/)

## CUDA: First Example

```
#include<cuda.h>
#include<iostream>

__global__voidsimpleKernel(int* data)
{
    //this adds a value to a variable stored in global memory
    data[threadIdx.x] += 2*(blockIdx.x+ threadIdx.x);
}

int main()
{
    const int numElems= 4;
    int hostArray[numElems], *devArray;
    
    //allocate memory on the device (GPU); zero out all entries in this device array 
    cudaMalloc((void**)&devArray, sizeof(int) * numElems);
    cudaMemset(devArray, 0, numElems* sizeof(int));
    
    //invoke GPU kernel, with one block that has four threads
    simpleKernel<<<1,numElems>>>(devArray);
    
    //bring the result back from the GPU into the hostArray
    cudaMemcpy(&hostArray, devArray, sizeof(int) * numElems, cudaMemcpyDeviceToHost);
    
    //print out the result to confirm that things are looking good 
    std::cout << "Values stored in hostArray: " << std::endl;
    for (int i = 0; i < numElems; i++)
        std::cout<< hostArray[i] << std::endl;
    
    //release the memory allocated on the GPU 
    cudaFree(devArray);
    return 0;
}
```

![](/files/-MUVq6mOrhPDgHSMp1rQ)

## GPU Execution Configuration

* Nomenclature
  * Host: The CPU executing the "master" thread
  * Device: GPU card, connected to the host through a PCIe connection
  * The host instructs the device to execute kernels
  * Defining the execution configuration: The process in which the host tells the device how many threads should each execute kernels

```
__global__ void kernelFoo(...); // declaration

dim3 DimGrid(100, 50);        // 2D grid structure, w/ total of 5000 thread blocks 
dim3 DimBlock(4, 8, 8);       // 3D block structure, with 256 threads per block 

kernelFoo<<<DimGrid, DimBlock>>>(...arg list...);
```

* The concept of "block" is important since it represents the entity that gets executed by an SM (stream multiprocessor)
* Threads in each block:
  * The threads can be organized as a 3D structure (x, y, z)
  * Max x- or y- dimension of a block is 1024
  * Max z- dimension of a block is 64
  * Max # threads per block is 1024
* Threads and blocks have indices
* 3D layout:
  * Most of the time people use 1D
  * This simplifies memory addressing when processing multi-dimensional data
    * Handling matrices
    * Solving PDEs on 3D subdomains

![](/files/-MUW0b9qDOcgAhvNRZu_)

![](/files/-MUW0oD0xmK2pbn-eWUi)

## Example: Matrix Multiplication

* Scope:
  * Only global memory (no shared memory)
  * Matrix will have a small dimension (one block of threads only)
  * Focus on `threadIdx` usage & memory transfer between host and device

![](/files/-MUWfvzK4fne3yMuTY4e)

![](/files/-MUWgPneHTyj58McVWZa)

### Code

Note that the following kernel is launched using `MatrixMulKernel<<<dimGrid, dimBlock>>>(Md, Nd, Pd)` where dimGrid is (1,1,1) and dimBlock is (WIDTH, WIDTH).

![Device-side kernel function](/files/-MUWgrOlSr6x139aQlz-)

* Words of wisdom: In GPU computing, we use as many threads as data items (tasks, jobs) we have to perform **(Number of threads == Number of data items)**
* Understanding what thread does what job is a very common source of error in GPU computing

Typically, in each kernel, we do ...

```
__global__ void multiply_ab(int* a, int* b, int* c, int size)
{
    int whichEntry = threadIdx.x + blockIdx.x * blockDim.x;
    if (whichEntry < size)  // ... this because ...
        c[whichEntry] = a[whichEntry] * b[whichEntry];
}
```

... because all blocks launched have the same number of threads, and we need to prevent out-of-bounds indexing. Say we have an array of 1493 elements and we launch two blocks of 1024 threads each, some threads will not do work.

![](/files/-MUWiFbY5zYcHSyT-2BM)

> That's probably one of the instances, probably many instances, when you regret that you took 759, because this is not fun.    -- Prof. Dan Negrut


# Lecture 10: GPU Scheduling Issues.

## Lecture Summary

* Wrap up GPU computing: generalities, execution configuration
* GPU computing: scheduling execution

## Using Multiple Blocks

![](/files/-MTag_62WJSGAkjketja)

## Execution Scheduling Issues

![Thread Index vs. Thread ID](/files/-MTahxcIWx52K-kVbS2C)

Scheduling questions:

* What is the order for the blocks to be executed?
* How is this execution process managed?
* When/How are the threads in a block executed?

Two levels of schedulers:

1. Device-level scheduler (NVIDIA GigaThread engine): Assigns (large numbers of) blocks to (small numbers of) SMs that signal that they have “excess capacity”
   1. Once a block is picked up for execution by one SM, it does not leave the SM before all threads in that block finish executing the kernel. Only when a block is finished & retired can we place another block on that SM. Thus, more SMs means a more expensive card.
2. SM-level scheduler (more interesting): Schedules the execution of the threads in a block onto the SM functional units

### SM-Level Scheduling

![Note that tensor cores are not present in older architectures](/files/-MTalNYiLac8hmBURtTZ)

* Each block of threads are divided into 32-thread warps
  * 32: Selected by NVIDIA
  * Warp: A group of 32 thread of consecutive IDs, basic scheduling unit on the SM
* SM hardware implements almost zero-overhead warp scheduling/switching

![SM Architecture Specifications (for one SM)](/files/-MTanyadX_4VkVSK9BzN)

* Thread IDs within a warp are consecutive and increasing

* But we cannot assume ordering among warps

* There are three possible states for warps:
  * Active warps (deployed on an SM)
  * Eligible warps (a subset of active warps)
  * Issued warps (a subset of eligible warps)

* Warp stalling: No new instruction issued at a clock cycle
  * Possible reasons
    * Instruction fetch
    * Memory dependency
    * Execution dependency
    * Synchronization barrier

* In execution configurations, we should have thread block sizes that result in mostly full warps

## Thread Divergence (pre-Volta)

Consider this:

```
__global__ void odd_even(int n, int* x)
{
    int i = threadIdx.x + blockDim.x * blockIdx.x;
    if( (i & 0x01) == 0 )
    {
        x[i] = x[i] + 1;
    }
    else
    {
        x[i] = x[i] + 2;
    }
}
// half of the threads in the warp execute the if clause, and the other half the else clause

```

![A visualization of what happens (execution moves forward for half of the threads each time in lockstep fashion)](/files/-MUWr8P_1JMOTjugaEYD)

* The performance decreases with the degree of divergence in warps, say a 32-case switch statement
* Solutions
  * Pre-Volta: a single program counter is shared amongst all 32 threads, combined with an active mask that specifies which threads of the warp are active at any given time
  * Post-Volta: enables equal concurrency between all threads, regardless of warp
    * Execution state (PC, program counter & S, call stack) are maintained per thread (as opposed to one per warp up until Pascal)

![](/files/-MUWtho4gdQUFJT7mq5U)


# Lecture 11: Execution Divergence. Control Flow in CUDA. CUDA Shared Memory Issues.

## Lecture Summary

* Last time
  * GPU Computing: Execution Scheduling
    * Block scheduler (at the GPU level)
    * Warp scheduler (at the SM level)
  * Thread Divergence
* Today
  * Aspects related to how GPU memory operations take place

## The NVIDIA GPU Memory Ecosystem

![From high vantage point (2 blocks w/ 2 threads each)](/files/-MU-kgq9_xEuiTHANh22)

Each thread can:

* R/W per-thread registers&#x20;
* R/W per-thread local memory&#x20;
* R/W per-block shared memory&#x20;
* R/W per-grid global memory&#x20;
* Read only per-grid constant memory&#x20;
* Read only per-grid texture memory&#x20;
* Read only per-grid surface memory

Some aspects of Local Memory:

* Physically, local memory does not exist
  * In reality, data stored in local memory is placed in cache or the global memory at run time or by the compiler
* It's specific to one thread and not visible to any other thread
* Local memory has the same latency as global memory, unless cached

Different memories:

* Global memory: Main means of communicating R/W data between host and device. cudaMalloc(), cudaFree(), and cudaMemcpy() operate here. Note that there are four types of cudaMemcpy transfers ({host/device} to {host/device}), and things happen over a PCIe connection.
* Texture and Constant memories: Constants initialized by host, contents available to all threads.&#x20;

Global, texture and constant memories are accessible by host (done at high latency, low bandwidth).

![](/files/-MUWvdlx1VKsrdzJ2o0d)

![](/files/-MUWvh-iq5_1M-3g8CTm)

![Memory Access Times](/files/-MU-mPQFGKffrBKhLvJA)

![Storage Locations](/files/-MU-mVa98p1fKr9ClwX-)

![The 3 most important GPU memory spaces](/files/-MU-mewdFW4fThoFzjy2)

## Case Studies: Matrix Multiplication, Revisited

Purpose:

* See an example where the use of multiple blocks of threads play a central role

* Highlight the use/role of the shared memory

* Point out the \_\_syncthreads() function call (synchronizes all threads in a block)

* The previous example: Low arithmetic intensity, a lot of unnecessary movements from global memory to device

* **Rule of thumb: If the data that you, as a thread, use can also be used by another thread in your block, then you should consider using shared memory**

* To use shared memory:
  * Partition data into data subsets (tiles) that each fits into shared memory
  * Handle each data subset (tile) with one thread block by:
    * Loading the tile from global memory into shared memory, using multiple threads to exploit memory-level parallelism
    * Performing the computation on the tile from shared memory; each thread can efficiently multi-pass over any data element of the tile

![](/files/-MU-pK_llmHlfZTlVScE)

![](/files/-MU-pOckb9xMGBF_cVqq)

* `__syncthreads()` synchronizes all threads in a block
  * Used to avoid RAW/WAR/WAW hazards when accessing shared or global memory
  * Be very careful when using it in a conditional
* 3 ways to set aside shared memory:
  * Statically, declare inside a kernel
  * Through the execution configuration (see code block below)
  * Dynamically, via CUDA driver API `cuFuncSetSharedSize()` (out of scope)

```
__global__ void MyFunc(float*) // __device__ or __global__ function 
{
    extern __shared__ float shMemArray[];
    // Size of shMemArray determined through the execution configuration
    // You can use shMemArrayas you wish here...
}

// invoke like this. Ns indicates the size in bytes to be allocated in shared memory
MyFunc<<< Dg, Db, Ns>>>(parameter);
```

![Example: Reversing an array using dynamic shared memory](/files/-MUWy0e-NWQno8IS6iby)

![How different technology fetches data into shared memory](/files/-MUWyXTZw2J_L2O5Wrvs)

* Each SM has shared memory organized in 32 memory banks
  * Successive 32-bit words map to successive banks
  * Each bank has a bandwidth of 32 bits per clock cycle
* ShMem and L1 cache draw on the same physical memory inside an SM

![](/files/-MUX12bQ1JhNcQJ2FIdc)


# Lecture 12: Global Memory Access Patterns and Implications.

## Lecture Summary

* Last time
  * Aspects related to how GPU memory operations take place
    * Registers, local memory, shared memory, global memory (texture & constant memories)
* Today
  * GPU mem operations: focus on shared memory
  * GPU mem operations: focus on global memory
  * How parallel computing makes memory operations tricky
  * Atomic operations
  * Things that determine the speed of execution of a kernel

## Banks

* Recap
  * Each SM has 32 banks
  * Each warp has 32 threads
  * At any point in time, the 32 banks are only accessed by threads in one warp
* Bank conflicts
  * No bank conflicts: Either linear addressing or random 1:1 permutation
  * Bank conflicts
    * N-way bank conflicts: a bank is accessed by N threads
    * Reading: "no conflict"
      * Broadcast: all threads in a warp access the same bank
      * Multicast: some threads in a warp access the same bank
  * For visualizations, see the slides

![An example of bank conflicts](/files/-MUZlIAGrHfJhPtNUh5E)

![Linear addressing](/files/-MUZm4dhjwRCpxLy2CLq)

### Example

![](/files/-MUZnKDfllr2i0JB2tpU)

![](/files/-MUZnMp63rpBSH7OYXQi)

## Getting the results right (broadly for parallel computing)

### Example

![We are all good](/files/-MUZoQJkRZUEyt1KOiOw)

![All hell break loose](/files/-MUZoUFne1jqlSoHD7vu)

### Data Hazards

* Three types of data hazards
  * RAW: Read-After-Write (j ought to read only after the write by i occurred)
  * WAR: Write-After-Read (j ought to write only after the read by i occurred)
  * WAW: Write-After-Write (j ought to write only after the write by i occurred)
* Moral of the story: The ordering of memory operations is important
* Types of memory consistency
  * Sequential consistency: All reads and all writes are in-order
  * Relaxed consistency: Some types of reordering are allowed
  * Weak consistency: Reads & writes arbitrarily reordered
* The `__threadfence()` family of functions: enforces that memory transactions for one thread can be seen by other threads
  * `__threadfence_block()`: Execution of the kernel by the calling thread pauses until all global and shared memory outstanding writes are visible to all threads in block
  * `__threadfence()`: Execution of kernel by a calling thread ensures all global and shared memory outstanding writes are visible to all threads in block AND all other threads in flight for global data
  * Not about synchronization, but about memory transaction
  * For an example, see the slides
* The volatile qualifier
  * If a variable located in global or shared memory is declared as volatile, the compiler assumes that its value can be changed or used at any time by another thread and therefore any reference to this variable compiles to an actual memory read or write instruction
  * W/o this keyword, the compiler optimizes instructions related to shared memory, and this keyword disables those optimizations
* volatile applies equally well to sequential computing
* `__threadfence()` is specific to parallel computing

## Getting the results fast (for specifically GPU computing)

* Issues
  * Not all global memory accesses are equally efficient (higher priority)
  * Not all shared memory accesses are equally efficient
* Two aspects of global memory access are relevant
  * The layout/pattern of the access
    * If threads that access global memory are neatly grouped, then we have a coalesced memory access, and this is good
    * If the threads are scattered all over the place, it impacts the effective bandwidth
  * The alignment of the data we are fetching from global memory
    * If all threads in a warp access data inside only one memory block, it's great
* Good memory accesses are coalesced and properly aligned

![Coalesced and aligned](/files/-MUZyWDq9xK3J1Op9SZb)

![Coalesced but not aligned](/files/-MUZyZAoBdF0oSW4IPxZ)


# Lecture 13: Atomic operations in CUDA. GPU ode optimization rules of thumb.

## Lecture Summary

* Last time
  * GPU mem operations: focus on shared memory
  * GPU mem operations: focus on global memory
  * How parallel computing makes memory operations tricky
    * RAW, WAW, WAR hazards
* Today
  * Atomic operations
  * Things that determine the speed of execution of a kernel
  * Case studies: parallel reduction on the GPU&#x20;

## Example: Element-Wise Matrix Addition

![](/files/-MU9jm3pzIjfGArlsVQt)

As `threadIdx.x` changes faster than `threadIdx.y`, we should have `C[j][i] = A[j][i] + B[j][i]` instead of `C[i][j] = A[i][j] + B[i][j]`.

## Example: CUDA Global Memory Access

![](/files/-MU9kFzFVaimYFzWEbdb)

1. Good!
2. Coalesced, but not aligned
3. Misaligned, non-coalesced
4. Level of indirection

Two ways to store data in global memory:

1. Array of structures (AoS)
2. Structure of arrays (SoA)

## Atomic Operations

* Atomic memory operations is a mechanism that alleviates race conditions/access coordination problems
* The order in which concurrent atomic updates are performed is not defined
* While the order is not clear, none of the atomically performed updates will be lost
* Performance becomes poor when many threads attempt to perform atomic operations on a small number of locations
* When to use:
  * Cannot fall back on normal memory operations because of possible race conditions
  * Use for infrequent, sparse, and/or unpredictable global communication
  * Use shared memory and/or customized data structures & algorithms to avoid synchronization whenever possible
* Difference between `__syncthreads()` and an atomic operation:
  * `__syncthreads()` establishes a barrier, i.e. of synchronization
  * Atomic operations instead tie to the idea of coordination in relation to operations that involve memory transactions. Threads need not synchronize their execution, it’s only that a certain memory operation in a kernel is conducted in an atomic fashion

## Resource Management Considerations

* "Used at capacity": SM executes the max number of warps it can possibly host
* Three factors come into play:
  * threads/block
  * registers/thread
  * shMem/block
* Occupancy != Performance (yet it's a pretty good proxy)

## CUDA Optimization: Rules of Thumb

### High Priority

1. To get the maximum benefit from CUDA, focus first on finding ways to parallelize sequential code. Expose fine-grain parallelism
2. Minimize data transfer between the host and the device, even if it means running some kernels on the device that do not show performance gains when compared with running them on the host CPU
3. Strive to have aligned and coalesced global memory accesses. Design your implementation such that global memory accesses are coalesced for that part of the red-hot parts of the code
4. Minimize the use of global memory. Prefer shared memory access where possible (consider tiling as a design solution)

### Medium Priority

1. Accesses to shared memory should be designed to avoid serializing requests due to bank conflicts
2. Strive for sufficient occupancy
3. Keep the number of threads per block a multiple of 32 to avoid wasted lanes
4. Use the fast math library whenever speed is very important, and you can live with a tiny loss of accuracy
5. Avoid thread divergence

## Some More Compiler-Related Stuff

![Compiling CUDA code with nvcc driver. PTX: Parallel Thread Execution, an ISA that exposes the GPU as a data-parallel computing device. It's like NVIDIA-specific Assembly.](/files/-MU9vzjvxSXw4z4Su2t4)

![](/files/-MU9wMyXHqIzEzLcYzMG)


# Lecture 14: CUDA Case Studies. (1) 1D Stencil Operation. (2) Vector Reduction in CUDA.

## Lecture Summary

* Last time
  * Atomic operations
  * Things that shape the speed of execution of a kernel
    * The concept of "occupancy" and what impacts it (how many threads per block, how many registers/thread, how much ShMem/block)
  * Rules of thumb, for good execution speed in GPU computing
  * The nvcc toolchain, and how code is sent to host or gpu compilers
* Today
  * Case studies: parallel reduction on the GPU & 1D convolution
  * Looking beyond today: some more GPU computing feature, but looking for a while into optimization features

![Application optimization process](/files/-MUaGzNUqo5nvkZ8652d)

## 1D Stencil Operation

![What the algorithm does](/files/-MUaJHU_DXQVJAav8PI6)

![Serial implementation](/files/-MUaF6vRP6ESWAOteYoK)

![Parallel implementation](/files/-MUaFBvQe_ao9L1H6S1U)

![nvprof pointed out spaces for optimizations](/files/-MUaFhSgzV3a-743xa61)

![Use pinned memory (pinned memory cannot be paged out by the OS)](/files/-MUaFmXuI4iYW3mtJnTM)

![Data partitioning example (overlapping compute & memory)](/files/-MUaHGbHf0eZozOEa_6H)

![Performance improvements](/files/-MUaHVUUxxYU36hw0EQF)

![Optimization summary](/files/-MUaHl8Jiw7XsP_gAIWm)

## Vector Reduction in CUDA

![What the algorithm does (summing all entries in an array)](/files/-MUaJQQHlwV7DAtdvMzz)

Problem: Ideally we want to synchronize across all thread blocks, but CUDA does not have global synchronization. Our workaround is to decompose into multiple kernels.

* Optimization goal: Reaching GPU peak performance
  * Choosing the right metric
    * GFLOP/s: for compute-bound kernels
    * Bandwidth: for memory-bound kernels
* Reductions have low arithmetic intensity (1 flop/2 elements loaded), so we should go for peak bandwidth

![Interleaved addressing: highly divergent warps are inefficient, and % operator is very slow](/files/-MUaZQUej2Ywkez82zyK)

![Change which thread works on what. New problem: shared memory bank conflicts](/files/-MUaZafkMRUMDusaJMYU)

![Sequential addressing](/files/-MUa_4Wm9gkhVOKy1Uit)

* Kernel 4: Replace single load w/ two loads and first add of the reduction
* Kernel 5: Loop unrolling (unroll last warp)
* Kernel 6: Completely unrolling (using templates)
* Kernel 7: Multiple elements per thread


# Lecture 15: CUDA Case Studies. (3) Parallel Prefix Scan on the GPU. Using Multiple Streams in CUDA.

## Lecture Summary

* Last time
  * Case studies: parallel reduction on the GPU & 1D convolution
  * Looking beyond today: some more GPU computing feature, but looking for a while into optimization features
* Today
  * One more cast study: parallel prefix scan
  * Using streams in GPU computing: increasing problem size; improving execution speeds

## Parallel Prefix Scan on the GPU

![Definition of the algorithm](/files/-MUagVRQSj9oQ_u7rnM0)

### Algo 1: Hillis & Steele (1986)

* Simple, but suboptimal (O(N\*log2(N)))

![](/files/-MUai1zJdx8Fh1Gx69o8)

![](/files/-MUahma-WBNSWHRrRell)

### Algo 2: Harris-Sengupta-Owen (2007)

* Convoluted, but O(N)
* Balanced trees: A common parallel algorithm pattern
  * Upsweep from roots to the main trunk, and then down sweep from trunk to root
  * "Tree": Just a concept--the actual data structure is not used

![The reduction/upsweep step](/files/-MUalCCjeF3jD1xH6QV1)

![The down sweep step. Sheesh, this is just...](/files/-MUalIN_NGokDWkN0SY3)

## CUDA Streams

* A CUDA-enabled GPU has 2 engines
  * An execution engine
  * A copy engine (which contains 2 sub-engines that can work simultaneously)
    * A H2D copy sub-engine
    * A D2H copy sub-engine
* Async execution
  * Examples: Kernel launches, D2D mem copies, mem copies by functions with the `Async` suffix, etc
* Overlapping Host <--> Device data transfer with device execution
  * Issue: The device execution stack is FIFO
    * Addressed by the usage of CUDA "streams"
* Concurrency can be managed through streams
  * Concurrency means one of two things:
    * The copy and the execution engines of GPU working at the same time
    * Several different kernels being executed at the same time on the GPU
* A stream is a sequence of CUDA commands issued by the host that executes on the GPU in issue-order
  * CUDA operations in different streams may run concurrently
  * CUDA operations from different streams may be interleaved
* As soon as a CUDA function is invoked, a default stream (stream 0) is created
* Create using `cudaStreamCreate()`, destroy using `cudaStreamDestroy()`

![](/files/-MUau-OIEYDDvXFpWp9r)


# Lecture 16: Streams, and overlapping data copy with execution.

## Lecture Summary

* Last time
  * Case study: Parallel prefix scan
  * Using streams in GPU computing
* Today
  * Wrap up streams in GPU computing: increasing problem size; improving execution speeds
  * Debugging & profiling GPU code: some nuts and bolts

## Streams

### Example 0

* Stream 1 & 2 are defined and initialized already
  * Use the two copy sub-engines at the same time: copy in (stream1), copy out (stream2)
  * Postpone launching of myKernel in stream2until the copy operation in stream1is completed

```
cudaEvent_t event;
cudaEventCreate(&event);                           // create event
cudaMemcpyAsync(d_in, in, size, H2D, stream1);     // 1) H2D copy of new input
cudaEventRecord(event, stream1);                   // record event
cudaMemcpyAsync(out, d_out, size, D2H, stream2);   // 2) D2H copy of previous result
cudaStreamWaitEvent(stream2, event);               // wait for event in stream1
myKernel<<<1000, 512, 0, stream2>>>(d_in, d_out);  // 3) GPU must wait for 1 and 2
someCPUfunction(blah, blahblah)                    // this gets executed right away

```

### Example 1

![](/files/-MUioE8afH4hCpCSuxzN)

![Stage 3 enqueues the set of GPU operations that need to be undertaken (the "chunkification")](/files/-MUioIE-U1H7NQ0bvAyj)

![Concurrency (manual pipelining)](/files/-MUiouqO4gVQSOgbUNib)

### Example 2.1

* Similar to example 1, but with two streams to increase the speed of execution
* This actually doesn't give a big speedup (62 ms -> 61 ms)

![](/files/-MUiqUBodq2ZSkIw6OH7)

![](/files/-MUiqX9GJQEKaLrtVY2q)

![Note that the kernel stays the same](/files/-MUiqZp1B9Dzr7wqlkKf)

![There is actually no overlap of copy & execution...](/files/-MUirjp6weTcbnpgAcXG)

### Example 2.2

![](/files/-MUirrUWb-2ttNJpko8u)

* Streams recap
  * Concurrency brings two flavors:
    * The copy and the execution engines of the GPU working at the same time
    * Several different kernels being executed at the same time on the GPU
* CUDA/GPU computing recap
  * Generally, any application that fits the SIMD paradigm can benefit from using GPUs
    * Good speedups at a small time and financial investment
  * Hardware is changing faster than software&#x20;

## Debugging & Profiling in CUDA

### cuda-gdb

* gdb but with more things that need our attention
* For more usage, see the slides
  * Program execution control
  * Thread focus
  * Program state inspection (stack trace, source variables, memory, HW registers, code disassembly)
  * Run-time error detection (cuda-memcheck)
  * Tips, best practices, and misc notes
* I still prefer `printf()`, change my mind. /s

### Profiling

* Nsight Compute (only focus on GPU; ncu to collect data, ncu-ui to visualize interactively)
* Nsight Systems (focus on the whole system)
* nvprof (being deprecated rn)


# Lecture 17: GPU Computing: Advanced Features.

## Lecture Summary

* Last time
  * Streams in GPU computing
  * Debugging & profiling
* Today
  * Use of unified memory in CUDA GPU Computing

## Unified Memory (Managed Memory) in CUDA

* cudaMemCpy
  * Available in release 1.0
  * Moves data between host and device (over PCI-E)
* cudaHostAlloc
  * Allocate host memory rather than malloc-ing -> improve host/device data transfer speed if host memory is not pageable
  * Pros
    * Faster device <--> host transfer
    * Enables the use of asynchronous memory transfer and kernel execution
    * Enables mapping of the host pinned memory into the memory space of the device
  * Cons
    * Large memory impacts system performance
    * Memory allocation speed using cudaHostAlloc is low
  * `cudaError_t cudaHostAlloc(void** pHst, size_t sz, unsigned int flag);`
    * Using the flag `cudaHostAllocMapped` maps the memory allocated on the host in the memory space of the device for direct access
  * **Zero-Copy (Z-C)** GPU-CPU interaction
    * We no longer need an explicit CUDA runtime copy call to move data onto the GPU
    * This balloons the device memory so that it includes main memory that physically resides on the host
    * However, this requires the runtime call to cudaHostGetDevicePointer(). The need for this is eliminated by the Unified Virtual Addressing (UVA) mechanism.
* UVA: GPU and CPU share the virtual memory space. UVAS: UV Address Space.
  * CUDA runtime can identify where the data is stored based on the pointer
  * Instead of `cudaMemcpyxxx`, now we can use a generic `cudaMemcpyDefault`
* Z-C: Use pointer within device function to access host data
* UVA
  * Data access: A GPU can access data on a different GPU
  * Data transfer: Copy data in between GPUs
* UM (Unified Memory): Like UVA, but enabled the CPU to access GPU memory
  * UM works in conjunction with a "managed memory pool"
  * `cudaMallocManaged`replaces the need for explicit memory transfers between host and device, and cudaMalloc / cudaHostAlloc
  * Data is stored on the device but migrated where needed
  * Makes writing code easier, and will probably run faster due to locality (for the casual programmer)
  * Still evolving

![Unified Memory simplifies things](/files/-MVJM7qMZ-W9tbM_RnD7)

## Review

1. cudaMemcpy
2. Z-C: Device could access memory on the host
3. UVA: Unified virtual space
4. UM: Processors can access each other's memory


# Lecture 18: GPU Computing with thrust and cub.

## Lecture Summary

* Last time
  * A three-stop journey noted in the evolution of the CUDA memory model
    * Z-C accesses on the host; the UVA milestone; the unified memory model that allowed to use of managed memory
* Today
  * GPU computing, from a distance (via thrust & CUB)

## Thrust

* Motivation
  * Increase programmer productivity
  * Do not sacrifice execution speed
* What is thrust?
  * A template library for parallel computing on GPU and CPU
  * Heavy use of C++ containers
  * Provides ready-to-use algorithms

### Namespaces, containers, iterators

* To avoid name collisions, use `thrust` vs. `std` namespaces
* 2 vector containers: host\_vector and device\_vector
  * Just like those in the C++ STL
  * Manage both host & device memory
  * Auto allocation & deallocation
* Iterators: Act like pointers for vector containers
  * Can be converted to raw containers
  * Raw pointers can also be wrapped with device\_ptr

### Algorithms

* Element-wise operations
  * for\_each, transform, gather, scatter
  * Example: SAXPY, **functor** using transform
* Reductions
  * reduce, inner\_product, reduce\_by\_key
* Prefix sums (scans)
  * inclusive\_scan, inclusive\_scan\_by\_key
* Sorting
  * sort, stable\_sort, sort\_by\_key

### General transformations. Zipping & fusing

![Problem at hand](/files/-MVK41A14RhCUuSPp4EQ)

* **Zipping**
  * Takes in multiple distinct sequences, zips into unique sequence of tuples
* **Fusing**
  * Just like zipping, but it's for reorganizing computation (instead of data) for efficient thrust processing
  * Increases the arithmetic intensity

### Thrust example: Processing rainfall data

Not covered in class

## CUB

* CUB: CUDA UnBound
* [CUB is on GitHub](https://github.com/NVIDIA/cub)
* thrust is built on top of CUB
* What CUB does
  * Parallel primitives
    * Warp-wide "collective" primitives
    * Block-wide "collective" primitives
    * Device-wide primitives
  * Utilities
    * Fancy iterators
    * Thread and thread block I/O
    * PTX intrinsics
    * Device, kernel, and storage management


# Lecture 19: Hardware aspects relevant in multi-core, shared memory parallel computing.

## Lecture Summary

* Last time
  * GPU computing via thrust & CUB
* Today
  * Final project proposal discussion
  * Parallel computing on the CPU: Hardware & OpenMP generalities

## Multi-core Parallel Computing with OpenMP

![Opportunities for efficiency gains](/files/-MWPrN6gEX0vz3CgMsfo)

* OpenMP targets parallelism on SMP architectures
* It is handy when
  * You have a multi-core processor, say 16 cores/socket (go beyond that and we suffer from diminishing returns due to overheads)
  * Might have multiple sockets, say 2
  * You have a good amount of system memory, say 64 GB
* Processes and threads are similar in the sense that they are both independent sequences of execution
  * OpenMP touches on threads, while MPI touches on processes
  * Threads of the same process run in a shared memory space and they have one translation page. Processes, on the other hand, run in separate memory spaces.
* We want to use OpenMP for both data parallelism and task parallelism
  * Data parallelism: The processing of a large amount of data elements can be done in parallel
  * Task parallelism: The execution of a collection of tasks can be performed in parallel

![Hello world for OpenMP](/files/-MWPuU_0b0W369Q-EOrI)

* The OMP parallel region is similar to a CUDA kernel: both are executed by threads
  * A major difference
    * Variables inside GPU kernel are truly local variables, stored in registers
    * OMP variables in a parallel region may or may not be visible to other threads executing the code of the parallel region: the scoping is tricky
* `#include <omp.h>`
* Most OpenMP constructs are compiler directives. In C/C++, they take the form of `pragmas`
* Programming model: A master thread spawns a team of threads

![](/files/-MWPv6oLzMgLG4JiWKzg)


# Lecture 20: Multi-core Parallel Computing with OpenMP. Parallel Regions.

## Lecture Summary

* Last time: OpenMP generalities
* This time: OpenMP nuts & bolts

## OpenMP

![Compiler directives examples (the directive goes behind \`#pragma omp\`)](/files/-MWPxtwjfZnxLR5FZYs_)

![User-level run time routines](/files/-MWPyGK8qdfNVpiiy1An)

![Environment variables. This helps with bypassing the run-time function calls, but using env vars does not allow for dynamic OpenMP behavior. A function call overrides an env var setting, though.](/files/-MWPzoPbfOsAB5BoHFZR)

* OpenMP: portable and scalable model for shared memory parallel applications
  * No need to dive deep and work with POSIX pthreads
  * Under the hood, the compiler translates OpenMPfunctions and directives to pthread calls
* Structured block and OpenMP construct are the two sides of the “parallel region” coin
* In a structured block, the only "branches" allowed are exit() function calls. There is an implicit barrier after each structured block where threads wait on each other.

![](/files/-MWQ3xI0kcbZyEtNFq20)

![](/files/-MWQ4KyzoRLHdvtw2S0e)

### Nested Parallelism

![](/files/-MWQ4wpwOB2h9VcxWNQG)

* The nested parallelism behavior can be controlled by using the OpenMP API
* The single directive identifies a section of the code that must be run by a single thread
  * The difference between single and master is that in single, the code is executed by whichever thread reaches the region first
  * Another diff is that for single, there is an implicit barrier upon completion of the region

![](/files/-MWQ6O-K0uwovOYdkqGB)

### Work Sharing

* Work sharing is a general term used in OpenMP to describe the distribution of work across threads
* The three main constructs for automatic work division are:
  * omp for
  * omp sections
  * omp task

### omp for

![](/files/-MWQ7rSOacz3mKGhmAWW)

* A #pragma omp for inside a #pragma omp parallel is equivalent to #pragma omp parallel for
* Most OpenMP implementations use default block partitioning, where each thread is assigned roughly n/thread\_count iterations. This may lead to load imbalance if the work per iteration varies
  * The schedule clause comes to the rescue!
  * Usage example: #pragma omp parallel for schedule(static, 8)

![](/files/-MWQ9Op2h05CJX-260_A)

![Effects of different schedules, assuming 3 threads](/files/-MWQ9lQXTJC4JTmbSkAx)

![Choosing a schedule](/files/-MWQ9un8IUUZ2x-N6Amt)

* OpenMP will only parallelize for loops that are in canonical form. Counterintuitive behavior may happen
* The collapse clause supports collapsing the embedded loops into one uber loop
  * For example, if the outer loop has 10 iters, the inner loop has 10^7 iters, and we have 32 threads: parallelizing the outer loop is bad (10<32), parallelizing the inner loop is good, but we can do better using collapse

![](/files/-MWQBmo703WkJX7H36YC)

### omp sections

![](/files/-MWQC0wYU64VC7X_6cjq)

![](/files/-MWQCApH6GluvFoFZZLA)

![](/files/-MWQCaCEH-vq5kTPLx-f)

![](/files/-MWQCcvbiRaXqkfBcdvx)


# Lecture 21: OpenMP Work Sharing.

## Lecture Summary

* Last time: OpenMP nested parallelism, work sharing (for loops, sections)
* Today
  * OpenMP: nested parallelism, work sharing (tasks)
  * OpenMP: variable scoping, synchronization, loose ends

## OpenMP Work Sharing

### omp sections

Ending example

![](/files/-MWSTn2X37MJRJRVI-tw)

![](/files/-MWSTr_VW11tfM0YFYDo)

### omp tasks

* Pros: Allows parallelization of irregular problems
  * Unbounded loops
  * Recursive algorithms
  * Producer/consumer
* Cons: Relatively tricky to deal with & introduce some overhead&#x20;
* Motivations
  * OpenMP started to be tailored for large array-based applications
  * For example, the parallelization of a dynamic list traversal cannot be done in OpenMP for a long time
  * Storing pointers to list elements in an array: High overhead for array construction (not easy to parallelize)
  * Using single nowait inside a parallel region: High cost of the single construct. Also, each thread needs to traverse the entire list to determine if another thread has already processed that element
* Who does what and when?
  * The developer
    * Uses a pragma to specify where & what the tasks are
    * Ensures that there are no dependencies (that is, tasks can be executed independently)
  * The OpenMP runtime system
    * Generates a new task whenever a thread encounters a task construct
    * Decide the moment of execution (can be immediate or delayed)
* Definition: A task is a specific instance/combo of executable code along w/ its data environment (the shared & private data manipulated by the task) and ICV (internal control variables: thread scheduling and environment variables, typically associated with OpenMP)
* Synchronization issues. Solution: use task barriers (`#pragma omp barrier`, `#pragma omp taskwait`) to ensure the completion of tasks.

![](/files/-MWSWArBj6G2eAqlx9TX)

![](/files/-MWSWfDdLK1Hrm_64t0x)

![](/files/-MWSXBEAWGJLJYw73vYk)

## OpenMP Variable Scoping Issues

* Threads have access to a pool of memory that is shared
* Threads can also have private data
* Basic rule: Any variable declared prior to a parallel region is shared in that parallel region
* The private clause reproduces for each thread variables declared private in the pragma
* There are also OpenMP variables treated as private by default
  * Stack (local) variables in functions called from within parallel regions
  * Loop iteration variables
  * Automatic variables within a statement block
* When in doubt, always explicitly indicate something to be private
* firstprivate: Specifies that each thread should have its own instance of a variable. Moreover, the variable is initializes using the value of the variable of the same name from the master thread
  * Usage: #pragma omp parallel num\_threads(4) firstprivate(i)&#x20;
* lastprivate: The enclosing context's version of the variable is set equal to the private version of whichever thread executes the final iteration of the work-sharing construct (for or section)
* Data scoping is a common source of errors in OpenMP. It is the programmer's responsibility to make sure data dependencies do not lead to race conditions

![](/files/-MWSZIRLN6GvZpV7Wtu8)

![Example of what's being shared and what's not](/files/-MWSf4FMi0EBydTdwIaK)

## OpenMP Synchronization

* Explicit barrier: #pragma omp barrier
* Implicit barriers: parallel, for, single, sections
* Unnecessary barriers hurt performance and can be removed with the nowait clause (applicable to for, single, sections)

![The nowait clause](/files/-MWScsVyj_zlTjnb_axv)

![The critical construct: prevents race conditions and protects access to shared, modifiable data](/files/-MWScy5Eon9-swpWx3bD)

![The critical construct in action. Note that naming the critical construct RES\_lock is optional but highly recommended](/files/-MWSdeBK4I5YQxxeN8OF)


# Lecture 22: OpenMP Work Sharing

## Lecture Summary

* Last time
  * OpenMP: Tasks, variable scoping, synchronization (barrier & critical constructs)
* Today
  * Wrap up synchronization
  * OpenMP rules of thumb
  * Parallel computing w/ OpenMP: NUMA aspects & how caches come into play

## Synchronization

* The atomic directive
  * A guarded memory access operation
  * Can only protect a single assignment
  * Applies only to simple update of memory
  * Is a special case of a critical section with significantly less overhead due to atomicity
* The reduction construct (see example down below)
  * Local copy of sum for each thread engaged in the reduction is private
    * Each local sum is initialized to the identity operand associated with the operator that comes into play. In this case, we have "+", so the init value is 0.
  * All local copies of sum are added together and stored in a "global" variable
  * \#pragma omp for reduction(op:list)
    * The variables in list will be shared in the enclosing parallel region&#x20;
* The simd directive
  * \#pragma omp for simd reduction(+:sum)

![The atomic directive](/files/-MXyWOilcthtpfSHN-jS)

![The reduction construct](/files/-MXyVxnynW2rk1-OQdD4)

## Performance Issues

* Common causes are:
  * Too much sequential code in your app
    * Seek to reduce amount of execution time where only one thread executes code
  * Too much communication
    * Difficult to pin down costly memory operations
  * Load imbalance
    * One thread gets too much work, while others idle waiting for it
    * For OpenMP for, one can use schedule(runtime)
      * Example: `setenv OMP_SCHEDULE "dynamic,5"`
  * Synchronization
    * Barriers can be expensive
    * Avoid them using
      * Careful use of the `nowait` clause
      * Parallelize at the outermost level possible
      * Use `critical` or `atomic`
      * Use other OpenMP facilities like `reduce`
  * Compiler (non-)optimizations
    * Sometimes the addition of parallel directives can prevent the compiler from performing sequential optimization
    * Symptom: parallel code running with 1 thread has longer execution and higher instruction count than sequential code

## NUMA

* Up to this point, we have been using the Symmetric Multi-Processing (SMP) model and we haven't been concerned about the mechanics of shared memory access
* In today's servers/clusters, nodes have many CPUs, each with many cores (this is called multi-socket configurations, as opposed to one chip per motherboard), and not all memory access are equal
* NUMA: Non-uniform memory access
  * Cost of memory access depends on which memory bank stores your data
* The NUMA factor: the ratio between the largest and shortest average amount of time for a thread running on a particular core to reach data in memory
  * A low NUMA factor is desirable (not much of a difference which bank data is stored)
  * Numa factor = 1: SMP system
  * Accessing memory outside a NUMA node: 20% slowdown for reads, 30% slowdown for writes
* NUMA aspects where OS comes into play
  * When a thread mallocs memory, how should this memory be allocated
  * Affinity: How the runtime/OS assigns a thread to a certain core
    * OMP\_PROC\_BIND: Allows you to dictate a distribution policy
      * master: Collocate threads with the master thread
      * close: Place threads close to the master in the places list
        * Useful if code is compute-bound and don't do many trips to main memory
        * Reduce synchronization costs (single, barrier, etc.)
      * spread (default): Spread out thread as much as possible
        * Useful if code is memory-bound as it improves aggregate system memory bandwidth
      * false: Set no binding
      * true: lock thread to a core
    * OMP\_PLACES: Allows you to control locations. OMP\_PLACES can assume one of these values
      * threads: Hardware thread, assuming hyper threading is on
      * cores: Core
      * sockets: Node (socket)
      * A place list: Defined by user, explicitly referencing the underlying hardware of the machine
    * An extensive list of examples can be found in the slides

![](/files/-MXyaETS9HvGPYAu_adg)

![OMP\_PLACES usage. The CPU ids can be found through \`numactl -H\` or \`lscpu\` ](/files/-MXyeZF1bhxe8BjtQtUh)


# Lecture 23: OpenMP NUMA Aspects. Caching and OpenMP.

## Lecture Summary

* Last time
  * Wrap up synchronization
  * OpenMP rules of thumb
  * Parallel computing with OpenMP: NUMA aspects
* Today
  * Parallel computing, multi-core: how caches come into play
  * Critical thinking, and similar tricks for speeding up your code

## Caches in a Multi-Core Setup

* Consistency vs. Coherence
  * Consistency establishes a set of rules that governs the collective actions of the threads relative to the entire system memory
    * Think of it this way: there are at least two memory entries that come up in the discussion
  * Coherence regards expected behavior that one memory location must display relative to transactions carried out by multiple threads running on multiple cores
    * Think of it this way: there is exactly one memory entry that comes up in the discussion
* Two established approaches for enforcing cache coherence
  * Directory-based: Directory acts as a filter through which any change to cache must pass. When an entry is changed, the directory either updates or invalidates the other caches with that entry
  * Snooping-based
    * Example: MESI protocol
      * 4 states: modified, exclusive, shared, invalid
* Assume each cache line can only exist in one of 3 states
  * Exclusive: the only valid copy in any cache
  * Read-only: A valid copy but other caches may contain it
  * Invalid: Out of date and cannot be used
* In this simplified coherency model,&#x20;
  * A read on an invalid or absent cache line will be cached as read-only or exclusive
  * A write on a line not in an exclusive state will cause all other copies to be marked invalid and the written line to be marked exclusive

![Snooping-based](/files/-MXylVuz9izNRNMKdG5D)

![4 states of the MESI protocol](/files/-MXym6YPakEnmkIjfeE5)

### False Sharing

* Each cache line is typically 64 bytes long and can store, for instance, 8 doubles or 16 ints. As soon as one entry in a cache line is changed, all the other values in cache line get dirty.
* False sharing happens when two threads are both writing into different locations within the same cache line
* Symptoms: Poor performance, high numbers of cache misses, unexpected load imbalance

## Critical Thinking

* This module brings together knowledge about
  * Compilers and how they work
  * Memory aspects: Pointers, hierarchy, latencies, bandwidths
  * Instruction Level Parallelism (pipelining, jump instructions, branch prediction, wide registers, etc.)


# Lecture 24: Critical Thinking. Code Optimization Aspects.

## Lecture Summary

* Last time
  * Parallel computing, multi-core: how caches come into play
  * Critical thinking, and similar tricks for speeding up your code
* Today
  * Critical thinking, and other similar tricks for speeding up your code

## Know your hardware

![Know your bandwidth/latency](/files/-MXzDYa0tV1dJtz9aCO7)

## Choose the right algorithm

* When working on a problem, there's more to performance than asymptotic complexity
  * Because asymptotic complexity is most often defined by the number of operations
  * Memory transactions are rarely considered: they are specific to the hardware
* Assess the arithmetic intensity associated with your problem

![](/files/-MXzESdkvKR6BfxCfv61)

* Simple optimization: Fusing transformations and do not bring data into the cache twice

## Compiler

* Aggressive optimizations done by the compiler might change the behavior of your code
* To help the compiler:
  * Allow it to see as much code as possible
  * Provide flags to convey information (e.g., the target architecture)
* There are a lot of amazing things covered in this lecture. The takeaways are:
  * Compilers are fantastic
  * Know them better to use them wisely
* A quick example is down below. Refer to the [slides](https://uwmadison.app.box.com/s/kapdp4qt18c6869dnaenlo35iytnlp2r) for a lot more fun facts

![](/files/-MXzI3IUdr6jxw-hCJ01)

![](/files/-MXzI64jS_qzpu4BFr0-)


# Lecture 25: Computing with Supercomputers.

## Lecture Summary

* Last time
  * Critical thinking, and similar tricks for speeding up your code
* Today
  * Wrapping up critical thinking
    * Case study

## Critical Thinking

* Once again, a lot of interesting stuff in the [slides](https://uwmadison.app.box.com/s/vc7xn1juqed1nyi1a323t0ytma55wiyd). Here's a quick summary of what's covered
* Basic optimizations
* Exploiting Instruction-Level Parallelism (ILP)
  * Hazards: Structural, data dependency, control
* Pipelining
* Loop unrolling
* Loop unrolling with reassociation
* Loop unrolling with separate accumulators
* Vector instructions (use fat registers to perform the same operation on all variables stored in register)
* Branch prediction
* On an unrelated note: I took [CS 61C @ Berkeley](https://cs61c.org/su19/) in summer 2019, and one of the project assignments is [Performance Programming](https://cs61c.org/su19/projects/proj4/) and the optimization techniques used include:
  * Profiling & Amdahl's Law
  * Unrolling & Other Optimizations
  * SIMD Instructions
  * OpenMP
* I transferred the credits of CS 61C to meet the requirements for CS 354 at UW-Madison. Looking back from the top of the mountain, CS 61C really covered most of the stuff in CS 252, 354, a bit of 352, 552, and 537, and even things in 759. I'm not saying that our courses are bad (the classes I mentioned above, particularly the more advanced ones, spent more time going in-depth) but damn. Berkeley is so good. So good.

## MPI

![Nomenclatures](/files/-MXzKxWKdKW5Va4xfc-d)

![HPC vs. HTC](/files/-MXzLPLgmuEsY6zoInvO)


# Lecture 26: MPI Parallel Programming General Introduction. Point-to-Point Communication.

## Lecture Summary

* Last time
  * Wrapped up “Critical thinking” segment. Went through a case study, saw more than 100X speed up
    * No function-call blockers; loop unrolling & re-association; dropped in the wide-register vectorization
  * Started discussion about parallel computing via message passing (multi-process parallel computing)
    * Covered the hardware aspects related to HPC
* Today
  * HPC via MPI: discuss the basic ideas/paradigms
  * MPI point-to-point communication

## MPI

### Introduction to message passing and MPI

* CUDA: A kernel (a small snippet of code) is run by all threads spawned via an execution configuration
* OpenMP: All threads execute an omp parallel region, work sharing
* MPI: The entire code is executed in parallel by all processses

![Hello world example](/files/-MX3gd3GKGnXGim_NkrY)

* MPI does branching based on the process rank
  * Very similar to GPU computing, where one thread does work based on its thread index
  * Very similar to OpenMP function omp\_get\_thread\_num()
* Each MPI process has its own program counter and virtual address space
  * The variables of each program have the same name but live in different virtual memories and assume different values
* MPI can be used whenever it is possible for processes to exchange messages:
  * Distributed memory systems
  * Network of workstations
  * One workstation with many cores
    * Data is passed through the main memory instead of a network
    * Different ranks share the same physical memory, but they are each tied to separate virtual memory spaces

![](/files/-MX3kYdc8uvhnJ4pv7Gj)

![](/files/-MX3hv6xDbFeVVP5Xegd)

![MPI vs. OpenMP](/files/-MX3l2r_nkd6Wd9V93Em)

![MPI vs. CUDA](/files/-MX3l9crH6uWsXn6vYnF)

### Point-to-Point (P2P) Communication

![](/files/-MX3skqX-QnLudR_KeN1)

* P2P: Simplest form of message passing communication
* One process sends a message to another process (MPI\_*Send, MPI\_*&#x52;ecv)
* `int MPI_Send(void *buf, int count, MPI_Datatype datatype, int dst, int tag, MPI_Comm comm)`
  * buf: starting point of the message with count elements, each described with datatype
  * dst: rank of the destination process within the comm communicator
  * tag: used to distinguish between different messages
* `int MPI_Send(void buf, int count, MPI_Datatype datatype, int dest, int tag, MPI_Comm comm, MPI_Status status)`
  * Envelope information is returned in an MPI\_Status object
* A custom communicator can be created using `MPI_Comm_create(MPI_COMM_WORLD, new_group, &MY_COMM_WORLD);`
* MPI data types and their C counterparts: see table below
* The order of messages is preserved, i.e. messages do not overtake each other
* Receiver can wildcard to received from any source/tag: MPI\_ANY\_SOURCE/MPI\_ANY\_TAG
* For a communication to succeed:
  * Sender must specify a valid destination rank
  * Receiver must specify a valid source rank
  * The communicator must be the same
  * Tags must match
  * Message data types must match
  * Receiver's buffer must be large enough
* MPI\_Send and MPI\_Recv are blocking: when a process sends, it does not stop until another process receives
* Eager mode vs. Rendezvous mode
  * Eager mode: Small messages, the content of the buffer is picked up right away by the MPI runtime
  * Rendezvous mode: Large amount of data, the sender function waits for the receiver to post a receive before the runtime facilitates the sending of the actual data of the message

![MPI data types](/files/-MX3tJjjY55EuLy0SLU_)


# Lecture 27: MPI Parallel Programming Point-to-Point communication: Blocking vs. Non-blocking sends.

## Lecture Summary

* Last time
  * HPC via MPI
  * MPI point-to-point communication: The blocking flavor
* Today
  * Wrap up point-to-point communication
  * Collective communication

## Point-to-point communication

* Different "send" modes:
  * Synchronous send: MPI\_SSEND
    * Risk of deadlock/waiting -> idle time
    * High latency but better bandwidth than bsend
  * Buffered (async) send: MPI\_BSEND
    * Low latency/bandwidth
  * Standard send: MPI\_SEND
    * Up to the MPI implementation to device whether to do rendezvous or eager
    * Less overhead if in eager mode
    * Blocks in rendezvous, switches to sync mode
  * Ready send: MPI\_RSEND
    * Works only if the matching receive has been posted
    * Rarely used, very dangerous
* Receiving, all modes: MPI\_RECV
* Buffered send
  * Reduces overhead associated with data transmission
  * Relies on the existence of a buffer. Buffering incurs an extra memory copy&#x20;
  * Return from an MPI\_Bsend does not guarantee the message was sent: the message remains in the buffer until a matching receive is posted

![Blocking options](/files/-MXy2Z8FtyKIQUNzADAX)

![Deadlocks](/files/-MXy3PpYijHCTsi8sDFs)

## Non-blocking point-to-point

* Blocking send: Covered above. Upon return from a send, you can modify the content of the buffer in which you stored data to be sent since the data has been sent
* Non-blocking send: The sender returns immediately, no guarantee that the data has been transmitted
  * Routine name starts with MPI\_I
  * Gets to do useful work (overlap communication with execution) upon return from the non-blocking call
  * Use synchronization call to wait for communication to complete
* MPI\_Wait: Blocks until a certain request is completed
  * Wait for multiple sends: Waitall, Waitany, Waitsome
* MPI\_Test: Non-blocking, returns quickly with status information
  * int MPI\_Test(MPI\_Request \*request, int \*flag, MPI\_Status \*status);
* MPI\_Probe: Allows for incoming messages to be queried prior to receiving them

![](/files/-MXy4XOzQrNTX04U96sL)

## Collective communications

* Three types of collective actions:
  * Synchronization (barrier)
  * Communication (e.g., broadcast)
  * Operation (e.g., reduce)
* [Writing distributed applications with PyTorch](https://pytorch.org/tutorials/intermediate/dist_tuto.html) is a good tutorial
* Broadcast: MPI\_Bcast
* Gather: MPI\_Gather
* Scatter: MPI\_Scatter
* Reduce: MPI\_Reduce
  * Result is collected by the root only
* Allreduce: MPI\_Allreduce
  * Result is sent out to all ranks in the communicator
* Prefix scan: MPI\_Scan
* User-defined reduction operations: Register using MPI\_Op\_create()

![Visualization of the operations, excerpted from the Distributed PyTorch documentation](/files/-MXy6PzKtebTz3GBgGH7)

![Predefined reduction operations](/files/-MXy7--ToeHA3hJ8GXQJ)


# Lecture 28: MPI Parallel Programming: MPI Collectives. Overview of topics covered in the class.

## Lecture Summary

* Last time
  * Wrap up p2p communication
  * Collective communication: Synchronizations, communications, operations
* Today
  * Collective communication: Operations, data types
  * Talk about the final exam. [Past exam](https://uwmadison.app.box.com/s/vkwsrao6og5ocyno7u491n03kybhzdyn)

## MPI Derived Types (did not have time to cover in class)

* Previously we sent/received a contiguous buffer of identical elements of predefined data types
* Now we want to send non-homogenous elements (structure) or chunks that are not contiguous in memory
* MPI Datatypes
  * Primitive datatypes: MPI\_CHAR, MPI\_FLOAT, MPI\_INTEGER, etc.
  * Derived datatypes: Can be constructed by four methods (contiguous, vector, indexed, struct)
* Typemaps
  * Used to describe an MPI derived type
  * Specifies a sequence of primitive data types, and a sequence of integers that represent the byte displacements, measured from the beginning of the buffer
    * Typemap = {(type0, disp0), ..., (typen, dispn)}
* Extent: distance, in bytes, from beginning to end of type
  * E.g., {(double,0),(char,8)} has extent 16
* Type signature: The sequence of primitive data types
  * E.g., for a data type with typemap {(double,0), (int,8), (char, 12)}, its signature is {double, int, char}

![MPI type-definition functions (constructors)](/files/-MXyB0muOSakJ8W4WwTL)

## Takeaways from ME 759

* Know your hardware
* Moving data around is expensive in energy and time
* Seek solution approaches that expose concurrency/parallelism in your problem
* Use the tools of the trade to work like a pro and then get paid like a pro
  * Debuggers, profilers, CMake, memory checkers, compiler flags, git, etc.
  * Do HPC

![](/files/-MXyOPIxK2-lcxw-gVVt)

![](/files/-MXyORs4sjCPh60Duv7e)

This has been such a fun ride.


# Cloud Computing Course Notes

## Acknowledgments

* These reading notes cover the class [Cloud Computing Specialization](https://www.coursera.org/specializations/cloud-computing) on Coursera, offered by UIUC
  * Financial aids are available for this class. There is also the option to audit the class (no access to projects, quizzes, and the course certificate)
* The reading note are titled using the following format: `f"{course_num}.{week_num} {title}"`

## Table of Contents

### Course 1: [Cloud Computing Concepts, Part 1](https://www.coursera.org/learn/cloud-computing?specialization=cloud-computing)

* [Week 1: Introduction to Clouds, MapReduce](/earlier-readings-and-notes/cloud-computing-course-notes/1.1-introduction-to-clouds-mapreduce)
* [Week 2: Gossip, Membership, and Grids](/earlier-readings-and-notes/cloud-computing-course-notes/1.2-gossip-membership-and-grids)
* [Week 3: P2P Systems](/earlier-readings-and-notes/cloud-computing-course-notes/1.3-p2p-systems)
* [Week 4: Key-Value Stores, Time, and Ordering](/earlier-readings-and-notes/cloud-computing-course-notes/1.4-key-value-stores-time-and-ordering)
* [Week 5: Classical Distributed Algorithms](/earlier-readings-and-notes/cloud-computing-course-notes/1.5-classical-distributed-algorithms)

### Course 2: [Cloud Computing Concepts: Part 2](https://www.coursera.org/learn/cloud-computing-2?specialization=cloud-computing)

### Course 3: [Cloud Computing Applications, Part 1: Cloud Systems and Infrastructure](https://www.coursera.org/learn/cloud-applications-part1?specialization=cloud-computing)

### Course 4: [Cloud Computing Applications, Part 2: Big Data and Applications in the Cloud](https://www.coursera.org/learn/cloud-applications-part2?specialization=cloud-computing)

* [Week 1: Spark, Hortonworks, HDFS, CAP](/earlier-readings-and-notes/cloud-computing-course-notes/4.1-spark-hortonworks-hdfs-cap)
* Week 2: Large Scale Data Storage
* Week 3: Streaming Systems
* Week 4: Graph Processing and Machine Learning

### Course 5: [Cloud Networking](https://www.coursera.org/learn/cloud-networking?specialization=cloud-computing)


# 1.1 Introduction to Clouds, MapReduce

See my reading notes on the MapReduce paper:

{% content-ref url="/pages/-Md-Vwd5mgUMg9Jdql2w" %}
[\[2004 OSDI\] MapReduce: Simplified Data Processing on Large Clusters](/machine-learning-systems/index/mapreduce-simplified-data-processing-on-large-clusters)
{% endcontent-ref %}


# 1.2 Gossip, Membership, and Grids

## Lesson 1: Gossip

### Multicast Problem

* In computer networking, multicast is group communication where data transmission is addressed to a group of destination computers simultaneously
* Multicast can be one-to-many or many-to-many distribution
* The difference between multicast and broadcast is that in broadcast, the packet is delivered to all the hosts connected to the network, whereas in multicast, the packet is delivered to intended recipients only.
* The multicast protocol typically sits at the application layer (i.e., does not deal with the underlying network)
* Challenges
  * **Fault-Tolerance**
    * Nodes may crash
    * Packets may be dropped
  * **Scalability**
    * Tens of thousands of nodes
* Simplest implementation: Centralized
  * The sender sends in a loop UDP/TCP packets
  * Problems
    * Not fault-tolerant: Sender may fail. Say it fails halfway through, only half of the receivers get the message
    * High overhead: Not scalable -> high O(N) latency
* Solution: Tree-Based implementation
  * Pro: For a good (balanced) tree, the height is O(log(N)) -> better latency
  * Con: High set up and maintenance costs
* Tree-Based multicast protocols
  * Build spanning trees to disseminate multicasts
  * Use ACKs or NAKs to repair multicasts not received
  * SRM: Scalable Reliable Multicast
    * Uses NAKs
    * Uses random delays (before sending out repair request) and exponential backoff (if sending out multiple NAKs, doubles the wait time every time they wait) to avoid NAK storms
  * RMTP: Reliable Multicast Transport Protocol
    * Uses ACKs
    * ACKs are only sent to designated receivers, which then re-transmit missing multicasts
* Studies show that despite these countermeasures, these protocols still suffer from O(N) ACK/NAK overheads, which motivated the development of gossip/epidemic protocols

### Gossip Protocols

* There are two "hyperparameters": t and b. Say we set t to be 5 seconds and b (fan-out) to be 2 nodes. In the following examples, we consider only 1 multicast message and only 1 sender.
* Periodically (every t seconds), a sender picks b random targets and sends them the multicast/gossip message. We can use UDP to transmit the messages as the gossip protocol itself is very reliable.
* Once a node receives its gossip, it is said to be "infected" and becomes a sender.&#x20;
* The gossip protocol is not synchronized across nodes: each node uses its local clock to send messages in rounds. When doing analyses, we typically assume them to be synchronized, though.
* Those above described the "push" gossip: once you have a multicast message, you start gossiping about it.
  * There is also a "pull" gossip that periodically polls randomly selected processes for new multicast messages that haven't been received.&#x20;
  * Another variant is the push-pull model. In this model, when sending out a pull query, the sender also includes some gossip messages it received recently.
* Multiple messages -> push a random subset/recently-received ones/high-priority ones

![In this example, we start with one sender at the bottom left corner](/files/-MdHa4TBwp1PB8cNNiLt)

### Gossip Analysis

* The simple push protocol:
  * Is lightweight in large groups
  * Spreads a multicast quickly
  * Is highly fault-tolerant
* Analyze using Epidemiology
  * Population: (n+1) nodes
  * The contact rate between any individual pair is ß
  * At any time, each individual is either uninfected (x) or infected (y)
  * x\_0 = n, y\_0 = 1. At all times, x + y = n + 1
* Model as continuous time process. Do some math, and the conclusion is that when t becomes very large (as time progresses), x goes to 0, and y goes to n + 1. I.e., eventually, everyone receives the gossip. We can also show that the gossip protocol is fast: it converges within a logarithmic number of rounds.
* Recap
  * Lightweight: Each node transmits no more than cblog(n) gossip messages
  * Low latency: Converges within clog(n) rounds
  * Reliability: All but 1/(n^(cb-2)) nodes receive the multicast
* While log(n) is not constant, it grows very slowly pragmatically (e.g., using base = 2, log(1000) \~= 10, log(all IPv4 addresses) = 32).
* Packet loss: with 50% packet loss, analyze with b /= 2. To achieve the same reliability as 0% loss rate, take twice as many rounds
* Node failure: with 50% nodes failing, analyze with n /= 2 and b /= 2. Same as above
* Fault tolerance: with failures, it is possible (but improbable) that the epidemic dies out quickly. If it happens, it happens early, but as gossips spread very fast, it is very difficult to kill a gossip after a few rounds (just like pandemics like COVID-19/rumors on the internet!)
* In all forms of a gossip, it takes O(log(n)) rounds before n/2 nodes get the gossip (because the fastest structure for a message to spread is a spanning tree). Thereafter, pull gossip is faster than push gossip. The second half of pull gossip finishes in time O(log(log(n))). Some more math is involved here...
* Gossip protocols are not topology-aware -> core switches may get overloaded (O(n)). In this example, there are two subnets/racks. If nodes select targets randomly, half of the gossips will go through the router. The fix is to have gossips prefer nodes in the local subnet using a higher probability and vice versa. E.g., in subnet i with n\_i nodes, pick gossip target in the subnet with probability (1 - 1/n\_i). With this fix, the router load becomes O(1), and the dissemination time is still O(log(n)).

![](/files/-MdUt4K7CIspJR4i7Cvc)

### Gossip Implementations

* Some implementations
* Example: NNTP Inter-Server Protocol

![Some implementations](/files/-MdUtP1jjZhx69AAI-0A)

![NTTP inter-server protocol](/files/-MdUtWJxV0P_bdtUIFZk)

## Lesson 2: Membership

### What is Group Membership List?

* In data centers, failures are the norm, not the exception. For example, if the rate of failure of one machine (OS/disk/motherboard/network, etc.) is once every 10 years, a DC with 12000 servers has a mean time to failure (MTTF) of \~7.2 hours!
* Thus, we need a mechanism to detect failures. Preferably, a failure detector program distributedly, and automatically detects failures and reports to your workstation.

![](/files/-MdUxauT0SlbsH0fSNbv)

* Two sub-protocols are needed by a membership protocol:
  * A failure detector
  * A mechanism to disseminate information about joins, leaves, and failures of processes

![](/files/-MdUzecvI6sYqRgTOO1n)

### Failure Detectors

* Two correctness properties for failure detectors:
  * **Completeness**: Each failure is detected (eventually by one other non-faulty process)
  * **Accuracy**: There is no mistaken detection
  * In reality, in lossy networks, it is impossible to achieve both 100% completeness and 100% accuracy. Otherwise, we can solve consensus (TODO: figure out what this is). In real life, failure detectors guarantee completeness while only guaranteeing partial/probabilistic accuracy.
* Two desirable properties:
  * **Speed**: Time until some process first detects a failure
  * **Scale**: Equal load on each member, network message load
* We want to satisfy all the above properties in spite of arbitrary, simultaneous process failures
* A very simple failure detection protocol, **centralized heartbeating**:
  * Heartbeats sent periodically
  * If a heartbeat is not received from process p within timeout, mark p as failed
  * All processes send heartbeats to one central process
  * Cons
    * The central process may fail
    * The central process may be overloaded in a large process group
* A variant: **ring heartbeating**
  * Cons
    * Unpredictable on simultaneous multiple failures: If both neighbors of a process p fail, before the neighbors are repaired, p may fail undetected
* A third variant: **all-to-all heartbeating**
  * Pros
    * Equal load per member
    * Guarantees completeness (as long as there is at least one non-faulty process in the group)
  * Cons
    * If there is one straggler process that receives packets at a longer delay than others, it may mark all other processes as failed, leading to a low accuracy/high false-positive rate
    * How to improve the robustness is covered in the next lecture

![](/files/-MdV4WZaGvsbDdSZxo9E)

### Gossip-Style Membership

* Gossip-style heartbeating: a more robust (better accuracy properties) variant of all-to-all heartbeating
  * Each node maintains a membership list with each entry being \[node address, hearbeat counter, local time]
  * Nodes periodically gossip their membership list
  * On receipt, the local membership list is updated
    * Those entries with a higher/newer heartbeat counter is updated. The new time is the current, local time at recepient nodes
  * When an entry is last updated/heartbeat has not increased more than T\_fail seconds ago, it is marked as failed
  * After T\_cleanup seconds, the member is deleted from the list
  * Without this two-stage cleanup mechanism, a failed node has its entry deleted right after it is detected as failed. However, other processes may not have detected the failure/deleted the entry, and it may be added back in a gossip.
* Analysis: tradeoff between false positive rate, detection time, and bandwidth
  * A single heartbeat takes O(log(N)) time to propagate
  * If bandwidth allowed per node is O(N), N heartbeats takes O(log(N)) time to propagate
  * If bandwidth allowed per node is O(1) (only a few sampled entries of the membership list), N heartbeats takes O(Nlog(N)) time to propagate (inversely proportional).
  * If the gossip period T\_gossip is decreased:
    * We have a higher bandwidth/send out gossips much quicker
    * We can have a shorter time for the failure detection time T\_fail and T\_cleanup
    * As a result, we have a higher false positive rate, as non-faulty nodes (that are mistakenly detected) are given slightly shorter time for their heartbeat to make it across

![](/files/-MdZA-R6ukQDkU6w36Yb)

### Which is the best failure detector?

* Metrics of comparisons
  * Completeness: we want it always guaranteed
  * Speed: denote the time to first detection of a failure to be T seconds
  * (In)Accuracy: denote as PM(T), probability of mistake in time T. In other words, this is the probability that a mistaken detection will be made in T time units.
  * Given the above requirements, we will compare the network message load, N\*L, across protocols
* All-to-all heartbeating
  * The load is linear per node: L = N / T (N heartbeats sent out every T time units)
* Gossip-based all-to-all heartbeating
  * Gossip period is every tg unit seconds, where O(N) gossip messages are sent
  * T = log(N) \* tg (gossip takes O(log(N)) rounds to propagate)
  * L = N / tg = N \* log(N) / T
  * Higher load compared to all-to-all heartbeating: better accuracy by using more messages
* What's the theoretical optimal?
  * Optimal L is independent of N (?!)
  * All-to-all and gossip-based protocols are sub-optimal (L = O(N / T))
  * Main reason: these two protocols mix up the failure detection and dissemination components. The keys to getting close to this optimal bound are:
    * Separate the two components
    * Use a non heartbeat-based failure detection component

![](/files/-MdZDu7twDZwgWGM_Yb0)

### Another Probabilistic Failure Detector

* SWIM: Scalable Weakly-consistent Infection-style Membership protocol
  * Instead of using heartbeating, we use pinging
  * Process pi runs the protocol every T time units (protocol period)
  * At the beginning of a protocol, pi randomly picks a process pj and sends a ping message
  * If everything goes well, pj responds with an ACK
  * If the ACK is not heard back (original ping/ACK is dropped), pi tries to ping pj again using indirect paths
    * pi sends pings to K randomly selected processes, each of which sends a direct ping to pj and sends an ACK back to pi. If one ACK is received by pi, then we are good
    * Otherwise, pi marks pj as failed
    * The reason for using indirect paths is that the pi-pj path may be congested, and it might be dropping more packets than other paths. Using indirect paths bypasses the potential congestion (spatial chance) and gives pj a second (temporal) chance.

![](/files/-MdZI2vY7Wd8qjoB2FeF)

![](/files/-MdZINTdASN9vO0dcRYw)

### Dissemination and suspicion

* Dissemination options
  * Multicast (Hardware/IP)
    * Unreliable
    * Multiple simultaneous multicasts
  * Point-to-point (TCP/UDP)
    * Expensive
  * Zero extra messages: Piggyback on Failure Detector messages
    * Infection-style dissemination
      * Maintain a buffer of recently joined/evicted processes
        * Piggyback from this buffer
        * Prefer recent updates
      * Buffer elements are garbage collected after a while
* Suspicion mechanism
  * False positives might be due to
    * Perturbed processes
    * Packet losses, e.g. from congestion
    * Indirect pinging may not solve the problem (correlated message losses near pinged host)
  * Solution: suspect a process before declaring it as failed in the group (see state diagram below)
  * To distinguish multiple suspicions of a process, use per-process incarnation numbers
    * Higher inc# notifications override lower inc#'s
    * Within an inc#: (Suspected, inc#) > (Alive, inc#)
    * (Failed, inc#) overrides everything else

![](/files/-MdZnCP_vS4yhJTGNiFl)

## Lesson 3: Grids

### Grid Applications

* Example: RAMS (Rapid Atmospheric Modeling System)
  * Compute-intensive computing (or HPC)
  * Can such programs be run without access to a supercomputer?
  * See picture below for a set of distributed computing resources in a grid

![Distributed computing resources. For example, in the UW-Madison (!) CS lab, at night, the workstations can be harvested for running Grid applications](/files/-MdZsA5Cm4Ge4blhsRUh)

* Grid applications...
  * May have several GBs of intermediate data
  * May take several hours/days
  * Have four stages: Init, Stage in, Execute, Stage out, Publish (optional)
  * Are computationally intensive, massivelly parallel
* The core problem comes down to scheduling and resource allocations

![Example application expressed as DAGs](/files/-MdZtEUjJoXpw-K6enSu)

### Grid Infrastructure

* 2-level scheduling infrastructure
  * Intra-site: for example, UW-Madison uses HTCondor protocol
    * Such protocols are responsible for:
      * Internal allocation & scheduling
      * Monitoring
      * Distribution and publishing of files
    * HTCondor:
      * Runs on a lot of workstations
      * When workstation is free, ask site's central server (or Globus) for tasks
      * If user hits a keystroke, the task is stopped (either killed or asked to reschedule)
  * Inter-site: e.g., Globus protocol
    * It is responsible for:
      * External allocation & scheduling
      * Stage in & stage out of files
    * Internal structures of different sites may be invisible to Globus
* Globus toolkit
* Grids are federated, i.e. no single entity controls the entire infrastructure
* Architectures & key concepts have a lot in common with those of clouds


# 1.3 P2P Systems

## Lesson 1: P2P Systems

### P2P Systems Introduction

* Why study P2P systems?
  * P2P systems are the first distributed systems that seriously focused on scalability (w\.r.t #nodes)
  * P2P techniques abound in cloud computing systems. E.g., key-value stores use Chord p2p hashing (consistent hashing)

### Napster

* When users upload files, the files are stored at client machines ("peers")
* The Napster servers store directory information (a list of \<filename, ip\_addr, port\_num>)
* Napster search
  * Client sends server keywords to search with
  * Server searches (using ternary tree algorithm) and returns a list of hosts \<ip\_addr, port\_num> to client
  * Client pings each host in the list to find transfer rates
  * Client fetches file from best host
* All communication uses TCP
* Joining a P2P system
  * Send an HTTP request to a well-known URL for that P2P service
  * Message routed to introducer, a well-known server that keeps track of some recently joined nodes in P2P system
  * Introducer initializes new peer's neighbor table
* Problems
  * Central servers: source of congestion/single point of failure
  * No security: plain messages and passwords
  * Indirect infringement: responsible for users' copyright violation

![Napster structure](/files/-MddILptSmYJcN87sZZF)

### Gnutella

* Different from Napster, Gnutella eliminates the servers and have clients act as servers (servents), such that client machines search and retrieve amongst themselves
* In the overlay graph (overlay in the sense that it is overlayed on top of the internet), peers being neighbors means that they know about each other's ip addr and port num, and can send them messages
* Gnutella routes different messages within the overlay

![](/files/-MdeDo75xXfmEaC_CeJg)

* There are five main message types in the Gnutella protocol
  * Query (search)
    * Queries are flooded out (forwarded to all peers except the peer from which the Query was received), TTL-restricted, and are forwarded only once
  * QueryHit (response to query)
    * A QueryHit messages contains:
      * Info about responder: \<port, ip\_addr, speed>
      * Results: \<fileindex, filename, fsize>
      * servent\_id: Unique identifier of responder (a function of its ip addr)
    * QueryHits are reverse-routed: If A sends B a Query and B got a hit, B sends to A a QueryHit
  * Push (used to initiate file transfer)
    * After QueryHits are received, the requestor chooses the "best" responder, and then initiates HTTP request directly to responder's ip\_addr:host
    * IRL, responders may be behind firewalls that rejects incoming connections
    * If a HTTP request fails, it routes a Push message via links in the overlay. The Push message contains ip\_addr:host at which the requestor can accept incoming connections. When the peer receives this Push message, it can generate an outgoing TCP connection (sends GIV, receives GET)
    * If the requestor is also behind a firewall, Gnutella gives up
      * Alternative: use a modified version of Gnutella to transfer the file via the overlay links themselves (this might be slow)
  * Ping (to probe network for other peers)
    * Peers initiate Pings periodically, and Pings are flooded out
  * Pong (reply to ping, contains address of another peer)
    * Pongs are routed along reverse paths
    * Pongs are used to keep neighbor lists fresh in spite of peers joining, leaving, and failing
* Problems
  * Ping/Pong constitute 50% of the traffic
    * Solutions: Multiplex, cache, and reduce frequency
  * Repeated searches with same keywords
    * Solutions: Cache query, QueryHits
  * Modem-connected hosts do not have enough bandwidth for passing Gnutella traffic
    * Solution: Use a central server to act as proxy for such peers
    * Another solution: FastTrack
  * Large number of freeloaders (only download files, never upload files)
    * In 2000, 70% of the users are freeloaders
  * Flooding causes excessive traffic
    * To maintain meta info about peers in order for more intelligent routing, use structures P2P systems (e.g., Chord)

![Gnuteella message header format](/files/-MdeF35r6MZyuyI6MqSr)

### FastTrack and BitTorrent

#### FastTrack

* Hybrid between Napster and Gnutella, takes advantages of "healthier" participants in the system
* Like Gnutella, but designate some peers as "supernodes"&#x20;
  * A supernode stores a directory listing a subset of nearby \<filename, peer pointer> (similar to Napster servers)
  * Supernode membership changes over time
  * Any node may become a supernode, provided it has earned enough reputation
    * E.g., reputation is affected by length of periods of connectivity and total number of uploads
  * A peer searches by contacting a nearby supernode

![FastTrack structure](/files/-MdkJzh8JktrXeW6Dk_8)

#### BitTorrent

* Files are split into blocks (32KB - 256KB)
* Download **Local Rarest First** block policy: Prefers early download of blocks that are least replicated among neighbors
* **Tit for tat** bandwidth usage: Provide blocks to neighbors that provided it the best download rates
  * Incentivizes nodes to provide good download rates
* **Choking**: Limit number of neighbors to which concurrent uploads <= a number (5), i.e. the best neighbors. Everyone else is choked
  * Prevents overloading of the upload bandwidth
  * Periodically (e.g., 10s) re-evaluate this set
  * Optimistic unchoke: Periodically (e.g., 30s) unchoke a random neighbor to keep the unchoked set fresh

![BitTorrent structure](/files/-MdkL9vy7Wn1FeZIe_R4)

### Chord

* [Original paper](https://pdos.csail.mit.edu/papers/chord:sigcomm01/chord_sigcomm.pdf)
* Distributed hash tables: objects = files
  * Performance concerns
    * Load balancing
    * Fault tolerance
    * Efficiency of lookups and inserts
    * Locality
  * Napster, Gnutella, and FastTrack are all DHTs
* Chord: Consistent hashing on nodes' addresses
  * SHA-1(ip\_addr, port) -> 160-bit string, truncated to m bits -> peer id
* Each node stores peer pointers
  * Successors
  * Finger tables
    * Used for routing queries quickly

![Comparative performance](/files/-MdkUrIsfjk5jid_gA68)

![Finger tables](/files/-MdkXFOIOP4ET9iSNGWO)

* Consistent hashing: With K keys and N peers, each peer stores O(K/N) keys&#x20;
* Storing files
  * Filenames are also mapped using the same consistent hash function
  * File is stored at first peer with id greater than or equal to its key (mod 2^m)
* Searching files
  * Takes O(log(N)) time

![Chord searching](/files/-MdkaPR5NO0Cpq-YBetv)

### Failures in Chord

* Solution 1: Maintain multiple (2log(N)) successor entries

![Yes, more math](/files/-Mdkg4HMEwmmyf4uT7Mz)

* Solution 2: Replicate file/key at r successors and predecessors
* Dealing with dynamic changes (P2P systems have a high rate of churn: peers joining, leaving, and failing)
  * Stabilization protocol is run by all nodes periodically (talk to neighbors to update finger table)
  * New peers may need to copy some files/keys from other nodes
  * A new peer affects O(log(N)) other finger entries in the system
    * Number of messages per peer join = O(log(N) \* log(N))
  * Concurrent peer joins/leaves/failures
    * Chord peers periodically run a stabilization algorithm that checks and updates pointers and keys, which ensures non-loopiness
  * Hash can get non-uniform -> bad load balancing
    * Solution: Virtual nodes (treat each node as multiple virtual nodes behaving independently)

### Pastry

* Just like Chord, assigns ids to nodes using a virtual ring
* Leaf set: Each node knows its successors and predecessors
* Routing table: Instead of "n+2^i" rule in Chord, use prefix matching -> log(N)
  * Consider a peer with id 01110100101. It maintains a neighbor peer with an id matching each of the following prefixes: {\*, 0\*, 01\*, 011\*, ..., 0111010010\*}
    * For each prefix, among all the potential neighbors, the neighbor with the shortest RTT is selected
    * Early hops/shorter prefixes have many more candidates -> likely to be closer -> hops are short, yet overall stretch (compared to direct Internet paths) stays short
  * When it needs to route to a peer (e.g., 01110111001), it forwards to a neighbor with the largest matching prefix (011101\*)
* Problems
  * O(log(N)) lookup hops may be high

### Kelips

* Constant lookup cost to DHT
* Instead of virtual rings, we use k (\~= sqrt(N)) affinity groups
* Each node is hashed (mod k) to a group
* A peer is neighbors with all other nodes in its affinity group
* Files are stored at whichever node uploaded them
  * Kelips decouples file replication/location from querying
  * Each filename hashed to a group
  * All nodes in the group replicate pointer information (i.e., \<filename, location>)
* Lookup
  * Find affinity group
  * Go to your contact for the file affinity group
    * If fails, try another neighbor to find a contact
  * Lookup = 1 hop (or a few under failures)
  * Memory cost: O(sqrt(N))
    * 1.93MB for 100K nodes, 10M files

![Kelips structure](/files/-MdkyZUYJebQia45C5SK)

![Kelips soft state](/files/-Mdl-YMczsXdkepMv20I)

### Summary

* Chord & Pastry & Kelips
  * Range of tradeoffs (memory vs. lookup cost vs. background bandwidth (in order to keep neighbors fresh))
    * Chord & Pastry use O(log(N)) for both memory & lookup
    * Kelips uses more memory (O(N^2)) & background bandwidth to provide O(1) lookup
  * All of them have provable properties

### One of the questions I in the discussion thread

> Hi, I have a question regarding question 8 in HW 3. My reasoning is as follows. Going down, we have 3 at level 3, 9 at level 4, 27 at level 5, making a total of 39. Going up, we have 1 at level 1, 2 at level 2, 6 at level 3, making a total of 9. Adding them up, it should be 48. I would appreciate it if someone can point out the mistake in my reasoning. Thanks in advance!
>
> To provide more context, the question is:
>
> A Gnutella topology looks like a balanced ternary tree with 4 levels of nodes, i.e., peers, as shown in the picture below. Thus, there is 1 root at Level 1, which has 3 children at Level 2, which each have 3 children at Level 3, which in turn each have 3 children at Level 4 – thus, there are a total of 40 nodes. If a child of the root (i.e., a Level 2 node in the tree) sends a Query message with TTL=3, then what are the number of nodes receiving the Query message, not including the originating node? Enter your answer as a numeric value in the text box below. (1 point)


# 1.4 Key-Value Stores, Time, and Ordering

## Lesson 1: Key-Value Stores

### Why Key-Value / NoSQL?

* The Key-Value Abstraction
  * Twitter: Tweet ID -> Info about tweet
  * Amazon: Item ID -> Info about it
  * Chase: Account # -> Info about it
* Kind of like a distributed dictionary/DHT in P2P systems
* Kind of like a database
  * Why not use relational DBMS? Mismatch with today's workloads
    * Data: Large and unstructured
    * Lots of random reads and writes from lots of clients
    * Sometimes write-heavy, while RDBMS are often optimized for reads
    * Foreign keys rarely needed
    * Joins frequent
* Demands of today's workloads
  * Speed
  * Avoid Single Point of Failure (SPoF)
  * Low Total Cost of Operation (TCO)
  * Fewer System Administrators
  * Incremental Scalability
  * Scale out, not up
    * Scale up: Grow the cluster capacity by replacing with more powerful machines
    * Scale out: Incrementally grow the cluster capacity by adding more COTS machines (Components Of The Shelf, sweet spot on the price curve)
      * This is cheaper, and we can phase in/out newer/older machines over a long duration
* NoSQL: Not Only SQL
  * Necessary API operations: get(key) and put(key, value)
  * There are tables like in RDBMS systems, but they may be unstructured/may not have schemas/don't always support joins/foreign keys, but they can have index tables
  * Storage: column-oriented storage
    * RDBMS stores an entire row together (on disk or at a server)
    * NoSQL systems store a column (or groups of columns) together
      * Entries within a column are indexed and easy to locate given a key (and vice versa)
      * This makes ranged searches within a column faster (as we don't need to fetch the entire database)
        * E.g., get me all the blog\_ids from the blog table that were updated within the past month&#x20;

![NoSQL structure](/files/-Mdyli6aWKLE4pc7c0QJ)

### Cassandra

* Data placement strategies
  * SimpleStrategy
    * RandomPartitioner: Chord-like hash partitioning
    * ByteOrderedPartitioner: Assigns ranges of keys to servers, easier for range queries
  * NetworkTopologyStrategy: For multi-DC deployments
    * Two/three replicas per DC
    * Per DC: The first replica is placed according to partitioner, then go clockwise until you hit a different rack
* Snitches: Maps IPs to racks and DCs
  * SimpleSnitch: Unaware of topology/rack
  * RackInferring: Assumes network topology by octet of server's IP address
    * 101.102.103.104 = x.\<DC octet>.\<rack octet>.\<node octet>
  * PropertyFileSnitch: Uses a config file
  * EC2Snitch: EC2 region = DC, availability zone = rack
* Writes
  * Client sends write to one coordinator node in Cassandra cluster
  * Coordinator uses partitioner to send query to all replica nodes responsible for key
  * When X replicas respond, coordinator returns an acknowledgment to the client
    * X is specified by the client -- we'll come back to this later
  * Hinted Handoff mechanism: If any replica is down, the coordinator writes to all other replicas, and keeps the write locally until down replica comes back up, when it sends a copy of that write. When all replicas are down, the coordinator buffers the write locally

![Cassandra ring structure](/files/-MeFjbRTIqQTX6G2l6wm)

* When a replica nodes receives a write
  * Log it in disk commit log for failure recovery
  * Make changes to appropriate memtables, in-memory representations of multiple key-value pairs. Memtables are flushed to disk when they are full/old.
  * Data files: An SSTable (Sorted String Table), list of key-value pairs sorted by key
  * Index file: An SSTable of (key, position in data SSTable) pairs
  * Efficient search: Bloom filters!
* Bloom filters: Large bit maps
  * Checking for existence in set is cheap
  * Some probabilities of false positives (an item not in set reported as being in there -> incur a slight overhead for going into the SSTable), but never false negatives
  * The bit map starts with all zeros. On insert, we use k hash functions to map a key to k indexes. Then, all hashed bits are set to 1 for those k indexes (if they hadn't been set already)

![Bloom filters](/files/-MeFp383iLsXOBYCKyvi)

* Compaction
  * Each server periodically merges SSTables by merging updates for a key
* Delete
  * Instead of deleting right away, add a tombstone to the log, and eventually, it will be deleted by the compaction process
* Reads
  * Coordinator contacts X replicas
  * When X replicas respond, the coordinator returns the latest-timestamped value from those X replicas
  * The coordinator also fetches values from other replicas
    * This checks consistency in the background, initiating a read repair if any two values are different
    * This mechanism seeks to eventually bring all replicas up-to-date
  * A row may span across multiple SSTables -> reads need to touch multiple SSTables -> reads are slower than writes
* Membership
  * Any server could be the coordinator -> every server needs to know about all the servers in the cluster, and the list of servers needs to be updated as servers join/leave/fail
  * Cassandra uses gossip-style membership
* Suspicion mechanism: Sets timeouts on a server-by-server basis
* Reads/writes are orders of magnitudes faster than MySQL. But what did we lose?

### The Mystery of X-The Cap Theorem

* In a distributed system, at most two out of these three can be satisfied:
  * Consistency: All nodes see the same data at any time/reads return the latest value written by any client
    * Thousands of people booking the same flight
  * Availability: The system allows operations all the time & operations return quickly
    * Amazon: each extra ms of latency implies a $6M yearly loss
  * Partition-tolerance: The system continues to work in spite of network partitions (within/across DCs)

![CAP tradeoff](/files/-MeJP2LefcjHhjHe_oOz)

* RDBMS provides ACID: Atomicity, Consistency, Isolation, Durability
* KV Stores provides BASE: Basically Available Soft-state Eventual Consistency (prefers availability over consistency)
* In Cassandra, for each operation, a client is allowed to choose a consistency level
  * ANY: Any server (may not be replica)
    * Fastest (coordinator caches write & replies quickly)
  * ALL: All replicas
    * Strong consistency but slow
  * ONE: At least one replica
    * Faster than ALL but no failure tolerance (if all replica fails)
  * QUORUM: Quorum across all replicas in all DCs
    * Quorum = majority (>50%)
    * Any two quorums intersect
    * Faster than ALL while still guaranteeing strong consistency
  * More quorum-related
* Quorum (N = total number of replicas)
  * Read consistency level: R <= N, coordinator waits for R replicas to respond before sending result to client, while in background, coordinator checks for consistency of remaining (N-R) replicas
  * Write consistency level: W <= N. Two flavors: (1) Coordinator blocks until quorum is reached, (2) Async: just write and return
  * For strong consistency:
    * W + R > N (write & read quorums intersect in at least one server among replicas of a key)
    * W > N / 2 (two write quorums ..., which keeps the latest value of the write)

![Selecting values of W and R based on workloads](/files/-MeJTt4gz_AfLCjO_UkD)

### The Consistency Spectrum

* Cassandra offers eventual consistency: If writes to a key stop, all replicas of key will converge

![Left: Faster reads/writes. Right: More consistency.](/files/-MeJYMG20kA9_piwbxn8)

* Per-key sequential: Per key, all operations have a global order
* CRDT: Commutative Replicated Data Types, commutated writes give same result
  * Servers do not need to worry about consistency/ordering
* Red-Blue: Rewrites client operations and split them into red/blue ops
  * Red ops: Need to be executed in the same order at each DC
  * Blue ops: Can be commutated in any order across DCs
* Casual: Reads must respect partial order based on information flow
* Strong consistency models: Linearizability/Sequential consistency

### HBase

* API functions
  * Get/Put (row)
  * Scan (row range, filter) - range queries
  * MultiPut
* Prefers consistency over availability (unlike Cassandra)

![HBase architecture](/files/-MeJbGYUQHeCw2hRQ6cU)

![HFile structure](/files/-MeJc6xTtQtusO03TOgG)

* HBase uses write-ahead log (before writing to memstore) to ensure strong consistency
* Cross-datacenter replication: Single master + other slave clusters replicate the same tables

## Lesson 2: Time and Ordering

### Introduction and Basics

* Time synchronization is required for both correctness and fairness
* Challenges
  * End hosts in Internet-based systems like clouds have their own clocks
  * Processes in Internet-based systems follow an asynchronous system model (no bounds on message delays/processing delays)
* Clock skew/drift: Relative difference in clock values/frequencies of two processes
* MDR: Maximum Drift Rate of a clock. Between any pair of clocks, given a max acceptable skew M, need to synchronize every M / (2 \* MDR) time units

![Terminologies](/files/-MeJg2D04aVwVQrZp12N)

* Consider a group of processes:
  * External synchronization: Each process's clock is within a bound D of a well-known external clock (e.g., UTC)
  * Internal synchronization: Every pair of processes have clocks within bound D
  * External sync. within D implies internal sync. within 2D

### Cristian's Algorithm

* Process P synchronizes with a time server S
* Problem: Time response message is inaccurate, the inaccuracy being a function of message latencies (and since latencies are unbounded, the inaccuracy cannot be bounnded)
* Cristian's Algorithm measures the RTT of message exchanges
* The actual time at P when it receives response is between `[t + min2, t + RTT - min1]`
  * min1 = P -> S latency, min2 = S -> P latency
* Cristian's Algorithm sets its time to `t + (RTT + min2 - min1) / 2` (halfway through this interval)
  * Error is now bounded, being at most `(RTT - min2 + min1) / 2`

### NTP

* NTP = Network Time Protocol
* Each client is a leaf of the tree, each node synchronizes with its parent

![](/files/-MeJtCdhgZS-mcchVrZF)

* Suppose child is ahead of parent by oreal, and suppose one-way latency of message i is Li, and suppose offset `o = (tr1 - tr2 + ts2 - ts1) / 2`, then
  * `tr1 = ts1 + L1 + oreal`
  * `tr2 = ts2 + L2 - oreal`
  * Then, `oreal = o + (L2 - L1) / 2`, and then,
  * `|oreal - o| < |(L2 - L1) / 2| < |(L2 + L1) / 2|`, making the error bounded by RTT
* Can we avoid clock synchronization and still be able to order events?

### Lamport Timestamps

* As long as timestamps obey causality, we can assign to events timestamps that are not absolute time
* Happens-before is denoted as ->

![](/files/-MeKB7wHtaNYLC0kOBeC)

* Rules for assigning timestamps
  * Each process uses a local counter that is initialized as 0
  * A process increments its counter when a send/instruction happens
  * A send(message) event carries its timestamp
  * For a receive(message) event, the counter is updated by max(local clock, message timestamp) + 1
    * To obey the causality order
* Lamport timestamps are not guaranteed to be ordered or unequal for concurrent events
  * `E1 -> E2` implies `timestamp(E1) < timestamp(E2)`
  * `timestamp(E1) < timestamp(E2)` implies `{E1 -> E2} OR {E1 and E2 are concurrent}`
* Can we tell if two events are concurrent or casually related? -> Vector clocks!

![Lamport example 1](/files/-MeKR_6kAkunPLwb1koR)

![Lamport example 2](/files/-MeKRXC_2b8nlt7CIf6A)

![Lamport example 3 (from the final exam)](/files/-MgdwdUxBgJU_1zuq2v5)

### Vector Clocks

* N processes
* Each process uses a vector of integer clocks: Process i maintains `Vi[1, ..., N]`
* jth element of vector clock at process i, `Vi[j]`, is i's knowledge of latest events at process j
* Rules for assigning vector timestamps
  * On an instruction or send event at process i, it increments only the ith element of its vector clock
  * Each message carries the send event's vector timestamp
  * When process i receives a message
    * `Vi[i] += 1`
    * `Vi[j] = max(Vmessage[j], Vi[j]) for j != i`
* Casually-related:
  * `VT1 = VT2` iff `VT1[i] = VT2[i]` for all i = 1, ..., N
  * `VT1 <= VT2` iff `VT1[i] <= VT2[i]` for all i = 1, ..., N
  * Two events are casually related iff `VT1 < VT2`
    * i.e., iff `VT1 <= VT2` & there exists j such that `1 <= j <= N & VT1[j] < VT2[j]`
  * Two events are concurrent iff `NOT(VT1 <= VT2) AND NOT (VT2 <= VT1)`
    * Denote as `VT1 ||| VT2`

![Vector clock example 1](/files/-MeKResdtGnk-7v7Mgya)

![Vector clock example 2](/files/-MeKRhjDHUbPkgi2ZSny)

![Vector clock example 3 (from the final exam)](/files/-MgdwjedkuFCTY23iTAl)


# 1.5 Classical Distributed Algorithms

## Lesson 1: Snapshots

### What is Global Snapshot?

* In the cloud: Each application or service is running on multiple servers which handle concurrent events and interact with each other. Thus, the ability to obtain a "global photograph" of the system is important
* **Global snapshot = global state = individual state of each process/channel in the distributed system**
* First solution: Synchronize clocks of all processes
  * Ask all processes to record their states at know time t
  * Problems
    * Time synchronization always has error
    * Does not record the state of messages
  * Causality is enough!

### Global Snapshot Algorithm

* System model:
  * N processes in the system
  * Two uni-directional communication channels between each ordered process pair
  * Channels are FIFO
  * No failures
  * All messages arrive intact and are not duplicated
* Requirements
  * Snapshot should not interfere with/block normal application actions
  * Each process is able to record its own state
  * The global state is collected in a distributed manner
  * Any process may initiate the snapshot
* Chandy-Lamport Global Snapshot Algorithm
  * First, Initiator Pi records its own state
    * The initiator process creates special messages called "Marker" messages
    * For all other processes j, Pi sends out a Marker message on outgoing channel Cij (N-1 channels in total)
    * Starts recording the incoming messages on each of the incoming channels at Pi: Cji for j = 1 to N excluding i
  * Whenever a process Pi receives a Marker message on an incoming channel Cki
    * If this is the first Marker Pi is seeing
      * Pi records its own state first
      * Marks the state of channel Cki as "empty"
      * For j = 1 to N except i, Pi sends out a Marker message on outgoing channel Cij
      * Starts recording incoming messages on each of the incoming channels at Pi: Cji for j = 1 to N except i and k
    * Else (if it has already seen a Marker message):
      * Mark the state of channel Cki as all the messages that have arrived on it since recording was turned on for Cki
  * The algorithm terminates when
    * All processes have received a Marker (to record their own state)
    * All processes have received a Marker on all the N-1 incoming channels (to make sure each process has its state recorded)
  * Then, optionally, a central server collects all these partial state pieces to obtain the full global snapshot

### **Consistent Cuts**

* Cut = time frontier at each process and at each channel
* Events at the process/channel that happen before the cut are "in the cut"
* Consistent cut: A cut that obeys causality
  * A cut is consistent iff for each pair of events (e, f) in the system s.t. event e is in the cut C, and if f -> e, then f is also in the cut C

![](/files/-MgUOidJ7YBpdCYyNHhw)

* Any run of the Chandy-Lamport Global Snapshot algorithm creates a consistent cut
  * Proof by contradiction

### Safety and Liveness

* Liveness: Guarantee that something good will happen eventually
* Safety: Guarantee that something bad will never happen
* Can be difficult to satisfy both in an asynchronous distributed system
  * Failure detector: Completeness/liveness & accuracy/safety cannot both be guaranteed in an asynchronous distributed system
  * Consensus: Decisions/liveness and correct decisions/safety cannot both be guaranteed by any consensus protocol in an asynchronous distributed system
* Liveness w\.r.t. a property Pr in a given state S means
  * S satisfies Pr, or there is some causal path of global states from S to S' where S' satisfies Pr
* Safety w\.r.t. a property Pr in a given state S means
  * S satisfies Pr, and all global states S' reachable from S also satisfy Pr
* Chandy-Lamport algorithm can be used to detect global properties that are stable (once true, stays true forever afterwards)

## Lesson 2: Multicast

### Multicast Ordering

* Different communication forms
  * Multicast: Message sent to a group of processes
  * Broadcast: Message sent to all processes anywhere
  * Unicast: Message sent from one sender process to one receiver process
* FIFO Ordering
  * Multicasts from each sender are received in the order they are sent, at all receivers
  * Doesn't care about multicasts from different senders
* Casual Ordering
  * Multicasts whose send events are causally related must be received in the same causality-obeying order at all receivers
  * Concurrent multicasts are ok to be received in different orders at different receivers
  * Casual Ordering -> FIFO Ordering (the reverse is not true)
* Total Ordering/Atomic Broadcast
  * Ensures all receivers receive all multicasts in the same order
  * Doesn't care about the order of multicast sending
  * May need to delay delivery of some messages at sender

### Implementing Ordering

* FIFO Ordering
  * Each receiver Pi maintains a per-sender sequence number Pi\[1...N], initially all zeros
  * Pi\[j] is the latest sequence number Pi has received from Pj
  * Send multicast from Pj:
    * Pj\[j] += 1
    * Include new Pj\[j] in the multicast message
  * Receive multicast (Pi receives from Pj with sequence number S in the message):
    * If S == Pi\[j] + 1, then
      * Deliver message to application
      * Pi\[j] += 1
    * Else
      * Buffer this multicast until the above condition is true

![FIFO ordering example 1 (from the final exam)](/files/-Mgdx6U2_RJ3x9E2WHz2)

![FIFO ordering example 2 (from the final exam)](/files/-MgdxL_Daz-505IhuZIN)

![FIFO ordering example 3 (from the final exam)](/files/-MgdxbD_qtxNPFirYJq9)

* Casual Ordering
  * Each receiver Pi maintains a per-sender sequence number Pi\[1...N], initially all zeros
  * Send multicast from Pj:
    * Pj\[j] += 1
    * Include new **entire vector** Pj\[1...N] in the multicast message
  * Receive multicast (Pi receives from Pj with vector M\[1...N], buffer it until both:)
    * This message is the next one Pi is expecting from Pj, i.e. M\[j] = Pi\[j] + 1
    * All multicasts, anywhere in the group, which happened-before M, have been received at Pi, i.e.
      * For all k != j, M\[k] <= Pi\[k] (Receiver satisfies causality)
    * When the above two conditions are met, deliver M to application and set Pi\[j] = M\[j]

![Casual ordering example 1 (from the final exam)](/files/-Mgdxsm1w01sA-uxpR2l)

![Casual ordering example 2 (from the final exam)](/files/-Mgdy17fKV4YCErzEIkV)

![Casual ordering example 3 (from the final exam)](/files/-MgdyDuIgVRm9e1dn3I_)

* Total Ordering: Sequencer-based approach

![](/files/-MgV83P5UsyVHgUkrgx-)

### Reliable Multicast

* Reliable multicast loosely says that every process in the group receives all multicasts
  * Reliability is orthogonal to ordering
* When it comes to process failures, the definition becomes vague
* Definition: Need all correct/non-faulty processes to receive the same set of multicasts as all other correct processes

### Virtual Synchrony

* Each process maintains a membership list, called a View
* An update to this membership list is called a View Change
* Virtual synchrony guarantees that all view changes are delivered in the same order at all correct processes
  * A multicast M is said to be "delivered in a view V at process Pi" iff Pi receives view V, and then some time before Pi receives the next view, it delivers multicast M
* Views may be delivered at different physical times at processes, but they are delivered in the same order
* Virtual synchrony ensures that
  * The set of multicasts delivered in a given view is the same set at all correct processes that were in that view
    * **What happens in a view, stays in that view**
  * The sender of the multicast message also belongs to that view
  * If a process Pi does not deliver a multicast M in view V while other processes in the view V delivered M in V, the Pi will be forcibly removed from the next view delivered after V at the other processes
* TODO: Add some examples

## Lesson 3: Paxos

* TODO


# 4.1 Spark, Hortonworks, HDFS, CAP

## Spark

### Apache Spark

* Motivation: Traditional MapReduce & classical parallel runtimes cannot solve iterative algorithms efficiently
  * Hadoop: Repeated data access to HDFS, no optimizations to data caching & data transfers
  * MPI: No natural support for fault tolerance; programming interface is complicated
* Apache Spark: Extend the MapReduce model to better support two common classes of analytics apps:
  * Iterative algorithms (ML, graphs)
  * Interactive data mining
* Why are current frameworks not working?
  * Most cluster programming models use acyclic data flow (from stable storage to stable storage)
  * Acyclic data flow is inefficient for apps that repeatedly reuse a working set of data
* Solution: **Resilient Distributed Datasets (RDDs)**
  * Advantages
    * Allow apps to keep working sets in memory for efficient reuse
    * Retains the attractive properties of MapReduce (fault tolerance, data locality, scalability)
    * Supports a wide range of applications
  * Properties
    * Immutable, partitioned collections of objects
    * Created through parallel transformations (map, filter, groupBy, join) on data in stable storage
    * Can be cached for efficient reuse

### Example Spark Applications

![Log mining](/files/-MgdaZIQmn72Ul6Epq8L)

![Logistic regression](/files/-MgdagZZgm_wOw9_id6k)

### RDD Fault Tolerance

![RDDs maintain lineage information that can be used to reconstruct lost partitions](/files/-Mgdb6Q5zFpQOkdWf2EG)

## Big Data Distros (Distributions)

### Hortonworks

* Connected data strategy
  * HDP: Apache Hadoop is an open-source framework for distributed storage and processing of large sets of data on commodity hardware. Hadoop enables businesses to quickly gain insight from massive amounts of structured and unstructured data
  * HDF: Real-time data collection, curation, analysis, and delivery of data to and from any device, source or system, either on-premise and in the cloud

![Hortonworks Data Platform (HDP)](/files/-Mgddl1alNQL7se_yHRj)

* HDP tools
  * Apache Zeppelin: Open web-based notebook that enables interactive data analytics
  * Apache Ambari: Source management platform for provisioning, managing, monitoring, and securing Apache Hadoop clusters
* HDP data access
  * YARN: Data Operating System
    * MapReduce: Batch application framework for structured and unstructured data
    * Pig: Script ETL data pipelines, research on raw data, and iterative data processing
    * Hive: Interactive SQL queries over petabytes of data in Hadoop
    * Hbase Accumulo: Non-relational/NoSQL database on top of HDFS
    * Storm (Stream): Distributed real-time large volumes of high-velocity data
    * Solr (Search): Full-text search and near real-time indexing
    * Spark: In-memory
  * Data management: HDFS
* HDF
  * Apache NiFi, Kafka, and Storm: Provide real-time dataflow management and streaming analytics

### Cloudera

![Cloudera Enterprise Data Hub (EDH)](/files/-MgdhYILVQW3M5Gw2ayD)

### MapR

* Platforms for big data
  * MapReduce (Hadoop written in C/C++)
  * NFS
  * Interactive SQL (Drill, Hive Spark SQL, Impala)
  * MapR-DB
  * Search (Apache Solr)
  * Stream Processing (MapR Streams)

## HDFS

### HDFS

* HDFS properties
  * Synergistic w/ Hadoop
  * Massive throughput
  * Throughput scales with attached HDs
  * Have seen very large production clusters (Facebook, Yahoo)
  * Doesn't even pretend to be POSIX compliant
  * Optimized for reads, sequential writes, and appends
* How can we store data persistently? Ans: Distributed File System replicates files
* Distributed File System
  * Datanode Servers
    * A file is split into contiguous chunks (16-64MB), each of which is replicated (usually 2x or 3x)
    * Sends heartbeat and BlockReport to namenode
  * Replicas are placed: one on a node in a local rack, one on a different node in the local rack, and one on a node in a different rack (lots of back-ups)

![HDFS architecture](/files/-MgdkHcPaKzeFoNq8aVb)

* Master node (namenode in HDFS) stores metadata, and might be replicated
  * Client libraries for file accesses talk to master to find datanode chunk, and then connect directly to datanode servers to access data
* Replication pipelining: Data is pipelined from datanode to the next in the background
* Staging: A client request to create a file does not reach namenode immediately. Instead, HDFS client caches the data into a temporary file -> once the data size reaches a HDFS block size, the client contacts the namenode -> namenode inserts the filename into its hierarchy and allocates a data block for it -> namenode responds to the client with the identity of the datanode and the destinations of the replicas/datanodes for the block -> client flushes from local memory

### YARN and Mesos

* Mesos: Built to be a scalable global resource manager for the entire datacenter

![Mesos architecture](/files/-MgdoJo_AbvEfUWf8ic6)

![Mesos resource offer mechanism](/files/-MgdoS6nygcz_vnVPMAn)

* YARN: Created out of the necessity to scale Hadoop

![YARN ResourceManager](/files/-Mgdnf8ei7n7-SoybUkJ)

![The insides of YARN](/files/-Mgdntfe7Jz9S8-8eCFx)

* Project myriad: Composites Mesos and YARN
  * Mesos framework and a YARN scheduler that enables Mesos to manage YARN resource requests

![](/files/-Mgdouk_cyV2Nu2wOBxl)


# 4.2 Large Scale Data Storage

## MapReduce

### Motivation

* Challenges w/ traditional programming models (MPI)
  * Deadlock is possible: Blocking communication can cause deadlock
  * Large overhead from communication mismanagement
  * Load imbalance
  * Hard to code
* Challenges with commodity clusters
  * Web datasets can be very large
  * Standard architectures are emerging -- how to organize computations on this storage?
* Solutions
  * Use distributed storage
    * 6-24 disks attached to a blade, 32-64 blades in a rack connected by Ethernet
  * Push computations down to storage
* Stable storage becomes a first order problem. Answer: Distributed File System
  * Typical usage pattern
    * Huge files (100s of GB to TB)
    * Data is rarely updated in place
    * Reads and appends are common

### TODO

## CAP Theorem & Eventual Consistency

### CAP & Eventual Consistency

* Consistency models for distributed systems: ACID, BASE, Paxos
* CAP theorem (Eric Brewer, 2002; started as conjecture, proven in 2002?): You can have just two of Consistency, Availability, and Partition Tolerance
  * Consistency: All nodes see the same data at the same time
  * Availability: A guarantee that every request receives a response about whether it was successful or failed
  * Partition tolerance: The system continues to operate despite arbitrary message loss or failure of part of the system
* Data centers should weaken consistency for faster response

### ACID & BASE

* ACID
  * Atomicity: Even if transactions have multiple operations, does them to completion (commit) or rolls back so that they leave no effect (abort)
  * Consistency: A transaction that runs on a correct database leaves it in a correct/consistent state
  * Isolation: It looks as if each transaction ran all by itself. Basically says "we'll hide any concurrency"
  * Durability: Once a transaction commits, updates can't be lost or rolled back

### Zookeeper & Paxos

## Distributed Key-Value Store

## Scalable Databases

## Publish-Subscribe Queues (Kafka)


# Operating Systems Papers - Index

## Meta stuff

* Reading lists
  * [CS 736 @ UW-Madison: Advanced Operating Systems](/earlier-readings-and-notes/index/cs-736-uw-madison-fall-2020-reading-list)
  * [CS 262a @ Berkeley: Advanced Topics in Computer Systems](https://ucbrise.github.io/cs262a-fall2020/)
  * [OSTEP (Operating Systems: Three Easy Pieces)](http://pages.cs.wisc.edu/~remzi/OSTEP/) is written by the brilliant Remzi & Andrea and each chapter is followed by a lovely reading list about the topic covered in the chapter.
  * Some reading notes by individuals:
    * [Zeyuan Hu's paper reading notes](https://zhu45.org/), with a focus on database systems
    * [An open-source reading notes in CN](https://github.com/dyweb/papers-notebook), with a focus on virtualization and distributed systems

![Source: http://pages.cs.wisc.edu/\~remzi/OSTEP/](/files/-MROIqbSLWD5Q1pCzp-D)

* Some other stuff
  * [CS 262a @ Berkeley Fall 2020 class summary slides](https://ucbrise.github.io/cs262a-fall2020/notes/26-Class-Summary.pdf)
  * [Systems Benchmarking Crimes](https://www.cse.unsw.edu.au/~gernot/benchmarking-crimes.html)

## Table of Contents

### File and Storage Systems

| Title                                                                                                                                                                                                                                       | Venue                                    |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- |
| [FFS: A Fast File System for UNIX](/earlier-readings-and-notes/index/ffs-a-fast-file-system-for-unix)                                                                                                                                       | ACM Transactions on Computer Systems ‘84 |
| [NFS: Sun's Network File System](/earlier-readings-and-notes/index/nfs-suns-network-file-system)                                                                                                                                            | USENIX '86                               |
| [RAID: A Case for Redundant Arrays of Inexpensive Disks](/earlier-readings-and-notes/index/raid-a-case-for-redundant-arrays-of-inexpensive-disks)                                                                                           | SIGMOD ‘88                               |
| [LFS: The Design and Implementation of a Log-Structured File System](/earlier-readings-and-notes/index/lfs-the-design-and-implementation-of-a-log-structured-file-system)                                                                   | ACM Transactions on Computer Systems ‘92 |
| [SnapMirror: File-System-Based Asynchronous Mirroring for Disaster Recovery](/earlier-readings-and-notes/index/snapmirror-file-system-based-asynchronous-mirroring-for-disaster-recovery)                                                   | FAST '02                                 |
| [Venti: A New Approach to Archival Storage](/earlier-readings-and-notes/index/venti-a-new-approach-to-archival-storage)                                                                                                                     | FAST '02                                 |
| [ARC: A Self-Tuning, Low Overhead Replacement Cache](/earlier-readings-and-notes/index/arc-a-self-tuning-low-overhead-replacement-cache)                                                                                                    | FAST '03                                 |
| [RDP: Row-Diagonal Parity for Double Disk Failure Correction](/earlier-readings-and-notes/index/rdp-row-diagonal-parity-for-double-disk-failure-correction)                                                                                 | FAST '04                                 |
| [Data Domain: Avoiding the Disk Bottleneck in the Data Domain Deduplication File System](/earlier-readings-and-notes/index/data-domain-avoiding-the-disk-bottleneck-in-the-data-domain-deduplication-file-system)                           | FAST '08                                 |
| [Mnemosyne: Lightweight Persistent Memory](broken://pages/-MQduO-uYpAkF3YaRyBI)                                                                                                                                                             | ASPLOS '11                               |
| [A File is Not a File: Understanding the I/O Behavior of Apple Desktop Applications](/earlier-readings-and-notes/index/a-file-is-not-a-file-understanding-the-i-o-behavior-of-apple-desktop-applications)                                   | SOSP '11                                 |
| [OptFS: Optimistic Crash Consistency](/earlier-readings-and-notes/index/optfs-optimistic-crash-consistency)                                                                                                                                 | SOSP '13                                 |
| [All File Systems Are Not Created Equal: On the Complexity of Crafting Crash-Consistent Applications](/earlier-readings-and-notes/index/all-file-systems-are-not-created-equal-on-the-complexity-of-crafting-crash-consistent-applications) | OSDI '14                                 |
| [The Unwritten Contract of Solid State Drives](/earlier-readings-and-notes/index/the-unwritten-contract-of-solid-state-drives)                                                                                                              | EuroSys '17                              |
| [From WiscKey to Bourbon: A Learned Index for Log-Structured Merge Trees](/earlier-readings-and-notes/index/from-wisckey-to-bourbon-a-learned-index-for-log-structured-merge-trees)                                                         | OSDI '20                                 |
| [CheckFreq: Frequent, Fine-Grained DNN Checkpointing](/machine-learning-systems/machine-learning-systems-index/checkfreq-frequent-fine-grained-dnn-checkpointing)                                                                           | FAST '21                                 |

### Process Synchronization and Scalability ( 🥵 )

| Title | Venue |
| ----- | ----- |
|       |       |

### Scheduling

| Title                                                                                                                                                                                                                         | Venue          |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- |
| [Scheduler Activations: Effective Kernel Support for the User-Level Management of Parallelism](/earlier-readings-and-notes/index/scheduler-activations-effective-kernel-support-for-the-user-level-management-of-parallelism) | ACM SIGOPS '91 |
| [Lottery Scheduling: Flexible Proportional-Share Resource Management](/earlier-readings-and-notes/index/lottery-scheduling-flexible-proportional-share-resource-management)                                                   | OSDI '94       |
| [Resource Containers: A New Facility for Resource Management in Server Systems](/earlier-readings-and-notes/index/resource-containers-a-new-facility-for-resource-management-in-server-systems)                               | OSDI '99       |
| [The Linux Scheduler: A Decade of Wasted Cores](/earlier-readings-and-notes/index/the-linux-scheduler-a-decade-of-wasted-cores)                                                                                               | EuroSys '16    |
| [Monotasks: Architecting for Performance Clarity in Data Analytics Frameworks](/earlier-readings-and-notes/index/monotasks-architecting-for-performance-clarity-in-data-analytics-frameworks)                                 | SOSP '17       |
| [Gandiva: Introspective Cluster Scheduling for Deep Learning](/machine-learning-systems/machine-learning-systems-index/gandiva-introspective-cluster-scheduling-for-deep-learning)                                            | OSDI '18       |
| [Tiresias: A GPU Cluster Manager for Distributed Deep Learning](/machine-learning-systems/machine-learning-systems-index/tiresias-a-gpu-cluster-manager-for-distributed-deep-learning)                                        | NSDI '19       |
| [Themis: Fair and Efficient GPU Cluster Scheduling](/machine-learning-systems/machine-learning-systems-index/themis-fair-and-efficient-gpu-cluster-scheduling)                                                                | NSDI '20       |

### OS Structure and Virtual Machines

| Title                                                                                                                                                                                                     | Venue    |
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| [Disco: Running Commodity Operating Systems on Scalable Multiprocessors](/earlier-readings-and-notes/index/disco-running-commodity-operating-systems-on-scalable-multiprocessors)                         | SOSP '97 |
| [Memory Resource Management in VMWare ESX Server](/earlier-readings-and-notes/index/memory-resource-management-in-vmware-esx-server)                                                                      | OSDI '02 |
| [ReVirt: Enabling Intrusion Analysis through Virtual Machine Logging and Replay](/earlier-readings-and-notes/index/revirt-enabling-intrusion-analysis-through-virtual-machine-logging-and-replay)         | OSDI '02 |
| [Biscuit: The benefits and costs of writing a POSIX kernel in a high-level language](/earlier-readings-and-notes/index/biscuit-the-benefits-and-costs-of-writing-a-posix-kernel-in-a-high-level-language) | OSDI '18 |
| [LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation](/earlier-readings-and-notes/index/legoos-a-disseminated-distributed-os-for-hardware-resource-disaggregation)                | OSDI '18 |

## To Read

### To move from local note to GitBook

#### File and Storage Systems

* [ ] Mnemosyne: Lightweight Persistent Memory
* [ ] Level Hash: Write-Optimized and High-Performance Hashing Index Scheme for Persistent Memory

#### Process Synchronization and Scalability

* [ ] Monitors: An Operating System Structuring Concept
* [ ] Mesa: Experiences with Processes and Monitors in Mesa
* [ ] Scalability Analysis: An Analysis of Linux Scalability to Many Cores
* [ ] Scalable Commutativity: The Scalable Commutativity Rule: Designing Scalable Software for Multicore Processors
* [ ] (Delegation/RCL) Remote Core Locking: Migrating Critical-Section Execution to Improve the Performance of Multithreaded Applications
* [ ] Shuffle Locks: Scalable and Practical Locking with Shuffling
* [ ] Arachne: Core-Aware Thread Management

#### Scheduling

* [ ] SEDA: An Architecture for Well-Conditioned, Scalable Internet Services
* [ ] TAM: Principled Schedulability Analysis for Distributed Storage Systems using Thread Architecture Models

#### OS Structure and Virtual Machines

* [ ] THE: The Structure of "THE" Multiprogramming System
* [ ] Nucleus: The Nucleus of a Multiprogramming System
* [ ] Exokernel: An Operating System Architecture for Application-Level Resource Management
* [ ] Arrakis: The Operating System is the Control Plane
* [ ] UNIX: The UNIX Time-Sharing System

### To read

* [ ] seL4: Formal Verification of an OS Kernel


# CS 736 @ UW-Madison Fall 2020 Reading List

Imported from https\://canvas.wisc.edu/courses/205576/pages/paper-list. The reading list was put together by Prof. Andrea Arpaci-Dusseau.

This semester, we are reading many of the paper that the OS community has placed into the SIGOPS [Hall of Fame (Links to an external site.)](https://www.sigops.org/award-hof.html). The SIGOPS Hall of Fame Award was instituted in 2005 to recognize the most influential Operating Systems papers that were published at least ten years in the past.   We've marked those papers on our reading list that are in the Hall of Fame.

[Schedule](https://canvas.wisc.edu/courses/205576/pages/schedule)

## File and Storage Systems&#x20;

1. **Background: Traditional Local File Systems -- FFS and LFS**
   1. &#x20;**FFS -** [**Questions,**](https://canvas.wisc.edu/courses/205576/pages/ffs-questions) **Background:** [**Disk Questions**](https://canvas.wisc.edu/courses/205576/pages/disk-questions)\
      McKusick, M.K., Joy, W\.N., Leffler, S.J., and Fabry, R.S. , [**A Fast File System for UNIXLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/ffs.ps) **,** ACM Transactions on Computer Systems, Vol. 2, No. 3, August 1984, pp. 181-197.  SIGOPS Hall of Fame Award&#x20;
   2. &#x20;**LFS -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20LFS?titleize=0)\
      Rosenblum, M. and Ousterhout, J.  [**The Design and Implementation of a Log-Structured File SystemLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/lfs.ps) **,** ACM Transactions on Computer Systems, Vol. 10, No. 1, February 1992, pp. 26-52.  SIGOPS Hall of Fame Award
2. &#x20;**Background: Storage Technology -- RAID**&#x20;
   1. &#x20;**RAID** [**- Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20RAID?titleize=0)\
      Patterson, D., Gibson, G., and Katz, R., [**A Case for Redundant Arrays of Inexpensive Disks (RAID)Links to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/raid.ps) Proceedings of the 1988 ACM SIGMOD Conference on Management of Data, Chicago IL, June 1988.  SIGOPS Hall of Fame Award
   2. &#x20;**RDP (No questions yet)**\
      [**Row-Diagonal Parity for Double Disk Failure Correction (Links to an external site.)**](https://www.usenix.org/conference/fast-04/row-diagonal-parity-double-disk-failure-correction)[**, (Links to an external site.)**](https://www.usenix.org/conference/fast-04/row-diagonal-parity-double-disk-failure-correction) Proceedings of USENIX File and Storage Technology (FAST), 2004, FAST Test of Time Award
3. &#x20;**Measurement**
   1. &#x20;**iBench** [**- Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20iBench?titleize=0)\
      *Tyler Harter, Chris Dragga, Michael Vaughn, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau* [ **A file is not a file: understanding the I/O behavior of Apple desktop applications  Links to an external site.**](http://research.cs.wisc.edu/wind/Publications/ibench-sosp11.pdf)SOSP '11 Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles Pages 71-83 SOSP Best Paper, UW-Madison Authors
4. &#x20;**Archival Storage and Deduplication-**[**Questions**](https://canvas.wisc.edu/courses/205576/pages/questions-archival-storage)
   1. &#x20;**SnapMirror** [\
      **SnapMirror: File-System-Based Asynchronous Mirroring for Disaster Recovery,**](broken://pages/-MNUAYOat_8e1WeejSh9)2002\
      FAST Test of Time Award
   2. &#x20;**Venti**[\
      **Venti: A New Approach to Archival Storage**,](broken://pages/-MNUAYOat_8e1WeejSh9) 2002\
      FAST Test of Time Award
   3. &#x20;**Deduplication**[\
      **Avoiding the Disk Bottleneck in the Data Domain Deduplication File System,**](broken://pages/-MNUAYOat_8e1WeejSh9)2008\
      FAST Test of Time Award
5. &#x20;**Caching**
   1. **ARC (No questions yet)**[\
      **ARC: A Self-Tuning, Low Overhead Replacement Cache,**](broken://pages/-MNUAYOat_8e1WeejSh9) 2003 FAST Test of Time Award
6. &#x20;**Crash Consistency**
   1. &#x20;**Alice -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Alice?titleize=0)\
      Thanumalayan Sankaranarayana Pillai, Vijay Chidambaram, Ramnatthan Alagappan, Samer Al-Kiswany, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, [**All File Systems Are Not Created Equal: On the Complexity of Crafting Crash-Consistent Applications** Links to an external site.](http://research.cs.wisc.edu/adsl/Publications/alice-osdi14.pdf)Proceedings of the 11th Symposium on Operating Systems Design and Implementation (OSDI '14) Broomfield, CO, October 2014. UW-Madison Authors
   2. &#x20;**OptFS -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20OptFS?titleize=0)\
      Vijay Chidambaram, Thanumalayan Sanakaranarayana Pillai, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau [**Optimistic Crash Consistency Links to an external site.**](http://research.cs.wisc.edu/adsl/Publications/optfs-sosp13.pdf)Symposium on Operating System Principles, SOSP 2013 , UW-Madison Authors
7. &#x20;**SSDs  and  Key-Value  Stores**
   1. &#x20;[**Unwritten SSD Contract**\
      *Jun He*Links to an external site.](http://www.cs.wisc.edu/~jhe/)*,* [*Sudarsun KannanLinks to an external site.*](http://www.cs.wisc.edu/~sudarsun/)*,* [*Andrea C. Arpaci-DusseauLinks to an external site.*](http://www.cs.wisc.edu/~dusseau/)*,* [*Remzi H. Arpaci-DusseauLinks to an external site.*](http://www.cs.wisc.edu/~remzi/)  [**The Unwritten Contract of Solid State Drives** Links to an external site.](http://research.cs.wisc.edu/adsl/Publications/eurosys17-he.pdf)Proceedings of the 20th European Conference on Computer Systems (EuroSys '17) Belgrade, Serbia, April 2017. UW-Madison Authors
   2. &#x20;**Bourbon (preprint in hotcrp)**\
      *Yifan Dai, Yien Xu, Aishwarya Ganesan, Ramnatthan Alagappan, Brian Kroth, Andrea Arpaci-Dusseau, and Remzi Arpaci-Dusseau*. From WiscKey to Bourbon: A Learned Index for Log-Structured Merge Trees. *In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI’20), October 2020.*  UW-Madison Authors
      1. **Optional Background: WiscKey -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/questions-wisckey)[\
         *Lanyue Lu*Links to an external site.](http://www.cs.wisc.edu/~ll/)*,* [*Thanumalayan Sankaranarayana PillaiLinks to an external site.*](http://www.cs.wisc.edu/~madthanu/)*,* [*Andrea C. Arpaci-DusseauLinks to an external site.*](http://www.cs.wisc.edu/~dusseau/)*,*[*Remzi H. Arpaci-DusseauLinks to an external site.*](http://www.cs.wisc.edu/~remzi/)  [**WiscKey: Separating Keys from Values in SSD-conscious Storage** Links to an external site.](http://research.cs.wisc.edu/adsl/Publications/wisckey-fast16.pdf)Proceedings of the 14th USENIX Conference on File and Storage Technologies (FAST '16) UW-Madison Authors
8. &#x20;**Persistent Memory**
   1. &#x20;**Mnemosyne**\
      *Haris Volos, Andres Jaan Tack, Michael M. Swift.* [*Mnemosyne: Lightweight Persistent MemoryLinks to an external site.*](https://pages.cs.wisc.edu/~swift/papers/asplos11_mnemosyne.pdf)*, ASPLOS '11: Proceedings of the 16th International Conference on Architectural, UW-Madison Authors*
   2. &#x20;**Level Hash -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/index-for-persistent-memory)\
      *Pengfei Zuo, Yu Hua, and Jie Wu*, [**Write-Optimized and High-Performance Hashing Index Scheme for Persistent Memory  (Links to an external site.)**](https://www.usenix.org/conference/osdi18/presentation/zuo)*Huazhong University of Science and Technology, OSDI'18*
9. &#x20;Graph Processing - Don't read
   1. [**Links to an external site.**](http://www.cs.wisc.edu/~jhe/)**GraphChi -** [**Questions for all 3 papers**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Graph?titleize=0)\
      *Aapo Kyrola and Guy Blelloch and Carlos Guestrin* [**GraphChi: Large-Scale Graph Computation on Just a PC.**  (Links to an external site.)](https://www.usenix.org/system/files/conference/osdi12/osdi12-final-126.pdf)USENIX Symposium on Operating Systems Design and Implementation (OSDI'12).
   2. **Xstream**\
      *Amitabha Roy, Ivo Mihailovic, Willy Zwaenepoel* [**Xstream: Edge-centric graph processing using streaming partitions.   (Links to an external site.)**](https://infoscience.epfl.ch/record/188535/files/paper.pdf)Symposium on Operating Systems Principles (2013).
   3. &#x20;**FlashGraph**\
      *Da Zheng and Disa Mhembere and Randal Burns and Joshua Vogelstein and Carey E. Priebe and Alexander S. Szalay,* [**FlashGraph: Processing Billion-Node Graphs on an Array of Commodity SSDs,  (Links to an external site.)**](https://www.usenix.org/system/files/conference/fast15/fast15-paper-zheng.pdf)Conference on File and Storage Technologies (FAST 2015)

## Process Synchronization and Scalability

1. &#x20;**Background: Monitors, Theory and Practice-** [**Questions: Monitors**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Monitors?titleize=0)&#x20;
   1. &#x20;**Monitors**\
      C.A.R. Hoare  [**Monitors: An Operating System Structuring Concept Links to an external site.**](http://pages.cs.wisc.edu/~remzi/Classes/736/Fall2010/Papers/hoare-monitors.pdf)Communications of the ACM 17, 10, October 1974, pp. 549-557 &#x20;
   2. &#x20;**Mesa**\
      Butler W. Lampson, David D. Redell [**Experiences with Processes and Monitors in Mesa Links to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/mesa.ps)Communications of the ACM, 23 2, February 1980, pp. 105-117.  SIGOPS Hall of Fame Award
2. &#x20;**OS Scalability: Measurement and Redesign**
   1. [  (Links to an external site.)](http://doi.acm.org/10.1145/1368506.1368525)**Measurement -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Scalability%20Measurement?titleize=0)\
      Silas Boyd-Wickizer, Austin T. Clements, Yandong Mao, Aleksey Pesterev, M. Frans Kaashoek, Robert Morris, and Nickolai Zeldovich  [**An Analysis of Linux Scalability to Many Cores (Links to an external site.)**](http://people.csail.mit.edu/nickolai/papers/boyd-wickizer-scaling.pdf) In Proceedings of the 9th Symposium on Operating Systems Design and Implementation (OSDI), Vancouver, Canada, October 2010
   2. &#x20;**Scalable Commutativity -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Commutativity?titleize=0)\
      Austin T. Clements, M. Frans Kaashoek, Nickolai Zeldovich, Robert T. Morris, and Eddie Kohler [**The Scalable Commutativity Rule: Designing Scalable Software for Multicore Processors. (Links to an external site.)**](http://people.csail.mit.edu/nickolai/papers/clements-sc.pdf) In Proceedings of the 24th ACM Symposium on Operating Systems Principles (SOSP), Farmington, PA, November 2013.
3. &#x20;**Alternate Locking Primitives**
   1. &#x20;**Delegation -** [**Questions: Delegation**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Delegation?titleize=0)\
      Jean-Pierre Lozi and Florian David and Gael Thomas and Julia Lawall and Gilles Muller, [**Remote Core Locking: Migrating Critical-Section Execution to Improve the Performance of Multithreaded** **Applications,**  (Links to an external site.)](https://www.usenix.org/system/files/conference/atc12/atc12-final237.pdf)USENIX Annual Technical Conference (ATC'12), 2012. &#x20;
   2. &#x20;**Shuffle Locks**\
      Sanidhya Kashyap,  Irina Calciu, Xiaohe Cheng, Changwoo Min, Taesoo Kim, [**Scalable and Practical Locking with Shufflin** (Links to an external site.)](https://dl.acm.org/doi/10.1145/3341301.3359629)[**g,** ](broken://pages/-MNUAYOat_8e1WeejSh9) SOSP'19

## Scheduling

1. &#x20;**Background: Threads and Events**
   1. &#x20;**Scheduler Activations -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Scheduler%20Activations?titleize=0)\
      Anderson, T., Bershad, B., Lazowska, E., and Levy, H. [**Scheduler Activations: Effective Kernel Support for the User-Level Management of ParallelismLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/scheduler.pdf) ACM Transactions on Computer Systems, Vol. 10, No. 1, February 1992, pp. 53-79.&#x20;
   2. &#x20;**SEDA**\
      Matt Welsh, David Culler, Eric Brewer (UC Berkeley) [**SEDA: An Architecture for Well-Conditioned, Scalable Internet Services (Links to an external site.)**](http://www.sosp.org/2001/papers/welsh.pdf) SOSP'01
2. &#x20;**Background: Local CPU Schedulers and Resource Tracking**
   1. &#x20;**Lottery Scheduling -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Lottery?titleize=0)\
      Waldspurger, C.A. and Weihl, W\.E. [**Lottery Scheduling: Flexible Proportional-Share Resource Mangement Links to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/lottery-osdi94.ps)Proceedings of the First Symposium on Operating Systems Design and Implementation, Monterey CA, November 1994, pp. 1-11.
   2. &#x20;**Resource Containers -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Resource%20Containers?titleize=0)\
      Banga, G., Druschel, P,. Mogul, J. [**Resource Containers: A New Facility for Resource Management in Server SystemsLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/rc-osdi99.ps.gz) Proceedings of the Third Symposium on Operating System Design and Implementation (OSDI-III), New Orleans, LA, February, 1999, 45-58. &#x20;
3. &#x20;**Measurement: Linux and System Services**
   1. &#x20;**Linux Scheduler**\
      *Jean-Pierre Lozi (Université de Nice Sophia-Antipolis), Baptiste Lepers (Ecole Polytechnique Fédérale de Lausanne), Justin Funston (University of British Columbia), Fabien Gaud (Coho Data), Vivien Quéma (Grenoble INP / ENSIMAG), Alexandra Fedorova* [**The Linux Scheduler: A Decade of Wasted Cores.**  (Links to an external site.)](https://dl.acm.org/citation.cfm?doid=2901318.2901326)*Eurosys 2016*
   2. &#x20;**TAM -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20TAM?titleize=0)\
      Suli Yang, Jing Liu, Andrea C. Arpaci-Dusseau, and Remzi H. Arpaci-Dusseau [**Principled Schedulability Analysis for Distributed Storage Systems using Thread Architecture Models**  (Links to an external site.)](https://www.usenix.org/conference/osdi18/presentation/yang)*(OSDI'18) UW-Madison Authors*
      1. &#x20;*Optional Background*\
         **Split-Level I/O Scheduling**[\
         Links to an external site.](http://www.cs.wisc.edu/~suli/)*Suli Yang, Tyler Harter, Nishant Agrawal, Salini Selvaraj Kowsalya, Anand Krishnamurthy, Samer Al-Kiswany, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau*  [**Split-Level I/O Scheduling** Links to an external site.](http://research.cs.wisc.edu/adsl/Publications/split-sosp15.pdf)Proceedings of the 25th ACM Symposium on Operating Systems Principles (SOSP '15) UW-Madison Authors
4. &#x20;**Current System Scheduling**
   1. &#x20;***Monotasks -*** [***Questions***](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Monotasks?titleize=0)\
      *Kay Ousterhout (UC Berkeley); Christopher Canel (Carnegie Mellon University); Sylvia Ratnasamy (UC Berkeley); Scott Shenker* [***Monotasks: Architecting for Performance Clarity in Data Analytics Frameworks (Links to an external site.)***](https://dl.acm.org/authorize?N47256)***,** SOSP'17*
   2. &#x20;**Arachne -** [**Questions: Arachne**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Arachne?titleize=0)\
      *Henry Qin, Qian Li, Jacqueline Speiser, Peter Kraft, and John Ousterhout,Stanford University,*[  (Links to an external site.)](https://www.usenix.org/conference/osdi18/presentation/qin)[**Arachne: Core-Aware Thread Management  (Links to an external site.)**](https://www.usenix.org/conference/osdi18/presentation/qin)[*OSDI'18* (Links to an external site.)](https://www.usenix.org/conference/osdi18/presentation/qin)
5. &#x20;***Current System Scheduling 2***
   1. &#x20;**Themis**[ **(Links to an external site.)**](http://shivaram.org/publications/themis-nsdi2020.pdf)[\
      &#x20;*(Links to an external site.)*](http://shivaram.org/publications/themis-nsdi2020.pdf)*Kshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman, and Aditya Akella, University of Wisconsin-Madison;*&#x41;mar Phanishayee,*Microsoft Research;*&#x53;huchi Chawla,*University of Wisconsin-Madison,* [Themis: Fair and Efficient GPU Cluster Scheduling (Links to an external site.)](http://shivaram.org/publications/themis-nsdi2020.pdf)- NSDI 2020, UW-Madison Authors

## OS Structure and Virtual Machines

1. &#x20;**Background: Layered vs. Extensible Kernels**
   1. &#x20;**THE -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20THE?titleize=0)\
      Edsger W. Dijkstra[**The Structure of the "THE" Multiprogramming SystemLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/theTHE.pdf) Communications of the ACM 11(5), May 1968.  SIGOPS Hall of Fame Award
   2. &#x20;**Nucleus -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Nucleus?titleize=0)\
      Per Brinch Hansen, [**The Nucleus of a Multiprogramming SystemLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/nucleus.pdf) Communications of the ACM 13(4), April 1970 &#x20;
2. &#x20;**Microkernels: Concepts and Measurements**
   1. &#x20;**Exokernel -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Exokernel?titleize=0)\
      Dawson R. Engler, M. Frans Kaashoek, and James O’Toole Jr [**Exokernel: An Operating System Architecture for Application-Level Resource Management (Links to an external site.)**](https://pdos.csail.mit.edu/6.828/2008/readings/engler95exokernel.pdf) SOSP '95 Proceedings of the fifteenth ACM symposium on Operating systems principles
   2. &#x20;**Arrakis**\
      Simon Peter, Jialin Li, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, and Thomas Anderson, *University of Washington;* Timothy Roscoe, *ETH Zürich* [Arrakis: The Operating System is the Control Plane, (Links to an external site.)](https://www.usenix.org/system/files/conference/osdi14/osdi14-paper-peter_simon.pdf) OSDI'14
      1. Optional Background - **Barrelfish**\
         Andrew Baumann, Paul Barham, Pierre-Evariste Dagand, Tim Harris, Rebecca Isaacs, Simon Peter, Timothy Roscoe, Adrian Schüpbach, and Akhilesh Singhania. [**The Multikernel: A new OS architecture for scalable multicore systems** (Links to an external site.)](http://www.barrelfish.org/publications/barrelfish_sosp09.pdf)[. (Links to an external site.)](http://www.barrelfish.org/publications/barrelfish_sosp09.pdf) In *Proceedings of the 22nd ACM Symposium on OS Principles*, Big Sky, MT, USA, October 2009
3. &#x20;**Monolithic, Disaggregation, and HLLs**
   1. &#x20;**UNIX**\
      Ritchie, D.M. and Thompson, K. [**The UNIX Time-Sharing SystemLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/unix-cacm.ps.gz) Communications of the ACM, Vol. 17, No. 7, July 1974, pp. 365-375.  SIGOPS Hall of Fame Award
   2. &#x20;**Disaggregation**\
      *Yizhou Shan, Yutong Huang, Yilun Chen, and Yiying Zhang,* [**LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation,** (Links to an external site.)](https://www.usenix.org/conference/osdi18/presentation/shan)OSDI 2018
   3. &#x20;**HLLs -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Biscuit?titleize=0)\
      *Cody Cutler, M. Frans Kaashoek, and Robert T. Morris, MIT CSAIL,* [**The benefits and costs of writing a POSIX kernel in a high-level language,** (Links to an external site.)](https://www.usenix.org/conference/osdi18/presentation/cutler)OSDI 2018
4. &#x20;**Virtual Machines**
   1. &#x20;**Disco -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20Disco?titleize=0)\
      Edouard Bugnion, Scott Devine, Mendel Rosenblum. [**Disco: Running Commodity Operating Systems on Scalable MultiprocessorsLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/disco.ps.gz) Proceedings of The Sixteenth Symposium on Operating Systems Principles (October 1997).  SIGOPS Hall of Fame Award
   2. [  (Links to an external site.)](http://doi.acm.org/10.1145/1368506.1368525)**ESX -** [**Questions**](https://canvas.wisc.edu/courses/205576/pages/Questions%3A%20ESX?titleize=0)\
      Carl A. Waldspurger  [**Memory Resource Management in VMware ESX Server  Links to an external site.**](http://pages.cs.wisc.edu/~remzi/Classes/736/Fall2010/Papers/esx-osdi02.pdf)In Proc. Fifth Symposium on Operating Systems Design and Implementation (OSDI ’02), Dec. 2002  SIGOPS Hall of Fame Award
   3. &#x20;**Revirt**\
      George W. Dunlap, Samuel T. King, Sukru Cinar, Murtaza A. Basrai, and Peter M. Chen.\
      [**ReVirt: Enabling intrusion analysis through virtual-machine logging and replay (Links to an external site.)**](http://doi.acm.org/10.1145/844128.844148)**.**\
      In Proceedings of the 5th Symposium on Operating Systems Design and Implementation (OSDI '02), 2002, 211-224. SIGOPS Hall of Fame Award
      1. &#x20;**Optional Overview**\
         Bugnion, Nief, Tsafir, [**Hardware and Software Support for Virtualization**](https://canvas.wisc.edu/courses/205576/files/13611425/download?wrap=1)  Synthesis Lectures on Computer Architecture

## Testing, Debugging, and Design

1. &#x20;**Profiling and Binary Code**
   1. &#x20;**KernInst**\
      *Ariel Tamches and Barton P. Miller,* "Fine-Grained Dynamic Instrumentation of Commodity Operating System Kernels",3rd Symposium on Operating Systems Design and Implementation (OSDI),New Orleans, Louisiana, February 1999. UW-Madison Authors
      1. &#x20;**Optional**
         1. *Nathan E. Rosenblum, Xiaojin (Jerry) Zhu and Barton P. Miller,* "Who Wrote This Code? Identifying the Authors of Program Binaries", 2011 European Symposium on Research in Computer Security (ESORICS), Leuven, Belgium, September 2011. UW-Madison Authors
         2. &#x20;[*Xiaozhu Meng (Links to an external site.)*](https://www.researchgate.net/profile/Xiaozhu_Meng) *and Barton P. Miller,* Binary Code is Not Easy, International Symposium on Software Testing and Analysis, 2016 UW-Madison Authors
2. &#x20;**Symbolic Execution and Debugging Experience**
   1. &#x20;**KLEE**\
      *Cristian Cadar, Daniel Dunbar, and Dawson Engler.* **KLEE: Unassisted and Automatic Generation of High-Coverage Tests for Complex Systems Programs.** In OSDI’08, SIGOPS Hall of Fame Award
   2. **Debug**\
      Kirk Glerum, Kinshuman Kinshumann, Steve Greenberg, Gabriel Aul, Vince Orgovan, Greg Nichols, David Grant, Gretchen Loihle, and Galen Hunt.[**Debugging in the (Very) Large: Ten Years of Implementation and Experience**](broken://pages/-MNUAYOat_8e1WeejSh9)**. In SOSP ’09,** SIGOPS Hall of Fame Award
3. &#x20;**Summary of System Design**
   1. **Hints**\
      Butler Lampson [Hints for Computer System Design (Links to an external site.)](http://portal.acm.org/citation.cfm?doid=800217.806614),  Proceedings of the Ninth ACM Symposium on Operating Systems Principles, pp. 33-48, October 1983, Bretton Woods, NH, USA.  SIGOPS Hall of Fame Award

## Other Relevant SIGOPS Hall of Fame Papers (not covered)

1. Daniel G. Bobrow, Jerry D. Burchfiel, Daniel L. Murphy and Raymond S. Tomlinson.\
   [Tenex, A Paged Time Sharing System for the PDP-10 (Links to an external site.)](http://dl.acm.org/citation.cfm?id=361271) \
   Communications of the ACM 15(3), March 1972. [SIGOPS Hall of Fame Award (Links to an external site.)](http://doi.acm.org/10.1145/1368506.1368525)
2. Daley, R.C., and Dennis, J.B. \
   [**Virtual Memory, Processes, and Sharing in MULTICSLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/multics.pdf) \
   Communications of the ACM, Vol. 11, No. 5, May 1968, pp. 306-312. (Multics paper in [SIGOPS Hall of Fame (Links to an external site.)](http://doi.acm.org/10.1145/1368506.1368525))
3. &#x20;R. Rashid and A. Tevanian and M. Young and D. Golub and R. Baron and D. Black and W. Bolosky and J. Chew, \
   [**Machine-Independent Virtual Memory Management for Paged Uniprocessor and Multiprocessor ArchitecturesLinks to an external site.**](http://pages.cs.wisc.edu/~dusseau/Classes/CS736/Papers/mach-vm.pdf) [**SIGOPS Hall of Fame Award (Links to an external site.)**](http://doi.acm.org/10.1145/1368506.1368525)\
   Proceedings of the 2nd International Conference on Architectural Support for Programming Languages and Operating System (ASPLOS), 1987.  (Mach in [SIGOPS Hall of Fame (Links to an external site.)](http://doi.acm.org/10.1145/1368506.1368525))
4. J. Liedtke. \
   [On micro-kernel construction (Links to an external site.)](http://doi.acm.org/10.1145/224056.224075).\
   In Proceedings of the 15th ACM symposium on Operating Systems Principles (SOSP '95), December 1995, 237-250. [SIGOPS Hall of Fame Award (Links to an external site.)](http://doi.acm.org/10.1145/1368506.1368525)


# All File Systems Are Not Created Equal: On the Complexity of Crafting Crash-Consistent Applications

## One-line Summary

This paper presents a comprehensive study of file system persistence properties and modern application crash inconsistency vulnerabilities. Two tools, BOB and ALICE, are presented to analyze FS-level and application-level vulnerabilities.

## Paper Structure Outline

1. Introduction
2. Persistence Properties
   1. An Example
   2. Study and Results
      1. Atomicity
      2. Ordering
   3. Summary
3. The Application-Level Intelligent Crash Explorer (ALICE)
   1. Usage
   2. Crash States and APMs
      1. Logical Operations
      2. Abstract Persistence Models
      3. Constructing crash states.
   3. Finding Application Requirements
   4. Static Vulnerabilities
   5. Implementation
   6. Limitations
4. Application Vulnerabilities
   1. Workloads and Checkers
   2. Overview
      1. Databases and Key-Value Stores
      2. Version Control Systems
      3. Virtualization and Distributed Systems
   3. Vulnerabilities Found
   4. Common Patterns
      1. Atomicity across System Calls
      2. Atomicity within System Calls
      3. Ordering between System Calls
      4. Durability
      5. Summary
   5. Impact on Current File Systems
   6. Evaluating New File-System Designs
   7. Discussion
5. Related Work
6. Conclusion

## Background & Motivation

To provide crash consistency for update-in-place file systems, journaling is performed. A high-level overview of journaling:

* Intuition
  * Before updating the file system, write a note describing the update
  * Make sure note is safely on disk
  * Once the note is safe, update the file system
  * If interrupted, read the note and redo updates
* Protocol
  * Write the data (no pointers to it) - Optional
  * Write the note: Journal Metadata
  * Make sure the note is durably written: Journal Commit
  * Update the in-place metadata: Checkpointing
  * Replay the note: Recovery

The motivation for this work is that applications may not be aware of consistency guarantees for different file systems or even the same file system with different configurations. The authors categorize file system persistency properties and study the differences among widely deployed file systems (ext3, ext4, btrfs). It then studies application-level crash inconsistency vulnerabilities.

## Design and Implementation

### BOB (Block Order Breaker)

BOB analyzes syscall persistence properties on a **file system**. Here's how BOB works:

1. Runs user-level workloads stressing the property
2. Records block-level trace of the workload
3. Reconstructs disk-states possible on a power-loss
   1. All states possible if disk-cache does not re-order
   2. A few states where disk-cache re-orders
4. Run FS recovery, verify property on each disk-state (atomicity & ordering)

### ALICE (Application-Level Intelligent Crash Explorer)

![](/files/-MPpdshAVord0XpYbhNF)

ALICE analyzes **application** update protocols and finds crash vulnerabilities (across all file systems). ALICE runs an application and collects its syscall trace (which represents an execution of the application's update protocol). The traces are then converted into a sequence of logical operations, which is then used to generates possible disk states according to characteristics of these syscalls and produces all possible intermediate disk states. If any of these disk states violate any application invariant, then this is considered a crash vulnerability.

## Evaluation

### File systems

![File systems vary in their persistence properties](/files/-MPpdphN7aP-kwg0SPKd)

The conclusion for the file system study is that applications should not rely on persistence properties. Also, testing applications on a specific file system is not enough.

### Applications

![](/files/-MPpgLGocImV69v54-LZ)

![](/files/-MPpgwxjEANItXbKOR2H)

![A total of 60 vulnerabilities are found.](/files/-MPpgSx3lmcxwy3l0fh4)

Applications from different domains are studied (relational & non-relational databases, version control, distributed services, virtualization). Many of the vulnerabilities will bring trouble over some modern file system configurations (e.g., content-atomic appends over no-delayed-allocation file systems).

## New Vocabulary

* APM: Abstract Persistence Model

## Links

* [Paper PDF](https://www.usenix.org/system/files/conference/osdi14/osdi14-paper-pillai.pdf)
* [Presentation Video at OSDI '14](https://www.usenix.org/conference/osdi14/technical-sessions/presentation/pillai)
* [Presentation Slides at OSDI '14](https://www.usenix.org/sites/default/files/conference/protected-files/osdi14_slides_pillai.pdf)


# ARC: A Self-Tuning, Low Overhead Replacement Cache

## One-line Summary

ARC is a new cache management policy.

## Paper Structure Outline

1. Introduction
   1. The Problem
   2. Our Contributions
   3. A Brief Outline of the Paper
2. Prior Work: A Brief Review
   1. Offline Optimal
   2. Recency
   3. Frequency
   4. Recency and Frequency
   5. Temporal Distance Distribution
   6. Caching using Multiple Experts
   7. Ghost Caches
   8. Summary
3. A Class of Replacement Policies
   1. Double Cache and a Replacement Policy
   2. A New Class of Policies
   3. LRU
4. Adaptive Replacement Cache
   1. Fixed Replacement Cache
   2. The Policy
   3. Learning
   4. Scan-Resistant
   5. Extra History
5. Experimental Results
   1. Traces
   2. OLTP
   3. Two Traces: P8 and P12
   4. ARC and 2Q
   5. ARC and MQ
   6. ARC and LRU
   7. ARC is Self-Tuning and Empirically Universal
   8. A Closer Examination of Adaptation in ARC
6. Conclusions

## Background & Motivation

The goals/metrics of this work are:

* High hit rate
* Low overhead
  * computational & space
  * locking overhead (high concurrency)
* Scan resistant: A large file does not blow out the cache
* No parameters to tune (self-tuning)
* Online algorithm: Handles any workload; Changes dynamically
* Balance between recency & frequency
* Empirically universal: Performs as well as fixed replacement policies

Some assumptions are:

* The replacement policy (selecting the page to page out) is investigated
* No prefetching (assume all demand paging)
* Only look at read policy (no write)
* The cache receives a continuous stream of requests for pages

## Existing Algorithms

![](/files/-MPlhJ8N3SAfhKiIhopl)

### Offline Optimal

* Replaces the furthest page in the future
* Upper bound on the achievable hit ratio by any online policy

### LRU

* Replaces the least recently used page
* Not scan resistant
* High concurrency (lock overhead)

### LFU

* Replaces the least frequently used page
* High implementation complexity
* Does not adapt to changes in access patterns (pages with previous high frequency counts may no longer be useful)

### LRU-k (LRU-2 in paper, most common)

* Track time of last 2 accesses of each page
* Replaces the page with the least recent penultimate reference
* Expensive algorithm (for maintaining a priority queue)

### [2Q](http://www.vldb.org/conf/1994/P439.PDF) (VLDB '94, used in PostgreSQL)

* Keeps three working sets: Current working set, previous working set, and the long term working set.
* Scan resistant
* Easy to implement
* Takes into account both recency and frequency
* 2Q has some sizing parameters (K\_in and K\_out)

### MQ (Multiple queues, USENIX '01)

![Course notes by Prof. Andrea Arpaci-Dusseau](/files/-MPlrJY9F6Bs-_4bpWEW)

* Focused on a replacement algorithm for buffer caches (second level on the storage system, below the traditional OS buffer cache)&#x20;
* Put items in m LRU queues according to their frequency. Q\_i contains pages that have been seen at least 2^i times but no more than 2^(i+1)-1 times recently
* An expiration time is associated with every item
* Maintains Q\_out: Ghost cache, contains references instead of actual data
* Not robust under a wider range of workloads
* Higher overhead than LRU, ARC, 2Q (need to check time stamps of pages on every request)

## Design and Implementation

![](/files/-MPlseDGK5MEutZ6IktG)

ARC is:

* Parameter-less
* Self-tuning
* Simple to implement
* Scan resistant
* Considers both recency and frequency

Two LRU lists are maintained: L1 contains pages accessed once recently (recency), partitioned into a top portion T1 and a bottom portion B1. Only T1 is in cache. L2 contains pages accessed at least twice recently (frequency), partitioned into a top portion T2 and a bottom portion B2. Only T2 is in cache. The middle line between the two lists can be shifted.

* Hit in T1 or T2: MRU to T2
* Miss in B1: MRU to T2, increase p, move T1
* Miss in B2: MRU of T2, decrease p
* Miss everywhere (not in B1/B2): MRU in T1, replace some LRU (complicated)

![](/files/-MPlxU723gUMSrsFsc__)

## New Vocabulary

* Demand paging: A disk page is copied into physical memory only if an attempt is made to access it and that page is not already in memory.

## Links

* [Paper PDF](https://www.usenix.org/legacy/events/fast03/tech/full_papers/megiddo/megiddo.pdf)
* [Presentation slides by the authors @ University of Houston](http://www2.cs.uh.edu/~paris/6360/PowerPoint/ARC.ppt)
* Thanks to Jane Chen for the paper review notes!


# A File is Not a File: Understanding the I/O Behavior of Apple Desktop Applications

## One-line Summary

The I/O patterns for home-user applications are studied.

## Paper Structure Outline

1. Introduction
2. Case Study
3. iBench Task Suite
   1. Representative
   2. Easy to Use
4. Analysis of iBench Tasks
   1. Nature of Files
      1. File Types
      2. File Sizes
   2. Access Patterns
      1. File Accesses
      2. Sequentiality
      3. Preallocation
   3. Transactional Properties
      1. Durability
      2. Atomic Writes
   4. Threads and Asynchronicity
5. Related Work
6. Discussion and Conclusions

## Background & Motivation

Previously, the design and implementation of file and storage systems have been inspired by workload studies. Nowadays, most file system users use home-user applications, thus an application study of typical I/O behavior of modern home-user applications is necessary.

## Design and Implementation

The authors presented a collection of applications, the iBench task suite. The applications consist of two Apple software suites: iWork (Pages, Numbers, Keynote) and iLife (iPhoto, iTunes, iMovie). 34 typical tasks (importing songs, editing movies, etc.) are analyzed.

An instrumentation framework built upon the DTrace tracing system is presented. DTrace monitors system calls made by each traced application and allows people to examine stack traces, in-kernel functions such as page-ins and page-outs, and other details required to ensure accuracy and completeness.

![](/files/-MPaJ-Ko8RER5dV-rv7K)

## Evaluation

The following main conclusions are drawn:

* **A file is not a file**: The files that appear to users are in actuality small file systems containing many sub-files. For example, a Microsoft .doc file is actually a FAT file system containing pieces of the document.
* **Sequential access is not sequential** (Fig. 8 & 9): Pure sequential access is rare.
* Auxiliary files dominate: Most files are helper files that applications use to provide a rich GUI, support multiple languages, and record history and metadata.
* **Writes are often forced** (Fig. 11 & 12): Most written data is explicitly forced to disk by the application. The implication is that frequent fsyncs reduces the benefits of buffering writes in memory.
* **Renaming is popular** (Fig. 13): Atomic operations, in particular `rename()`, is often used to present a consistent view of files to users. The implication is that traditional file locality does not work: placing a file on disk based on its parent directory does not work as expected when the file is first created in a temporary location and then renamed.
* **Multiple threads perform I/O** (Fig. 1): Threads are required to perform long-latency operations in the background to keep the GUI responsive.
* **Frameworks influence I/O**: Modern libraries put more code between applications and the underlying file system, as compared with traditional UNIX-style applications that invoke system calls directly.

![](/files/-MPaOGxVCA8S3pdIfLbT)

![](/files/-MPaJGHLADJVn6SH6RbN)

![](/files/-MPaJKE6jxKoRUFt4SNm)

![](/files/-MPaOXtG57zeS8WjN_Hb)

![](/files/-MPaJNk_qjo_xFCV5v2D)

![](/files/-MPaOaPc-is3YFX7hZ2h)

![](/files/-MPaJTCOs6A-bCKpiheN)

## Links

* [Paper PDF](https://research.cs.wisc.edu/wind/Publications/ibench-sosp11.pdf)
* [34 traces from the iBench task suite open-sourced by ADSL](https://research.cs.wisc.edu/adsl/Traces/ibench/)
* Thanks to Jiaxin Lin, a fellow classmate, for her paper reading notes!


# Biscuit: The benefits and costs of writing a POSIX kernel in a high-level language

## One-line Summary

This paper analyzes (duh) the benefits and costs of writing a POSIX kernel in a high-level language, Go.

## Paper Structure Outline

1. Introduction
2. Related work
3. Motivation
   1. Why C?
   2. Why an HLL?
4. Overview
5. Garbage Collection
   1. Go's collector
   2. Biscuit's heap size
6. Avoiding heap exhaustion
   1. Approach: reservations
   2. How Biscuit reserves
   3. Static analysis to find s
      1. Basic MAXLIVE operation
      2. Handling loops
      3. Kernel threads
   4. Limitations
   5. Heap exhaustion summary
7. Implementation
8. Evaluation
   1. Biscuit's use of HLL features
   2. Potential to reduce bugs
   3. Experimental Setup
   4. HLL tax
   5. GC delays
   6. Sensitivity to heap size
   7. Go versus C
      1. Ping-pong
      2. Page-faults
   8. Biscuit versus Linux
   9. Handling kernel heap exhaustion
   10. Lock-free lookups
9. Discussion and future work
10. Conclusions

## Background & Motivation

The main reason for using low level languages like C to implement a kernel is that C supports low-level techniques that can help performance (pointer arithmetic, explicit memory allocation, etc.).

High level languages (HLL), on the other hand, have some potential advantages compared to C:

1. Automatic memory management: reduces programmer effort and use-after-free bugs
2. Type-safety: detects bugs
3. Runtime typing and method dispatch: helps with abstraction
4. Language support for threads and synchronization eases concurrent programming

With the idea of exploring the possibility of using a HLL to implement a monolithic POSIX-style kernel in mind, the authors present Biscuit, a kernel written in Go, which has good performance.

## Design and Implementation

The Biscuit kernel is written using 27583 lines of Go, 1546 lines of assembly, and no C. Biscuit provides 58 syscalls and it has enough POSIX compatibility to run some existing server programs (NGINX, Redic, etc.).

### Garbage collection

Biscuit uses Go's collector, which suspends ordinary execution on all cores ("stop-the-world" pause of \~10μs) twice during a collection. This hurts tail latency the most, and it's especially bad for machines that are dependent on pauses (e.g., datacenters).

### Avoiding heap exhaustion

Heap exhaustion refers to live kernel data completely filling the RAM allocated for the heap. Waiting for memory in allocator might lead to deadlocks; Checking and handling allocation failure (like C kernels) is difficult to get right, and Go does not expose failed allocations. Biscuit uses reservations as a solution:

![](/files/-MP4uHFfo8tDzmUwUL3X)

A syscall does not start until either it can reserve enough heap memory or a killer thread frees up some memory. To execute a syscall,&#x20;

```
reserve()
    (no locks held)
    evict, kill
    wait...
sys_read()
    ...
unreserve()
```

## Evaluation

Some missing features of Biscuit:

* Scheduling priority (relies on Go runtime scheduler)
* Does not handle large multicore machines or NUMA
* Does not swap or page out to disk
* Does not implement reverse page mappings (revoke shared pages)
* Security features (Users, access control lists, address space randomization)
* 58 out of 300-400 syscalls

The experiments used three kernel-intensive applications: CMailbench, NGINX, and Redis.

![The use of Go improved the outcome of 40 out of 65 Linux execute-code bugs in the CVE database: 8 won't happen at all and 32 will cause runtime error + panic.](/files/-MP4wAx-RAx7das4jzq8)

![The performance of Biscuit is in the same league as Linux](/files/-MP4wyzyMM4J04x31h32)

![Measurement of HLL tax. Prologue cycles are the most expensive. The cost of the GC cycles increases with the size of live kernel heap.](/files/-MP4xEwQz26NFy9xdZyb)

{% hint style="info" %}

* Tput: throughput in application requests per second
* Prologue cycles: the fraction of total time used by compiler-generated code at the start of each function that checks whether the stack must be expanded, and whether the garbage collector needs a stop-the-world pause
* Safety cycles: the cost of runtime checks for nil pointers, array and slice bounds, divide by zero, and incorrect dynamic casts
* Alloc cycles: the time spent in the Go allocator, examining free lists to satisfy allocation requests (but not including concurrent collection work)
  {% endhint %}

![The performance of code paths are compared using two benchmarks: ping-pong and page-faults. Go has 5% - 15% performance tax.](/files/-MP4xkrdYGHuA1mNXlNl)

## New Vocabulary

* [Goroutines](https://www.geeksforgeeks.org/goroutines-concurrency-in-golang/)
* CVE: Common Vulnerabilities and Exposures
* Code path: the set of specific instructions that are actually executed during a single run of a program or program fragment.

## Links

* [Paper PDF](https://www.usenix.org/system/files/osdi18-cutler.pdf)
* [Presentation Audio at OSDI '18](https://www.usenix.org/conference/osdi18/presentation/cutler)
* [Presentation Slides](https://www.usenix.org/sites/default/files/conference/protected-files/osdi18_slides_cutler.pdf)


# Data Domain: Avoiding the Disk Bottleneck in the Data Domain Deduplication File System

## One-line Summary

The Data Domain Deduplication File System introduces some techniques to relieve the disk bottleneck in previous works. The techniques combined can remove 99% of the disk accesses for deduplication.

## Paper Structure Outline

1. Introduction
2. Challenges and Observations
   1. Variable vs. Fixed Length Segments
   2. Segment Size
   3. Performance-Capacity Balance
   4. Fingerprint vs. Byte Comparisons
3. Deduplication Storage System Architecture
   1. Content Store
   2. Segment Store
   3. Container Manager
4. Acceleration Methods
   1. Summary Vector
   2. Stream-Informed Segment Layout
   3. Locality Preserved Caching
   4. Accelerated Segment Filtering
5. Experimental Results
   1. Results with Real World Data
   2. I/O Savings with Summary Vector and Locality Preserved Caching
   3. Throughput
   4. Discussion
6. Related Work
7. Conclusions

## Background & Motivation

The throughput for data deduplication in [Venti](/earlier-readings-and-notes/index/venti-a-new-approach-to-archival-storage) was bad. This work aims to resolve those issues.

## Design and Implementation

The three main performance enhancement techniques are:

1. Summary vectors
2. Stream-informed layout
3. Locality-preserved caching

The Data Domain File System also featured a layered file system architecture.

### Layered file system

![](/files/-MPjX1Xi6GzgqmVmML2W)

* Content Store: Manages data in files; Breaks data into segments; Does the fingerprinting
* Segment Store: Maps segment descriptors (fingerprints) to data; Performs deduplication; Keeps track of references, updates segment index; Compresses segments and pushes to container layer
* Container Manager: Provides container abstraction; Write fixed-sized containers entirely

### Summary vector (bloom filter)

![](/files/-MPjdTusuanu8pTD13QZ)

This avoids going to disk and writing data if the data already exists. Bloom filters are used to check the fingerprint. It may produce false positives but no false negatives (i.e. the summary vector will not tell you that an existing fingerprint doesn't exist), so it's only a slight performance problem.&#x20;

### Stream-informed layout

![CS 736 course notes by Prof. Andrea Arpaci-Dusseau](/files/-MPjb0z7IsGThLEHVz44)

The motivation is that fingerprints are random, distributed, but the same group of fingerprints is likely to occur at the same time in the future. The idea is to dedicate containers to holding segments & fingerprints in the logical order. This improves locality and fewer number of I/Os by the container manager.

### Locality-preserved caching

![CS 736 course notes by Prof. Andrea Arpaci-Dusseau](/files/-MPjamQnk2H7dYkmLH6k)

An extension of the idea that motivated stream-informed layout. We divide the cache into container-sized units and cache on the granularity of containers. If we need one of the items in the container, we will likely need all of them as well.

### Procedures of a segment write

1. Check if the fingerprint is in the segment cache. If yes, then we are done. Else, continue to step 2.
2. Check the summary vector to see if the data is already written.
   1. If not (new data), then append it to the current container (for locality with the other data in the container). Once the container is full, it is passed off to the container manager.
   2. If yes (old data), it might be a false positive so we check if it really exists.
      1. Check the index, if it's really in there (duplicate), insert it into the segment cache along with all the other fingerprints in the container.
      2. Otherwise, go to step 2.1 (append to the current container).

## Evaluation

![\~100-fold performance improvement using the combination of the two techniques](/files/-MPjdaDi6pzit9x5u3Wj)

## New Vocabulary

* [Bloom filter](https://www.youtube.com/watch?v=kfFacplFY4Y\&ab_channel=SpanningTree): A space-efficient probabilistic data structure that is used to test whether an element is a member of a set.

## Links

* [Paper PDF](https://www.usenix.org/legacy/events/fast08/tech/full_papers/zhu/zhu.pdf)
* [Presentation Video at FAST '08](https://www.usenix.org/conference/fast-08/avoiding-disk-bottleneck-data-domain-deduplication-file-system)
* [Presentation Audio at FAST '08](https://c59951.ssl.cf2.rackcdn.com/legacy_media/fast08/tech/full_papers/zhu/zhu.mp3)

{% file src="/files/-MPkHyAjthcnHXZBzloY" %}
Prof. Andrea Arpaci-Dusseau's course notes on Data Deduplication
{% endfile %}


# Disco: Running Commodity Operating Systems on Scalable Multiprocessors

## One-line Summary

Disco uses virtual machines to run multiple commodity operating systems on large-scale shared-memory multiprocessors. Disco VMM hides NUMA-ness from non-NUMA aware OSes, requires low effort to implement, and introduces moderate overhead due to virtualization.

## Paper Structure Outline

1. Introduction
2. Problem Description
3. A Return to Virtual Machine Monitors
   1. Challenges Facing Virtual Machines
4. Disco: A Virtual Machine Monitor
   1. Disco's Interface
   2. Implementation of Disco
      1. Virtual CPUs
      2. Virtual Physical Memory
      3. NUMA Memory Management
      4. Virtual I/O Devices
      5. Copy-on-write Disks
      6. Virtual Network Interface
   3. Running Commodity Operating Systems
      1. Necessary Changes for MIPS Architecture
      2. Device Drivers
      3. Changes to the HAL
      4. Other Changes to IRIX
   4. SPLASHOS: A Specialized Operating System
5. Experimental Results
   1. Experimental Setup and Workloads
   2. Execution Overheads
   3. Memory Overheads
   4. Scalability
   5. Dynamic Page Migration and Replication
6. Related Work
   1. System Software for Scalable Shared Memory Machines
   2. Virtual Machine Monitors
   3. Other System Software Structuring Techniques
   4. ccNUMA Memory Management
7. Conclusions

## Background & Motivation

The motivation is to enable existing commodity operating systems to handle Non-Uniform Memory Access (NUMA) architectures. Instead of modifying existing operating systems to run on scalable shared-memory multiprocessors, an additional layer (VM monitor) is inserted between the hardware and the OS.

![Course notes by Prof. Andrea. Left: SMP (symmetrical multiprocessor uniform memory access machine), right: cc-NUMA](/files/-MQA2fE9_jzOzxEgT_68)

Cache-coherent Non-Uniform Memory Architecture (cc-NUMA) makes hardware scalable, while SMP ensures the same performance to all memory from everywhere. Both ensure correctness, though.

## Design and Implementation

![Disco is a layer between OSes and hardware](/files/-MQA3gfQerincnJtMsnw)

The advantages of using virtual machines in the context of this work are:

* The Disco layer understands the NUMA architecture
* It's a portability layer
* Monitors are smaller and easier to understand & trust than operating systems
* Allows to run different OSes concurrently (almost unmodified)

The drawbacks of using virtual machines are:

* Overhead: cost of virtualizing
  * Time: VMM (Disco) acts as an emulator. Most instructions can just run, but privileged instructions + TLB instructions must be trapped & emulated
  * Space: Multiple copies (OS code & each OS's file cache) waste memory
* Resource management: Lack of information to make good policy decisions
  * Lost information about what is being used
    * CPU - idle thread
    * Memory - pages on the free list
* Communication and Sharing problems:
  * Hard to communicate between standalone VMs
  * Most OSes require exclusive access to disks

![High-level challenges of using virtual machines](/files/-MQA55VUAcB9oZXQ4tIu)

![How Disco virtualizes CPU. Three priviliged levels: user, supervisor, and kernel.](/files/-MQA5OI6DAXsRxT9DULG)

![How Disco virtualizes memory. Users generate virtual addresses, OS translates to physical addreses, Disco translates to machine addresses.](/files/-MQA66WQp-UcD3uPFbQ9)

![](/files/-MQA7uY2DNaUQ-Svm14k)

![Records copy-on-write to track shared data efficiently.](/files/-MQAD8k96VeiSOotNKtV)

![Send becomes additional mapping (emulate device); Copy becomes additional mapping](/files/-MQADB-JW4Qb--33GABY)

![Changes Disco made to IRIX to improve performance](/files/-MQADfpA09ZEKgQDzYoD)

## Evaluation

![](/files/-MQAFIpTg1I5uKynaZrm)

![Pmake & Database do a lot of syscalls, often traps into Disco, which then goes to kernel. The extra 16% overhead for those workloads is due to the extra work handling TLB misses. The kernel time being less is because Disco zeros the pages (does work instead of IRIX).](/files/-MQAFM0jMLb85-5q1opo)

![Disco does a good job sharing buffer cache space across VMs and sharing IRIX text.](/files/-MQAG2MPJtAnrwMxI1RN)

![No migration + replication, just looking at how much more scalable is Disco than IRIX due to optimizations of not having locks in which IRIX does a bad job at. IRIX on 8-processor cc-NUMA machine. 2VM -> 8VM actually improves because Disco does not have bad lock](/files/-MQAG52ltcDsjRMJnGfp)

![Much less time accessing remote memory, more local memory.](/files/-MQAI9HvVbhyRdF1GfYr)

This paper started off VMWare (which was founded by authors of Disco in 1998 and successfully commercialized this work) and revived virtual machines for the next 20 years. Now VMs are commodities, and every cloud provider and virtually every enterprise uses VMs today.

## New Vocabulary

* IRIX: A variety of UNIX System V with BSD extensions.
* NUMA: [What is NUMA?](https://www.kernel.org/doc/html/v4.18/vm/numa.html)

## Links

* [Paper PDF](https://bob.cs.ucdavis.edu/assets/ecs251/bugnion97.pdf)
* [Paper review notes from CS 443 @ Northwestern by Joseph Paris](https://users.cs.northwestern.edu/~fabianb/classes/cs-443-s05/review-disco-jparis.pdf)
* [Discussion panel from CS 736 @ UW-Madison](http://pages.cs.wisc.edu/~swift/classes/cs736-fa12/blog/2012/09/disco_running_commodity_operat.html)
* [Lecture slides from CS 262a @ Berkeley by Prof. Ion Stoica and Ali Ghodsi](https://ucbrise.github.io/cs262a-spring2018/notes/10-VMs-Disco-Xen.pdf)

{% file src="/files/-MQ9ak0bKd1b2skrAZVU" %}
Prof. Andrea's notes on Disco
{% endfile %}


# FFS: A Fast File System for UNIX

## One-line Summary

As a reimplementation of the UNIX file system, FFS addresses the problem of low data throughput rates/bandwidth usage by making the file system "disk aware".

## Paper Structure Outline

1. Introduction
2. Old file system
3. New file system organization
   1. Optimizing storage utilization
   2. File system parameterization
   3. Layout policies
4. Performance
5. File system functional enhancements
   1. Long file names
   2. File locking
   3. Symbolic links
   4. Rename
   5. Quotas
6. Acknowledgments

## Background & Motivation

![The old UNIX file system.](/files/-MPUbjqDfyiyjgEYCteK)

As a predecessor of FFS, the original UNIX operating system (very-simple file system, VSFS) was simple and easy-to-use. It had some major problems, though:

* The biggest problem was **terrible performance**: As measured in this paper, in some cases, the disk bandwidth utilization was a mere 3%. The throughput was also low.
* This was because of **low locality**: The old UNIX file system treated the disk like it was a random-access memory, so inodes and data blocks were scattered across the disk.
* Another problem was **file system fragmentation**: The free space was not carefully managed. As data come and go, the file system got fragmented in that the free list ended up pointing to blocks spread across the disk, so accessing a logically contiguous file required going back and forth across the disk. (This problem was originally solved by disk defragmentation tools.)
* The original **block size was too small** (512 bytes): A smaller size was good for minimizing internal fragmentation (waste within the block) but **bad for transfers** as each block required a positioning overhead.

To resolve these issues, the solution is to make the file system "disk aware".

## Design and Implementation

### Cylinder groups

![A disk is divided into several cylinder groups.](/files/-MPUbAPR9yVmbkTbQnyG)

![IRL, as disks hide details of their geometry from clients, modern file systems organize the drive into block groups, each of which is a consecutive portion of the disk's address space. Note that IRL each group will contain many more blocks.](/files/-MPUbXcTn9_DGW7SIm1z)

![What FFS keeps within a single cylinder group. ib and db are inote bitmap and data bitmap that tracks whether the inodes and data blocks of the group are allocated. Bitmaps replaces freelists as bitmaps are faster to update/lookup and are space efficient.](/files/-MPUbd3F-sSI0uq5Eqj5)

By putting two files in the same group, FFS ensures that accessing one after the other will not result in long seeks across the disk.

The superblock is duplicated and stored in different cylinder groups with a rotated location/offset. This reduces the chance of data loss due to corruption of the superblock (top platter damage).

### Layout policies

The basic mantra is simple: Keep related stuff together and keep unrelated stuff far apart.

* Placing directories: FFS finds the cylinder group with a low number of allocated directories and a high number of free inodes.
* Placing files: First, data blocks of a file in the same group as its inode are allocated. Second, file inodes are allocated in the same cylinder group with other files within the same directory.

A problem is the large files exception: Single, large files (> 48KB) can fill nearly all of a group, while ideally none of the cylinder groups should be completely full. The solution is to redirect block allocation to a different cylinder group when a file exceeds 48KB and at every megabyte thereafter.&#x20;

{% hint style="info" %}
The first spill over point at 48 kilobytes is the point at which a file on a 4096 byte block file system first requires a single indirect block. This appears to be a natural first point at which to redirect block allocation. The other spillover points are chosen with the intent of forcing block allocation to be redirected when a file has used about 25% of the data blocks in a cylinder group. In observing the new file system in day to day use, the heuristics appear to work well in minimizing the number of completely filled cylinder groups.
{% endhint %}

### Larger blocks & fragments within blocks

The size of a file system block is increased from 512 bytes to 4096 bytes (4 KB). This improves performance because:

1. Each disk transfer access more data
2. More files can be described without the need to access indirect blocks (since the direct blocks now contain more data)

### File system parameterization (rotationally optimal block placement)

![Think about this: In sequential reads, FFS firstly reads block 0. By the time the read is finished, block 1 had rotated under the head and to get to block 1, we now need a full rotation. FFS resolves this by figuring out the specific performance parameters of the disk and use those to decide on the exact staggered layout scheme.](/files/-MPUf-CYoxfWW1w3w-kg)

## Evaluation

![FFS is faster for both reads and writes and the disk bandwidth (\~3% to 47%)](/files/-MPURxe_6xh4ZKKwuGLf)

FFS also introduced some neat file system functional enhancements that are routines of today's operating systems and likely help FFS gain a stronger user base:

* **Long file names**: File names can now be of nearly arbitrary length (previously: 8 characters. now: 255 characters)
* **File locking**: Programs can now apply advisory shared or exclusive locks at the granularity of files
* **Symbolic links**: Users can now create an "alias" to any other file or directory on a system and thus are much more flexible
* **File renaming**: Atomic `rename()` operation
* **Quotas**: Restricts the amount of file system resources (#inodes & #disk blocks) that a user can obtain

## New Vocabulary

* Hard locks & advisory locks: A hard lock is always enforced when a program tries to access a file, while an advisory lock is only applied when it is requested by a program.

## Links

* [Paper PDF](https://dsf.berkeley.edu/cs262/FFS-annotated.pdf)
* [FFS in OSTEP](http://pages.cs.wisc.edu/~remzi/OSTEP/file-ffs.pdf)
* [FFS in CS 537 @ UW-Madison](http://pages.cs.wisc.edu/~shivaram/cs537-sp20-notes/ffs/cs537-ffs-notes.pdf)

{% file src="/files/-MPkIY8GLc3BWa0IuLET" %}
Prof. Andrea's slides on FFS and LFS
{% endfile %}


# From WiscKey to Bourbon: A Learned Index for Log-Structured Merge Trees

## One-line Summary

Bourbon is a log-structured merge tree that utilizes machine learning to provide fast lookups.

## Paper Structure Outline

1. Introduction
2. Background
   1. LSM and LevelDB
   2. WiscKey
3. Learned Indexes: a Good Match for LSMs?
   1. Learned Indexes: Beneficial Regimes
   2. Learned Indexes with Writes
4. Bourbon Design
   1. Learning the Data
   2. Supporting Variable-size Values
   3. Level vs. File Learning
   4. Cost vs. Benefit Analyzer
      1. Wait Before Learning
      2. To Learn a File or Not
   5. Bourbon: Putting it All Together
5. Evaluation
   1. Which Portions does Bourbon Optimize?
   2. Performance under No Writes
      1. Datasets
      2. Load Orders
      3. Request Distributions
   3. Range Queries
   4. Efficacy of Cost-benefit Analyzer with Writes
   5. Real Macrobenchmarks
      1. YCSB
      2. SOSD
   6. Performance on Fast Storage
   7. Performance with Limited Memory
   8. Error Bound and Space Overheads
6. Related Work
7. Conclusions

## Background & Motivation

The work is based on an existing LSM system, WiscKey, which is significantly faster than LevelDB and RocksDB. The authors analyzed WiscKey and derived five learning guidelines that aid an LSM system to successfully incorporate learned indexes. These guidelines are applied to build Bourbon, a learned-index implementation of WiscKey.

## Design and Implementation

### Five learning guidelines

1. Favor learning files at lower levels (which live longer)
2. Wait before learning a file (very short-lived files in every level due to consecutive compactions to newly generated files)
3. Do not neglect files at higher levels (serve more internal lookups)
4. Be workload- and data-aware (#lookups vary a lot in different scenarios)
5. Do not learn levels for write-heavy workloads (level changes quickly)

### Beneficial regimes

![Read-only experiments on WiscKey when data resides on memory and different storage devices. Works better with faster data access (InMemory > Optane > SATA SSD)](/files/-MQDP0SERMFxyjhW0nBq)

Learned indexes can only speed up indexing time.

### Bourbon Learning

![](/files/-MQDQ7GrJf0GJ7_2kwDd)

Bourbon uses piecewise linear regression (PLR) to model the data as it has low overheads during learning and lookups, and the space overhead is small as well. Bourbon can learn individual sstables files (file learning) or entire levels (level learning). Level learning can be beneficial for read-only workloads, while for mixed workloads, level learning performs worse than file learning.

### Cost-benefit analyzer (CBA)

The motivation for an online Cost vs. Benefit Analyzer (CBA) is to filter out short-lived files as they are not worth learning. Doing so wastes resources and has little benefit. CBA uses stats of previous files at the same level.

To filter out short-lived files, Bourbon waits for a time threshold, Twait, before learning a file. The max time to learn a file is \~40ms, thus Bourbon sets Twait to be 50ms. However, learning a long-lived file may not be beneficial (and vice versa). Intuitively, as long as the benefit of the model (B\_model) outweighs the cost of building the model (C\_model), learning a file is profitable.

#### Estimating the cost (C\_model)

If we assume learning happens in the background (using idle cores), then C\_model is 0. Bourbon takes a conservative approach, though, and assumes that the learning threads will interfere and cause some slow down. As a result, we define Cmodel to be equal to T\_build, the time to train the PLR model for a file. As T\_build is linearly proportional to the number of data points in a file, we define T\_build to be the product of (1) the number of data points in the file and (2) the avg time to train a data point (measured offline).

#### Estimating the benefit (B\_model)

Bourbon defines the benefit of learning a file to be:

$$
B\_{model} = (T\_b - T\_m) \* N
$$

{% hint style="info" %}

* T\_b: Average time for the lookup in baseline
* T\_m: Average time for the lookup in model paths
* N: Number of lookups the file serves in its lifetime
  {% endhint %}

Then, we divide the internal lookups into negative and positive ones as most negative lookups terminate at the filer.

$$
B\_{model} = ((T\_{n.b} - T\_{n.m}) \* N\_n) + ((T\_{p.b} - T\_{p.m}) \* N\_p)
$$

{% hint style="info" %}

* N\_n & N\_p: Number of negative and positive internal lookups, respectively
* T\_(n.b) and T\_(p.b): Time in the baseline path for negative and positive internal lookups, respectively
* T\_(n.m) and T\_(p.m): Model counterparts
  {% endhint %}

To estimate the number of lookups (N\_n & N\_p) and the time the lookups take (T\_{n.b} & T\_{p.b}), CBA keeps track of the statistics of files that (1) have lived their lifetimes and (2) are at the same level (as stats vary significantly across levels).

Estimation of the above quantities are done during T\_wait:

* T\_(n.b) and T\_(p.b): DuringTwait, lookups are served in the baseline path. These times are used to estimate Tn.b & Tp.b.
* T\_(n.m) and T\_(p.m): Estimated as the avg of those of all other files at the same level.
* N\_n & N\_p: Same as above but normalized by a factor f (f = s / s’: s is the size of the file, while s’ is the avg size of files at this level). While estimating the above quantities, short-lived files are filtered out.

If C\_model < B\_model, a file will be learned. If multiple files are chosen to be learned at the same time, they are put on a max priority queue so that files that would deliver the most benefit are prioritized.

Possible directions for future improvements include:

* Better estimations of N\_n, N\_p, T\_(n.m) & T\_(p.m)
* Develop a model of T\_build on-line or calculate C\_build such that it is not simply T\_build
* Sort the work queue by some function other than B\_model - C\_model

![Bourbon lookups](/files/-MQDQDIbhf-6lrf0pcpC)

## Evaluation

![](/files/-MQDQfLsZe9r0YQVtUTQ)

![Bourbon works better with datasets of fewer segments](/files/-MQDQihAJoNUDRJfDhM0)

![Bourbon works better with sequential loads than random loads, as random loads cause many negative lookups and Bourbon offers less gain for negative lookups](/files/-MQDQr0hCXwmH2WRsiDg)

![CBA: At lower write rates, learn most files, resulting in best foreground latency. At higher write rates, limited learning, resulting in significantly less learning time and a foreground latency close to always-learn policy. At all write rates: Minimal total CPU time.](/files/-MQDRBoCe4_tSm4dqyRw)

![](/files/-MQDRUeXXzItnC423r94)

![YCSB & SOSD: Read-only gains holds for real benchmarks; Have less gain when write rate is higher; Accelerate reads without affecting writes](/files/-MQDRXBZ0PSnrfVlB9Mo)

## New Vocabulary

* [Log-structured merge trees](https://en.wikipedia.org/wiki/Log-structured_merge-tree)

## Links

* [Paper PDF](https://www.usenix.org/system/files/osdi20-dai_0.pdf)
* [Presentation video at OSDI '20](https://www.youtube.com/watch?v=EUxEx5hwLXk)
* [Presentation slides at OSDI '20](https://www.usenix.org/sites/default/files/conference/protected-files/osdi20_slides_dai.pdf)
* Thanks to Yifan Dai and Yien Xu for the paper review notes!

{% file src="/files/-MQDI53\_q8JImOVdnYi9" %}
Prof. Andrea's course slides on the SSD paper and Bourbon
{% endfile %}


# LegoOS: A Disseminated, Distributed OS for Hardware Resource Disaggregation

## One-line Summary

The traditional monolithic server model in datacenters is having issues in resource utilization, elasticity, heterogeneity, and failure handling. LegoOS breaks down traditional OS functionalities into hardware components like Lego bricks and connects them with fast networks.

## Paper Structure Outline

1. Introduction
2. Disaggregate Hardware Resource
   1. Limitations of Monolithic Servers
   2. Hardware Resource Disaggregation
   3. OSes for Resource Disaggregation
3. The Splitkernel OS architecture
4. LegoOS Design
   1. Abstraction and Usage Model
   2. Hardware Architecture
   3. Process Management
      1. Process Management and Scheduling
      2. ExCache Management
      3. Supporting Linux Syscall Interface
   4. Memory Management
      1. Memory Space Management
      2. Optimization on Memory Accesses
   5. Storage Management
   6. Global Resource Management
   7. Reliability and Failure Handling
5. LegoOS Implementation
   1. Hardware Emulation
   2. Network Stack
   3. Processor Monitor
   4. Memory Monitor
   5. Storage Monitor
   6. Experience and Discussion
6. Evaluation
   1. Micro- and Macro-benchmark Results
   2. Application Performance
   3. Failure Analysis
7. Related Work
8. Discussion and Conclusion

## Background & Motivation

In datacenters, the monolithic server model has been used for decades. It's facing some limitations:

1. **Inefficient resource utilization**: With a server being the physical boundary of resource allocation, under-utilization occurs. See the figure below for an example.
2. **Poor hardware elasticity**: It's difficult to add/move/remove/reconfigure hardware components after they have been installed in a monolithic server.
3. **Coarse failure domain**: When a hardware component in a monolithic server fails, the whole server goes down.
4. **Bad support for heterogeneity**: As the monolithic server model tightly couples hardware devices with each other and with a motherboard, it is very difficult to make new hardware devices (GPU, TPU, DPU, NVM, NVMe-based SSDs, etc.) work with existing servers.

![Resource under-utilization](/files/-MPA7QH9fGOCMySrb1DX)

To break the server-centric monolithic server model, the authors suggested a hardware resource disaggregation architecture.

## Splitkernel

![As a backbone of LegoOS, Splitkernel disseminates traditional OS functionalities into loosely-coupled monitors (process monitor, memory monitor, and storage monitor) and offers resource allocation and failure handling of a distributed set of hardware components. ](/files/-MPA8xVpmGGTYB8hn_Ih)

## LegoOS Design

LegoOS' design targets three types of hardware components: processor, memory, and storage. We call them pComponent, mComponent, and sComponent.

### Hardware Architecture

![](/files/-MPAC0dFMVxywv_sDYD_)

1. **Separating process and memory functionalities**: All hardware memory functionalities (page tables, TLBs, MMU) are moved to mComponents. Only caches are left at the pComponent side. "With a clean separation of process and memory hardware units, the allocation and manage- ment of memory can be completely transparent to pCom- ponents. Each mComponent can choose its own memory allocation technique and virtual to physical memory ad- dress mappings (e.g., segmentation)."
2. **Processor virtual caches**: As all memory functionalities are moved to mComponents, pComponents will only see virtual addresses. To resolve this, LegoOS organizes all levels of pComponent caches as virtual caches. With virtual caches comes two potential problems: synonyms and homonyms. LegoOS resolves synonyms by not allowing writable inter-process memory sharing, and it resolves homonyms by storing an address space ID (ASID) with each cache line, and differentiate a virtual address in different address spaces using ASIDs.
3. **Separating memory for performance and for capacity**

### Process Management

> Let a thread run to the end with no scheduling or kernel preemption except when a pComponent has to schedule more threads than its cores. (Because LegoOS does not push for perfect core utilization when scheduling individual threads and instead aims to minimize scheduling and context switch performance overheads.)

LegoOS also process monitor configures and amnages ExCache. Finally, LegoOS supports Linux ABIs for backward compatibility and easy adoption of LegoOS.

### Memory Management

* Virtual memory space management: A two-level approach to manage distributed virtual memory spaces.
  * Higher level: Split each virtual memory ad- dress space into coarse-grained, fix-sized virtual regions, or vRegions (e.g., of 1 GB).
  * Lower level: Stores user process virtual memory area (vma) information, such as virtual address ranges and permissions, in vma trees.
* Physical memory space management: Each mComponent can choose their own way of physical memory allocation and own mechanism of virtual-to-physical memory address mapping.

### Storage Management

> LegoOS implements core storage functionalities at sComponents. To cleanly separate storage functionalities, LegoOS uses a stateless storage server design, where each I/O request to the storage server contains all the information needed to fulfill this request, e.g., full path name, absolute file offset, similar to the server design in NFS v2.

### Global Resource Management

LegoOS uses a two-level resource management mechanism:

* Higher level: Three global resource managers for process, memory, and storage resources are used. They perform coarse-grained global resource allocation and load balancing.
* Lower level: Each monitor can employ its own policies and mechanisms to manage its local resources.

## LegoOS Implementation

LegoOS supports 113 syscalls, 15 pseudo-files, and 10 vectored syscall opcodes. These Linux interfaces are sufficient to run many unmodified datacenter applications.

## Evaluation

![](/files/-MPA7k0t68Yy-b6nQbyB)

![](/files/-MPA7pnBgPRn60oXJXyF)

## New Vocabulary

* Monolithic server: A single server that contains all the hardware resources (typically a processor, some main memory, and a disk or an SSD) that are needed to run a user program.
* SLOC: Abbreviation for "Source Lines of Code".
* ABI: [Application Binary Interface](https://stackoverflow.com/a/2456882).
* [Synonyms & Homonyms](http://www.inf.ed.ac.uk/teaching/courses/car/Notes/2016-17/lecture09-virtual_memory.pdf): Synonyms happens when a physical address maps to multiple virtual addresses (and thus multiple virtual cache lines) as a re- sult of memory sharing across processes, and the update of one virtual cache line will not reflect to other lines that share the data. The homonym problem happens when two address spaces use the same virtual address for their own different data.
* Cache lines: A cache line is the unit of data transfer between the cache and main memory.

## Links

* [Paper PDF](https://www.usenix.org/system/files/osdi18-shan.pdf)
* [Presentation Video at OSDI '18](https://www.youtube.com/watch?v=GX74Q2-ZOQE)
* [Presentation Video at USENIX ATC '19](https://www.youtube.com/watch?v=KJqYHuL59_s)
* [Presentation Slides](https://www.usenix.org/sites/default/files/conference/protected-files/osdi18_slides_shan.pdf)
* [LegoOS on GitHub](https://github.com/WukLab/LegoOS)
* Thanks to Yuhao Zhang for the review notes!


# LFS: The Design and Implementation of a Log-Structured File System

## One-line Summary

A new file system structure that tries to use the disk sequentially using log-like structures.

## Paper Structure Outline

1. Introduction
2. Design for file systems of the 1990's
   1. Technology
   2. Workloads
   3. Problems with existing file systems
3. Log-structured file systems
   1. File location and reading
   2. Free space management: segments
   3. Segment cleaning mechanism
   4. Segment cleaning policies
   5. Simulation results
   6. Segment usage table
4. Crash recovery
   1. Checkpoints
   2. Roll-forward
5. Experience with the Sprite LFS
   1. Micro-benchmarks
   2. Cleaning overheads
   3. Crash recovery
   4. Other overheads in Sprite LFS
6. Related Work
7. Conclusion

## Background & Motivation

The motivation for LFS was because of the explosive improvements of system memories and processor speed and a relatively slow improvement of disk transfer bandwidth and access time, which made I/O a bottleneck. Also, the new workloads of doing small, random reads (and the poor performance of existing file systems in doing so compared to the sequential I/O, and the fact that RAID-5 is especially bad with small random reads) also contributed to the motivation. The idea, then, is to have the file system use the disk purely sequentially.

## Design and Implementation

The new type of file system proposed in this paper, Log-structured File System (LFS), has the following contributions:

* Log structure: LFS converts small synchronous writes into asynchronous sequential writes. When writing to disk, LFS first buffers all updates in an in-memory segment. When a segment is full, it is written to free locations on the disk in one long, sequential transfer.
* Garbage collection: Low-overhead segment cleaning mechanisms that increase the compactness of the contents on the disk.
* Crash recovery: This is fast and intuitive due to the nature of LFS. It occasionally checkpoints the imap to disk, and it uses two checkpoint regions to ensure consistency.

### Log structure

![The basic idea is to simply write all updates to the disk sequentially. D: data block, I: inode block](/files/-MPXM_DEm6xPCYywnNWE)

Writing to disk sequentially is not enough: Writing to two blocks right next to each other may actually incur the overhead of most of a rotation. What we actually want is to issue a large number of contiguous writes to the drive. LFS uses an ancient technique, write buffering (into large chunks called segments), to resolve this issue.&#x20;

![An example of buffering two sets of updates into a small segment: (1) four block writes to file j (2) one block added to file k. The entire seven blocks are committed at once.](/files/-MPXO0-K3EDqKJUA-nvF)

LFS uses the inode map (imap) to locate inodes. The imap takes an inode number and produces the disk address of the most recent version of the inode. Finally, we also have a checkpoint region (CR) that contains pointers to the latest pieces of the inode map. The CR is only updated periodically.

![Chunks of the inode map are placed right next to where it is writing all of the other new information. The CR is placed at a fixed location.](/files/-MPXQu5s94DUut44s9-Z)

### Garbage collection

While writes are efficient, old, garbage data are scattered across the disk. LFS keeps only the latest live version of a file by periodically cleaning the garbage. LFS reclaims segments, not individual inodes or data blocks. The LFS cleaner reads in M existing segments, compacts their contents into N new segments (N < M), and then writes the N segments to disk in new locations. The old M segments are then freed and are available for use for later writes.

#### Mechanism: Determining which blocks are alive

![The segment summary block (SS) stores for each data block, its inode number (which file it belongs to) and its offset (which block of the file this is).](/files/-MPXWHs5FeyQC2oK3BRE)

With the SS block and the imap, it is straightforward to determine whether a block is live or dead:

```
(N, T) = SegmentSummary[A]
inode = Read(imap[N})
if (inode[T] == A):
    // block D is alive
else:
    // block D is garbage
```

#### Policy: How often should the cleaner run, and which segments should it pick?

* **When**: Periodically/during idle time/when the disk is full
* **Which**: Challenging! In the original paper, the authors proposed hot and cold segments. Cold segments are cleaned sooner and hot segments are cleaned later. However, this is not perfect: see [this paper](https://homes.cs.washington.edu/~tom/pubs/lfs-adapt.html) for an explanation and a better heuristic.

### Crash recovery

1. **Checkpoint**: Upon recovery, reads CR to find most imap pointers and segment tail. To address the problem of crashing during checkpointing, two CRs are used, and only one is overwritten at a time.&#x20;
2. **Roll-forward**: Thank you, database community! This technique scans through the log segments that were written after the last checkpoint.

### Implementation

Sprite LFS is a prototype LFS that outperforms the current UNIX file systems by an order of magnitude for small-file writes while matching or exceeding UNIX performance for reads and large writes.

![](/files/-MPXKKWOnJDk5rFWeFlq)

## Evaluation

![LFS outperforms existing file systems drastically in small-file performance...](/files/-MPXZQDtx7Hp7jnnuXZB)

![...while exceeding/maintaining performance in most cases.](/files/-MPXZWRI3opTTVB5yX25)

![Cleaning overhead is low.](/files/-MPX_PcI7z3aKETfGkk_)

## Links

* [Paper PDF](https://people.eecs.berkeley.edu/~brewer/cs262/LFS.pdf)
* [LFS in OSTEP](http://pages.cs.wisc.edu/~remzi/OSTEP/file-lfs.pdf)
* [LFS in CS 537 @ UW-Madison](http://pages.cs.wisc.edu/~shivaram/cs537-sp20-notes/lfs/cs537-lfs-notes.pdf)

{% file src="/files/-MPkIY8GLc3BWa0IuLET" %}
Prof. Anrea's notes on FFS and LFS
{% endfile %}


# Lottery Scheduling: Flexible Proportional-Share Resource Management

## One-line Summary

Multiple clients hold various number of tickets. A winning ticket is randomly selected, and the client who holds the ticket wins.

## Paper Structure Outline

1. Introduction
2. Lottery Scheduling
   1. Resource Rights
   2. Lotteries
3. Modular Resource Management
   1. Ticket Transfers
   2. Ticket Inflation
   3. Ticket Currencies
   4. Compensation Tickets
4. Implementation
   1. Random Numbers
   2. Lotteries
   3. Mach Kernel Interface
   4. Ticket Currencies
   5. Compensation Tickets
   6. Ticket Transfers
   7. User Interface
5. Experiments
   1. Fairness
   2. Flexible Control
   3. Client-Server Computation
   4. Multimedia Applications
   5. Load Insulation
   6. System Overhead
6. Managing Diverse Resources
   1. Synchronization Resources
   2. Space-Shared Resources
   3. Multiple Resources
7. Related Work
8. Conclusions

## Background & Motivation

Existing schedulers do not have accurate control over computational service rates, are poorly understood, and are difficult to control. Current systems are limited due to the assumptions and overheads associated with existing fair share schedulers. In this work, the authors present lottery scheduling, a randomized scheduler that implements proportional-share resource management. It also provides good support for modular resource management. The idea of a lottery scheduling mechanism is really abstract and can be applied to many problem areas.

## Design and Implementation

![It really is this simple!](/files/-MOiiGn3pN-y5vU0FquV)

How lottery scheduling works on a high level is really intuitive. The interesting stuff is some of the low-level, modular mechanisms listed below:

### Ticket Currencies

![](/files/-MOijXVDf465nXABBcRy)

Currencies provides abstraction barriers across logical truct boundaries.

![With currencies, the inflation is contained/insulated within a currency.](/files/-MOio4gSl_A31ya7wZ6J)

### Ticket Transfers

This is basically transferring tickets from one client to another client. It is useful when a client blocks due to some dependency. Some other examples are (1) when another process is doing work on another's behalf (2) when waiting for another process (holding locks).

### Ticket Inflation

It is an alternative to explicit ticket transfers. This sounds like a bad idea because some clients can monopolize a resource by creating a large number of lottery tickets, but it is extremely useful among mutually trusting clients as inflation and deflation allows resource allocations to be adjusted without explicit communication.

### Compensation Tickets

If a client only consumes a fraction f of its allocated resources, it can be granted a compensation ticket that inflates its value by 1/f until the client starts its next quantum. Without compensation tickets, a client that does not utilize all of its allocated quantum may receive less than its entitled share of the processor. As an example, compensation tickets can be used when a process is blocked for a short period of time.

## Evaluation

Some interesting graphs:

![A 2:! ticket allocation leads to an actual 2.01:1 runtime ratio. However, a drawback we can notice here is that allocation is very random and unstable (notice the unit of the x-axis)](/files/-MOin-dp2RKliV15vwOG)

## Links

* [Paper PDF](https://www.usenix.org/legacy/publications/library/proceedings/osdi/full_papers/waldspurger.pdf)
* [Lottery and Stride Scheduling on YouTube](https://youtu.be/qAx4IxrOoAM) by CS4414 @ U of Virginia




---

[Next Page](/llms-full.txt/1)

