Authors: Chaithan Prakash (University of Wisconsin-Madison), Jeongkeun Lee (HP Labs), Yoshio Turner (Banyan), Joon-Myung Kang (HP Labs), Aditya Akella (University of Wisconsin-Madison), Sujata Banerjee (HP Labs), Charles Clark (HP Labs), Yadi Ma (HP Labs), Puneet Sharma (HP Labs), Ying Zhang (HP Labs)
Presenter: Chaithan Prakash from HB Labs
Paper: http://conferences.sigcomm.org/sigcomm/2015/pdf/papers/p29.pdf
Public Review: http://conferences.sigcomm.org/sigcomm/2015/pdf/reviews/346pr.pdf
PGA stands for Policy Graph Abstraction
It aims at providing an intuitive & simple API in which network operators can express and combine network policies. PGA is close to the kind of graph-based representations network operators would draw on a white board.
Network policy management is a challenge:
Different types of policies by different stakeholders have to be combined into a single coherent, well-formed policy that satisfies the constraints imposed by the individual policies.
In practice, this "intent-composition" is usually performed manually by network operators.
-> not scalable, error prone, low-level
High level languages such as Pyretic, NetKAT, ... provide composition operators; but they are not sufficient (?) to compose high-level policies.
PGA aims to provide:
* Simple, inutuitive abstractions
* portable policies (decoupled from network specifics)
* automatic policy composition
Model:
Nodes represent groups of end-points (EPG) that satisfy a certain membership predicate
Directed edges represent allowed communication
Edge attributes: classifiers/service chains/network functions, for example what port, byte counting, through load balancer
Labels: capture endpoint attributes: tenant, location, security status (logical labels)
Several graphs are composed into single graph by composition algorithm.
The authors have implemented a prototype in Python. It provides a GUI to input graphs, and uses one of the following two backends to deploy the resulting policies:
* Pox controller in reactive mode -> OpenFlow
* OpenStack Neutron
The runtime overhead (for executing policies) is in the sub-ms area.
Compilation scales:
(a) <10 min, < 1.2 GB for a graph with 1 million edges, but no network functions
(b) <14mins, < 20 GB when randomly inserting network functions into the graph.
There will be a live demo of the prototype tomorrow.
Questions (incomplete & paraphrased)
Q1: Complex policies combination, how do you know it's right?
A: Intuitive framework; no absolute guarantees. System raises flag if well-formed composition is impossible.
Tuesday, August 18, 2015
Session 3.2, Paper 3 [Experience Track 2]: Jupiter Rising: A Decade of Clos Topologies and Centralized Control in Google’s Datacenter Network
Authors: Arjun Singh (Google), Joon Ong (Google), Amit Agarwal (Google), Glen Anderson (Google), Ashby Armistead (Google), Roy Bannon (Google), Seb Boving (Google), Gaurav Desai (Google), Bob Felderman (Google), Paulie Germano (Google), Anand Kanagala (Google), Jeff Provost (Google), Jason Simmons (Google), Eiichi Tanda (Google), Jim Wanderer (Google), Urs Hoelzle (Google), Stephen Stuart (Google), Amin Vahdat (Google)
Presenter: Arjun Singh
Link to the Public Review by Ming Zhang
Summary:
Q/A (Paraphrased):
Q. You used DCTCP. What did you do for virtualized environments where you didn't have control over the end-host stack
A. I'm not too sure but I can't provide more details.
Q. It seems we are relearning what we learnt 30 years ago. Is it that earlier we focussed on "within" the switch Clos-type interconnections and now we are considering these distributed datacenters structured as Clos networks with centralized software for managing them
A. Yes
Q. What gains did DCTCP bring to your datacenter?
A. I don't know the precise numbers but the gains were quite significant.
Q. We in academia do not have access to a lot of things that companies like Google have. Would you like to comment on that?
A. Well, that is a good question. We also struggled with large-scale evaluation. As a result, we had to rely on virtualized testbed environments for testing new protocols and systems.
Q. Do you still have congestion hotspots? Is there more thirst for intra-datacenter bandwidth?
A. Even though we removed most of the congestion hotspots over time, the need for bandwidth is constantly growing.
Presenter: Arjun Singh
Link to the Public Review by Ming Zhang
Summary:
Large-scale datacenters operated by companies like Google, Facebook, and Amazon support hundreds of thousands of servers in a single facility. It is often useful to ask: How did they address the key challenges of scalability, manageability, cost, and evolvability as their datacenter networks grew over time. This paper studies the evolution of Google’s datacenter network during the last decade.
Arjun highlighted the three common themes across the five generations of evolution of Google’s datacenter networks. They are 1) the use of Clos topologies for achieving scalable performance and failure resilience, 2) use of centralized protocols for managing operational complexity, and 3) use of cheap off-the-shelf commodity devices.
Arjun showed how Google started with firehose 1.0 that provided few Tbps of aggregate capacity in 2004 to Jupiter that supports up to 1.3 Pbps of capacity. He walked through how they used the three key themes in Google's datacenter design that enabled them to achieve high performance, high availability, and ease of network management at relatively low costs. Finally, he talked about managing small on-chip buffers by leveraging ECN and DCTCP and providing high reliability using redundancy and diversity.
Q/A (Paraphrased):
Q. You used DCTCP. What did you do for virtualized environments where you didn't have control over the end-host stack
A. I'm not too sure but I can't provide more details.
Q. It seems we are relearning what we learnt 30 years ago. Is it that earlier we focussed on "within" the switch Clos-type interconnections and now we are considering these distributed datacenters structured as Clos networks with centralized software for managing them
A. Yes
Q. What gains did DCTCP bring to your datacenter?
A. I don't know the precise numbers but the gains were quite significant.
Q. We in academia do not have access to a lot of things that companies like Google have. Would you like to comment on that?
A. Well, that is a good question. We also struggled with large-scale evaluation. As a result, we had to rely on virtualized testbed environments for testing new protocols and systems.
Q. Do you still have congestion hotspots? Is there more thirst for intra-datacenter bandwidth?
A. Even though we removed most of the congestion hotspots over time, the need for bandwidth is constantly growing.
Session 3.2, Paper 2 [Experience Track 2]: End-User Mapping: Next Generation Request Routing for Content Delivery
Authors: Fangfei Chen (Akamai Technologies), Ramesh K. Sitaraman (UMass, Amherst & Akamai Technologies), Marcelo Torres (Akamai Technologies)
Presenter: Fangfei Chen
Link to Public Review by Ethan Katz-Bassett
Summary:
Presenter: Fangfei Chen
Link to Public Review by Ethan Katz-Bassett
Summary:
A key component in a content delivery network (CDN) is the mapping system that routes a client’s request to a server. Traditionally, the mapping systems have used the local recursive domain name servers (LDNS) as a proxy for determining clients' path conditions (called NS-based mapping approach). However, in several cases the network conditions of the LDNS and the client may be different. This paper provides the first large-scale study of the limitations of NS-based mapping and the experience of deploying EDNS0 (client-subnet extension to the DNS protocol); a solution that uses the client’s prefix to infer information about the client (their resulting system is called end-user mapping).
They collected data from over 3.76 million IP blocks. They found that more than 50% of the clients were within 100miles of their LDNS. However, there was a lot of diversity across countries. For instance, the average distance between clients and LDNS in India and Brazil was large. The author conjectured that this was because there weren't many (or perhaps no) public DNS servers within these countries.
Fangfei then talked about the performance benefits of using end-user mapping as observed by Akamai. In particular, they found that the download times improved by approximately 100ms. This is significant because better download times for the end-users are correlated with revenue. However, a major challenge they ran into was scalability. As there are more clients than LDNS servers generally, using EDNS0 meant that there was increase in DNS query traffic. It is a significant problem which requires more attention.
They collected data from over 3.76 million IP blocks. They found that more than 50% of the clients were within 100miles of their LDNS. However, there was a lot of diversity across countries. For instance, the average distance between clients and LDNS in India and Brazil was large. The author conjectured that this was because there weren't many (or perhaps no) public DNS servers within these countries.
Fangfei then talked about the performance benefits of using end-user mapping as observed by Akamai. In particular, they found that the download times improved by approximately 100ms. This is significant because better download times for the end-users are correlated with revenue. However, a major challenge they ran into was scalability. As there are more clients than LDNS servers generally, using EDNS0 meant that there was increase in DNS query traffic. It is a significant problem which requires more attention.
Q/A (Paraphrased):
Q. Why did you report results in terms of miles as opposed to RTTs
A. It was tricky to report them in terms of RTTs. Reporting them in miles was more direct and also helped in geo-locating the IP addresses
Q. Did you look at the costs at the ISP side of scalability concerns
A. We didn't look at the costs but such an analysis could be carried out with our data.
Q. Did you measure EDNS0 adoption?
A. No but we are encouraging ISPs to adopt it.
Q. Was there a rollout of new cache servers while conducting this study?
A. We did not specifically restrict the rollout of new cache servers.
Session 3.2, Paper 1 [Experience Track 2]: Large-scale measurements of wireless network behavior
Authors: Sanjit Biswas (Cisco Meraki), John Bicket (Cisco Meraki), Edmund Wong (Cisco Meraki), Raluca Musaloiu-E (Cisco Meraki), Apurv Bhartia (Cisco Meraki), Dan Aguayo (Cisco Meraki)
Presenter: Sanjit Biswas
Link to the Public Review by Kyle Jamieson
Summary:
In the end, Sanjit mentioned about one AP that had 10,000 nearby networks as well as about cable problems they frequently ran into. It is great that they have made their data publically available.
Q/A (Paraphrased):
Q. Most of Meraki's clients are enterprises. Do you expect the Netflix traffic to dominate for home networks or dorms?
A. Yes
Q. What was the latency between the radios and the collection server?
A. This was in the order of milliseconds.
Presenter: Sanjit Biswas
Link to the Public Review by Kyle Jamieson
Summary:
How big of an issue is interference in deployed WiFi networks? Which applications are most popular? How fast is the uptake of newer WiFi standards (e.g., 802.11n and 802.11ac)? How utilized are the WiFi frequency bands and how is this changing over time? These are some of the exciting questions that this paper attempts to answer by analyzing an extremely large-scale measurement dataset from Meraki’s deployed WiFi networks.
Using Meraki's centralized management platform, they collected time-series data of applications, clients, and device statistics. This data was collected from over 20,000 networks and about 5.58M clients across a variety of deployment types.
Sanjit highlighted some interesting insights during his talk. These included:
- YouTube, NetFlix, and iTunes combined (i.e, video/music) dominated the traffic and over a period of one year, all these applications observed at least 60% growth in their traffic
- Most data was downloaded by users using the Windows/MS OSes and smartphone traffic witnessed one of the largest growth in terms of traffic as well as number of devices.
- While overwhelming majority of clients used 802.11n capable devices, most only used the single stream mode (these were largely smartphones equipped with single antennas). In addition, 802.11ac was observed to be gaining traction.
- They found most interference to be from 802.11 traffic. In addition, they found that majority of the wireless links exhibited a wide range of delivery ratios. On average, they observed ~55 networks/AP
In the end, Sanjit mentioned about one AP that had 10,000 nearby networks as well as about cable problems they frequently ran into. It is great that they have made their data publically available.
Q/A (Paraphrased):
Q. Most of Meraki's clients are enterprises. Do you expect the Netflix traffic to dominate for home networks or dorms?
A. Yes
Q. What was the latency between the radios and the collection server?
A. This was in the order of milliseconds.
Q. Why was the channel utilization so low?
A. We actually think, it is fairly high.
A. We actually think, it is fairly high.
[Session 3.1, Experience track1] Pingmesh: A Large-Scale System for Data Center Network Latency Measurement and Analysis
Pingmesh: A Large-Scale System for Data Center Network Latency Measurement and Analysis
Authors: Chuanxiong Guo, Lihua Yuan, Dong Xiang, Yingnong Dang, Ray Huang, Dave Maltz, Zhaoyi Liu, Vin Wang, Bin Pang, Hua Chen, Zhi-Wei Lin, Varugis Kurien*
Microsoft, *Midfin systems
Presenter: Chuanxiong Guo
This work focuses on network measurement and analysis. It considers the data center network (DCN), and tries to tackle challenges in DCN operation, such as the source of incidents (network or not), tracking SLA etc. The work suggests that using latency measurements between any two servers these challenges can be addressed.
The system, dubbed Pingmesh, uses Pingmesh agents installed on each server. The agents are controlled by a Pingmesh controller, which defines and manages the measurements. The system is built on top of existing infrastructure (Autopilot, Cosmos/SCOPE). The storage and analysis is done on a scale that ranges from 10 minutes to a day.
The rest of the talk focused on the results collected and the lessons learned.
It started by exploring latency measurements, while comparing two data centers.
A following packet drop rate study has shown packet drops at the NIC and the ToR, and compared it between different data centers, showing that they differ (for intra and inter rack drop), and that a typical drop rate is in the order of 10^-4 to 10^-5.
Using Pingmesh, it was possible to detect black holes (deterministic packet drops) and silent random packet drops.
Pingmesh is not without limitations. Pingmesh can not tell which spine switch is dropping packets. Also, Pingmesh uses single packets for the measurements, which is not good for detecting network reachability and packet-level latency issues.
Pingmesh provides an always-on service that provides a full coverage. It provides scalability, leaving space for evolvability. As such, it is an interesting tool for tracking problems in the network, which is one of the hard problem today. Plus, it is running on a production system and for a long time (4 years), which is its main advantage.
Q&A:
Q: What are the thresholds beyond which the overhead of measurements is too high?
A: Pingmesh has a lot of overheads, that relate to processing, but in terms on networking pingmesh uses just Kbps, vs. Gbps available on the servers, so it is negligible.
Q: In the past, works like Planerlab used similar measurement techniques to predict how systems will behave. Will you be able to do what/if based on the huge amounts of data you collected?
A: We are currently focused on using the system for trouble shooting.
Q: Do you really need 2M probes to find that a server went down?
A: You can not predict what will happen next, so we try to make the measurements as comprehensible as possible, so we can catch failures quickly.
Session 3.1 - Paper 1: Inside the Social Network's (Datacenter) Network
Inside the Social Network's (Datacenter) Network
Authors: Arjun Roy (University of California, San Diego), Hongyi Zeng (Facebook), Jasmeet Bagga (Facebook), George M. Porter (University of California, San Diego), Alex C. Snoeren (University of California, San Diego)
Presenter: Arjun Roy
Summary:
This paper presents the production traffic workload characteristics of the Facebook datacenters and how the workload differs from the existing literature. Existing workload observations are mainly based on web search services, which does not capture all types of cloud services. Literature of datacenter traffic shows that: traffic is rack local and predictable over small timescales. The ability to design efficient deployments in datacenters by means of traffic workload knowledge motivates this work. Facebook datacenter architecture consists of multiple site, multiple datacenter within site and a fat tree topology consisting of clusters of servers connected together by top-of-rack switch, in turn connected to cluster switches, which are connected by Fat Cats aggregation switches. Servers have exclusive roles: cache, web, Hadoop and MultiFeed servers; rack typically contains same type of servers. Workloads are collected using Facebook-wide monitoring system using Fbflow(long term traffic storage, but of low resolution) and per-host packet-header traces using Port mirroring(short term traffic of higher resolution).
Following are the observations from traffic analysis:
- Although Hadoop deployments traffic is consistent with literature, other traffic is neither rack-local nor all-to-all(sparse rack locality and heavy inter-rack traffic observed in Front-end clusters).
- Despite this inter-rack traffic, most of the links are loaded with less than 10% traffic. Heavy hitters are generally instantaneous and are not persistent. Sub-second traffic is unpredictable and ephemeral.
- Traffic is bipartite due to colocation of servers of same type.
- Traffic is stable across all cache servers over long timescales(1-second timescale). Traffic distributions exhibit even spread across different levels(cluster, dataceenter, rack levels and so on). Hence traffic is stable over time and per destination.
These observations provide a chance to reduce inter-rack bandwidth.
Q&A:
Q: Any insights on why the heavy hitters didn't last for long, is it due congestion or some application the end hosts are doing?
A: Heavy hitters are ephemeral due to effect of load balancing. If heavy hitters are persistent, then that means the load balancer is not doing its job. Load balancer ensures that the cache center(which got heavy hit) does not get further requests for a significant amount of time to balance load properly.
Q: Was coordination between same type of service exhibiting locality similar to intra-rack locality?
A: We did look at locality. Since services are colocated at rack level, looking at rack level is exactly same as looking at type of service.
Q: Instead of blind colocation of servers of same type, why not mix servers together?
A: Locality patterns wouldn't work out well by doing colocation of different servers.
Q: Are aggregation and core switches as volatile with respect to heavy hitters as end hosts are?
A: This study did not cover this.
Q: Are incast traffic issues observed in the FB datacenters similar to observations in DCTCP?
A: Incast traffic require micro-second scale observations, which weren't part of this study.
Session 1, Paper 4: Central Control Over Distributed Routing
Authors: Stefano Vissicchio (UCLouvain), Olivier Tilmans (UCLouvain), Laurent Vanbever (ETH Zürich), Jennifer Rexford (Princeton University)
Presenter: Stefano Vissicchio
This paper presents Fibbing, an architecture that achieves both flexibility and robustness through central control over distributed routing. Fibbing relies on traditional link-state protocol like OSPF and IS-IS, and by introducing fake nodes and links, Fibbing is able to control forwarding information base (FIB) directly.
The most exciting thing about this paper is, as mentioned in the public review, “The paper reminds us that SDN is not architecturally about a particular wire protocol but about decoupling
the control and data plane, and how in fact there are multiple ways, with varying trade-offs, to achieve it.”
System Design is clearly described in the graph below borrowed from the paper:
A Fibbing prototype is evaluated over three dimensions,
- low CPU and memory overhead and no impact on convergence time introduced by Fibbing on existing routers,
- the efficiency of Fibbing’s augmentation algorithm in terms of speed and size of topology
- a case study on a real network consisting of 4 Cisco routers showing how Fibbing alleviate congestion.
Q&A:
Q: Do you ever get to trouble if working with different switches from different vendors?
A: No, we use standard IGP, and there is no fancy specific features used.
Q: Given the complexity of network dynamics, I can't tell if this is a bad idea or a brilliant idea. How do you debug Fibbing when things go wrong?
A: It's true that Fibbing is hard to roll-back. But Fibbing controller could automate the roll-back. Controller can push some information to improve debugging ability. IGP has a shared view of topology, and it will sync up. We don't think debugging ability is a problem.
- Follow-up Q: If the controller crashes, will router fall-back to standard?
- Follow-up A: Yes.
Q: OSPF is hard to enforce high-level policy. Will things be easy if use BGP high-level policy routing?
A: We can apply a modified version of Fibbing to BGP. Different protocols have values in different settings.
Q: How does Fibbing work with OSPF hierarchy?
A: We need one controller per zone, and coordination is needed. Maybe in future work.
Q: Does Fibbing complicate debugging?
A: To me, Fibbing is not degrading debugging too much. You just need to understand why the fake nodes are there.
Q: We can apply either segment routing or Fibbing to do policy routing. Which to choose?
A: In some cases, they are complimentary. The solution of Fibbing is simpler, because it needs less information from the hardware. If just used for optimizing traffic, Fibbing might be better.
Q: SDN also gives you fine-grained control, but I don't see it in Fibbing.
A: It's true. Fibbing is not very sweet for fine-grained control, but it still provides you some level fine-grained control, such as middle-box. One possibility is to deploy Fibbing in the core instead of edge.
Q: How is implementation complexity in practice?
A: There is one special field in OSPF message, which is leveraged to enable Fibbing. And it works with real Cisco and Juniper routers.
Q: Do you ever get to trouble if working with different switches from different vendors?
A: No, we use standard IGP, and there is no fancy specific features used.
Q: Given the complexity of network dynamics, I can't tell if this is a bad idea or a brilliant idea. How do you debug Fibbing when things go wrong?
A: It's true that Fibbing is hard to roll-back. But Fibbing controller could automate the roll-back. Controller can push some information to improve debugging ability. IGP has a shared view of topology, and it will sync up. We don't think debugging ability is a problem.
- Follow-up Q: If the controller crashes, will router fall-back to standard?
- Follow-up A: Yes.
Q: OSPF is hard to enforce high-level policy. Will things be easy if use BGP high-level policy routing?
A: We can apply a modified version of Fibbing to BGP. Different protocols have values in different settings.
Q: How does Fibbing work with OSPF hierarchy?
A: We need one controller per zone, and coordination is needed. Maybe in future work.
Q: Does Fibbing complicate debugging?
A: To me, Fibbing is not degrading debugging too much. You just need to understand why the fake nodes are there.
Q: We can apply either segment routing or Fibbing to do policy routing. Which to choose?
A: In some cases, they are complimentary. The solution of Fibbing is simpler, because it needs less information from the hardware. If just used for optimizing traffic, Fibbing might be better.
Q: SDN also gives you fine-grained control, but I don't see it in Fibbing.
A: It's true. Fibbing is not very sweet for fine-grained control, but it still provides you some level fine-grained control, such as middle-box. One possibility is to deploy Fibbing in the core instead of edge.
Q: How is implementation complexity in practice?
A: There is one special field in OSPF message, which is leveraged to enable Fibbing. And it works with real Cisco and Juniper routers.
How to Bid the Cloud
Authors: Liang Zheng, Carlee Joe-Wong (Princeton University), Chee Wei Tan (City University of Hong Kong & National University of Singapore), Mung Chiang (Princeton University), Xinyu Wang (City University of Hong Kong)
Paper (pdf)
Public review (pdf)
This paper studies the auction-based spot pricing of Amazon's Elastic Compute Cloud. Users can place bids (above a price set by Amazon) to have computation performed on EC2; jobs that are currently running can be outbid and interrupted. Hence, EC2 spots are ideal for jobs that can be suspended and the results of which are not needed immediately. The goal of this paper was twofold, to understand how the cloud provider sets their prices and to determine which prices users should bid.
The authors present two different types of bidding strategies, one-time bids and persistent bids. One-time bids are submitted once and then exit the system once they fall below the current spot price. The risk with one-time bids is that they be interrupted without completing. Persistent bids are resubmitted in each time period until the job finishes or is manually terminated by the user. This results in longer waiting and completion times. In contrast, one-time bids provide better control over bid completion times. The authors test these strategies to bid for MapReduce jobs. They propose placing a single one-time bid for the Master node, which prevents interruptions and using persistent bidding requests for slave nodes.
Q&A
Q: You talk about both a cloud provider price model and user price model. What if you cannot control the provider's price model? What if you have no knowledge about how Amazon prices? Will your model work?
A: We tried to model how the cloud provider sets a price to get more precise prediction of the spot price. If we do not have any knowledge of how Amazon, then we can still use probability distribution of spot price offered by Amazon to predict future prices. As a result, maybe it will not be as accurate, but it is still doable.
Q: Two questions. Do you provide any estimation or prediction models to the user that is currently bidding? In other words, apart from the current minimum bid, what other information is publicly available? Also in your map reduce model you can reduce your cost by taking some slaves online, but isn't this cost offset by the increased need for storage when the host goes offline?
A: To answer your second questions, for map reduce jobs, if some slave node goes offline, I think the slave nodes will upload some result to the master node. But when they jump off then there may be some results they haven't uploaded to the master node. This will be included in the recovery time. If another slave node takes over for the slave node that fails, then we implement this as a recovery/overhead time in our work. The other slave node will need more time to complete. Can you repeat your first question?
Q: Say I am a bidder, and I want to bid for some containers. I do know the current minimum bid price, but what other information is available to me? Is there an estimation model? For example, what is the probability my job is killed at some time for a given bid? Do you provide information on what the probability of a job being killed is if you bid some price. Say, if I bid 500 dollars I have 60% probability of being killed, if I bid 800 dollars I have 30% of my job being killed.
A: Yes, this information can be calculated by our model. The public information that is provided to users is spot instance prices of the past two months. From this, we can calculate the PDF of the spot price. Then we can calculate the probability by considering the total running time of the job.
Paper (pdf)
Public review (pdf)
This paper studies the auction-based spot pricing of Amazon's Elastic Compute Cloud. Users can place bids (above a price set by Amazon) to have computation performed on EC2; jobs that are currently running can be outbid and interrupted. Hence, EC2 spots are ideal for jobs that can be suspended and the results of which are not needed immediately. The goal of this paper was twofold, to understand how the cloud provider sets their prices and to determine which prices users should bid.
The authors present two different types of bidding strategies, one-time bids and persistent bids. One-time bids are submitted once and then exit the system once they fall below the current spot price. The risk with one-time bids is that they be interrupted without completing. Persistent bids are resubmitted in each time period until the job finishes or is manually terminated by the user. This results in longer waiting and completion times. In contrast, one-time bids provide better control over bid completion times. The authors test these strategies to bid for MapReduce jobs. They propose placing a single one-time bid for the Master node, which prevents interruptions and using persistent bidding requests for slave nodes.
Q&A
Q: You talk about both a cloud provider price model and user price model. What if you cannot control the provider's price model? What if you have no knowledge about how Amazon prices? Will your model work?
A: We tried to model how the cloud provider sets a price to get more precise prediction of the spot price. If we do not have any knowledge of how Amazon, then we can still use probability distribution of spot price offered by Amazon to predict future prices. As a result, maybe it will not be as accurate, but it is still doable.
Q: Two questions. Do you provide any estimation or prediction models to the user that is currently bidding? In other words, apart from the current minimum bid, what other information is publicly available? Also in your map reduce model you can reduce your cost by taking some slaves online, but isn't this cost offset by the increased need for storage when the host goes offline?
A: To answer your second questions, for map reduce jobs, if some slave node goes offline, I think the slave nodes will upload some result to the master node. But when they jump off then there may be some results they haven't uploaded to the master node. This will be included in the recovery time. If another slave node takes over for the slave node that fails, then we implement this as a recovery/overhead time in our work. The other slave node will need more time to complete. Can you repeat your first question?
Q: Say I am a bidder, and I want to bid for some containers. I do know the current minimum bid price, but what other information is available to me? Is there an estimation model? For example, what is the probability my job is killed at some time for a given bid? Do you provide information on what the probability of a job being killed is if you bid some price. Say, if I bid 500 dollars I have 60% probability of being killed, if I bid 800 dollars I have 30% of my job being killed.
A: Yes, this information can be calculated by our model. The public information that is provided to users is spot instance prices of the past two months. From this, we can calculate the PDF of the spot price. Then we can calculate the probability by considering the total running time of the job.
Q: There's an assumption that the computation results don't vary in value. For example, the computation results are not going to be useful if the conference is already over. Do you have any thoughts on what affect the value of computation over time would have on the model? Often, the busiest times to bid are more expensive because the results of the computation are more valuable at that time. A real life bidding strategy would have to take that into account the effects of running the job later during the off hours. Do you have any thoughts on how that would affect the bidding strategy?
A: Actually we did not find any correlation between the time and the spot price, because people from all over the world can bid at any time. But we do think that if a lot of people use our bidding strategies then this would affect the spot price. It should make the market efficient. We want to improve this in future work.
A: Actually we did not find any correlation between the time and the spot price, because people from all over the world can bid at any time. But we do think that if a lot of people use our bidding strategies then this would affect the spot price. It should make the market efficient. We want to improve this in future work.
Q: How will things change if you consider competitors? That if customers had the choice to select between different providers?
A: This is a good question which we never considered. If we have multiple cloud providers, then the user can have more choices to choose from by comparing the price they offered and the spot price distribution may be more variable. This will complicated model. I think that is an interesting point to study in the future.
A: This is a good question which we never considered. If we have multiple cloud providers, then the user can have more choices to choose from by comparing the price they offered and the spot price distribution may be more variable. This will complicated model. I think that is an interesting point to study in the future.
Poptrie: A Compressed Trie with Population Count for Fast and Scalable Software IP Routing Table Lookup
Authors: Hirochika Asai (The University of Tokyo), Yasuhiro Ohara (NTT Communications Corporation)
Paper (pdf)
Public review (pdf)
With the rapidly increasing size of routing tables, we need an efficient approach for IP routing table lookups. In this paper, the authors Poptrie, an IP routing lookup algorithm. The name was chosen because of how their algorithm leverages the popcount instruction and their use of the multiway trie. The authors were able to achieve a lookup rate of 148.8 Mlps (about 6.7 ns per lookup).
Poptrie uses a number of extensions to improve performance and reduce the memory footprint, such as compressing the leaf bit-vector, route aggregation, and direct pointing (like DXR). Using the leafvec extension, they were able to reduce the memory footprint to less than 1/3. Adding the route aggregation reduces the tree depth. Direct pointing extracts the nodes that corresponds to the most significant s bits, which specifies how many bits should be used as the index to the lookup array.
The authors tested Poptrie on both random and real traffic and compared against other algorithms. In their tests, D18R performed well on random traffic, but did poorly on on real traffic. Conversely, SAIL did poorly on random traffic, but performs well on real traffic. In contrast to both, Poptrie performed well on both random and real traffic patterns. The authors also pointed out that Poptrie is well suited for the future as the number of IPv6 prefixes increases, since it does well on datasets with longer prefixes.
Q&A
Q: Compared to work in the past, the most impressive thing here seems to be that you managed to compress the intermediate nodes and pointers. It seems like the big win is the deletion of all the duplicate space. Yes?
A: Yes. The direct pointing is significant for this. We couldn't see any big differences at shorter prefixes, but we could see very big differences at the deeper prefixes.
Q: Are there any other places such as firewall rules or encryption lookups that would also benefit that would benefit from these kind techniques for deletion of duplication.
A: Yes, I think so. Right now, we are focusing on IP routing table, but could extend this in future work to apply to access control lists, firewalls or string matching.
Q: You're using the pop count instruction in a general purpose processor, it's a standard instruction. You're benefiting a lot from that instruction. Normally, if you need high speed lookups you'd probably use ASICs or FPGAs perhaps. So, couldn't you put more things in the hardware?
A: Yes, this can implemented in hardware, but we focus on the software. Because with the advancement of hypervisors and virtual machine technologies, routing is often implemented in a hypervisor. In that case, we are not using dedicated hardware. This mechanism could be implemented in hardware but our focus is on software.
Q: I wonder if you had tried using memory prefetching to reduce the number of cycles for lookups? It should be able to reduce the number of cycles that are caused by cache misses.
A: The cache misses are the main factor that affects performance. Comparing 5th and 95th percentiles, we can see there are big fluctuations. It's very fast if there are no cache misses, but if there are cache misses we see bad performance. We looked into trying to quantify the number of CPU cache misses, but the kind of CPU we used does not report that. We want to look into that in the future. So instead we focused on measuring CPU cycles. We did try prefetching, but it doesn't help. We only have 6-7 nanoseconds. Since the prefetching requires more CPU cycles and time, it does not work well.
Paper (pdf)
Public review (pdf)
With the rapidly increasing size of routing tables, we need an efficient approach for IP routing table lookups. In this paper, the authors Poptrie, an IP routing lookup algorithm. The name was chosen because of how their algorithm leverages the popcount instruction and their use of the multiway trie. The authors were able to achieve a lookup rate of 148.8 Mlps (about 6.7 ns per lookup).
Poptrie uses a number of extensions to improve performance and reduce the memory footprint, such as compressing the leaf bit-vector, route aggregation, and direct pointing (like DXR). Using the leafvec extension, they were able to reduce the memory footprint to less than 1/3. Adding the route aggregation reduces the tree depth. Direct pointing extracts the nodes that corresponds to the most significant s bits, which specifies how many bits should be used as the index to the lookup array.
The authors tested Poptrie on both random and real traffic and compared against other algorithms. In their tests, D18R performed well on random traffic, but did poorly on on real traffic. Conversely, SAIL did poorly on random traffic, but performs well on real traffic. In contrast to both, Poptrie performed well on both random and real traffic patterns. The authors also pointed out that Poptrie is well suited for the future as the number of IPv6 prefixes increases, since it does well on datasets with longer prefixes.
Q&A
Q: Compared to work in the past, the most impressive thing here seems to be that you managed to compress the intermediate nodes and pointers. It seems like the big win is the deletion of all the duplicate space. Yes?
A: Yes. The direct pointing is significant for this. We couldn't see any big differences at shorter prefixes, but we could see very big differences at the deeper prefixes.
Q: Are there any other places such as firewall rules or encryption lookups that would also benefit that would benefit from these kind techniques for deletion of duplication.
A: Yes, I think so. Right now, we are focusing on IP routing table, but could extend this in future work to apply to access control lists, firewalls or string matching.
Q: You're using the pop count instruction in a general purpose processor, it's a standard instruction. You're benefiting a lot from that instruction. Normally, if you need high speed lookups you'd probably use ASICs or FPGAs perhaps. So, couldn't you put more things in the hardware?
A: Yes, this can implemented in hardware, but we focus on the software. Because with the advancement of hypervisors and virtual machine technologies, routing is often implemented in a hypervisor. In that case, we are not using dedicated hardware. This mechanism could be implemented in hardware but our focus is on software.
Q: I wonder if you had tried using memory prefetching to reduce the number of cycles for lookups? It should be able to reduce the number of cycles that are caused by cache misses.
A: The cache misses are the main factor that affects performance. Comparing 5th and 95th percentiles, we can see there are big fluctuations. It's very fast if there are no cache misses, but if there are cache misses we see bad performance. We looked into trying to quantify the number of CPU cache misses, but the kind of CPU we used does not report that. We want to look into that in the future. So instead we focused on measuring CPU cycles. We did try prefetching, but it doesn't help. We only have 6-7 nanoseconds. Since the prefetching requires more CPU cycles and time, it does not work well.
Session 1, Paper 2: A Declarative and Expressive Approach to Control Forwarding Paths in Carrier-Grade Networks
Authors: Renaud Hartert (UCLouvain), Stefano Vissicchio (UCLouvain), Pierre Schaus (UCLouvain), Olivier Bonaventure (UCLouvain), Clarence Filsfils (Cisco Systems Inc), Thomas Telkamp (Cisco Systems Inc), Pierre Francois (IMDEA Networks Institute)
Presenter: Stefano Vissicchio
This paper introduces a two-layer architecture in control plane to achieve scalable traffic engineering in carrier-grade networks. The solution supports declarativity and expressiveness by doing two things:
- A centralized optimizer, DEFO, which translates high-level goals into optimized paths
- A routing model called Middlepoint Routing (MR), which is less constrained than shortest-path routing and natively supports multi-path routing (contrary to end-to-end tunneling).
The two-layer architecture are optimization layer on top of connectivity layer. At optimization layer, it provides a domain specific language which enables network operators to declaratively define traffic engineering targets and constraints, and translates them into optimized paths. At connectivity layer, the configurations are initially done by the operators and later overwritten by optimized ones if generated by upper layer. DEFO computes paths based on Middlepoint Routing model, and midpoint selection is proved to be NP-hard. The authors propose a heuristic approach to compute solutions.
The evaluation is done on multiple topologies (including real-world ISP topologies, topologies from Rocketfuel project and synthetic ones) and multiple demand matrices. The result shows that DEFO is more efficient than incremental IGP-WO, and DEFO has low er optimization overhead compared to RSVP-TE.
To sum up, this paper is interesting in two aspects:
- declarativity and expressiveness achieved by a high-level domain specific language
- leveraging segment routing model enabling scalable traffic engineering in carrier-grade networks
Q: Do you allow subgraph to be shaped across different paths?
A: Yes. We have abstraction of demands, for each demand we have different graphs.
- Follow-up Q: What if there are nodes overlapped between top and bottom?
- Follow-up A: We can do that. It is not a strong constraint.
Q: About choice of optimization, traffic engineering (TE) is normally considered a linear programing (LP) problem, so why use constraint programming (CP)?
A: If you want to support different TE objectives, you need more structured approach. With CP we can support flexibility of our language in an easier way. We are not claiming that CP is better than LP. Another advantage with CP is that we can budget time.
- Follow-up Q: How do you cut off at 2 minutes?
- Follow-up A: It is generally fast with CP to compute a solution.
Q: 2 middlepoints are used in the paper, what about the performance and complexity when using more than 2?
A: 3 middlepoints do not bring much gain. More complexity for sure, but not sure if it will cause problems to the system.
Keynote - SDN for the cloud
SDN for the Cloud - Albert Greenberg
The importance of SDNs has gradually improved in recent years. The demand for increased flexibility and room for innovation in route control and traffic engineering has led to rise of centralized control network management to meet overall network wide goals. As a result of such innovations, the emergence of cloud has led to the rise of software-defined data center networking (where each customer can gain control of entire virtual networks).
Some work in this area has led to VL2 (SIGCOMM 2009). Cloud has the right scenarios such as huge leverage, huge degree of control to make changes in the right places and high scale fault tolerant distributed systems and data management.
At Azure, the team started from scratch as they updated the hardware from optics to host to NIC to physical fabric to WAN to Edge/CDN to ExpressRoute.
Most of the concepts presented in the papers have been put into practice in Microsoft cloud infrastructures. As a result of these improvements, modern Azure services can carry up to 1,400,000 SQL databases. Moreover, a typical Azure event hub sees as high as 1 trillion events per month. This clearly shows the level of scalability an Azure service can handle as a result of the ideas that have successfully been transferred from their research work.
He then discussed the following agenda:
**Cloud Design Principles
- Scale-out N-Active data plane
- Embrace and isolate failures
- Centralized control plane: drive network to target state
**Hyperscale Physical Networks
- VL2 -> Azure close fabrics with 40G NICS
- Outcome of > 10 years of history, with major revisions every six-months
- Scale-out
- Challenges of Scaling out
- Clos network management problem
- Solutions
- Infrastructure for graph and state tracking to provide an app platform
- Monitoring to drive out gray failures
- Azure Cloud Switch OS to manage the switches as we do servers
**Azure SmartNIC
- Use an FPGA for reconfigurable functions
- Programmed using Generic Flow Tables (GFT)
- SmartNIC can also do Crypto, Qos, storage acceleration, and more...
**Closing Thoughts
- Cloud scale, financial pressure unblocked SDN
- SDN realized through consistent application of principles of cloud design
- Microsoft Azure re-imagined networking, Created SDN and it paid off
Opening and The Award Reciepants
** The Opening SIGCOMM 2015 **
1. Overview of SIGCOMM 2015
1. Overview of SIGCOMM 2015
- 42 full-length papers
- Keynote by SIGCOMM award winner Albert Greenberg
- 4 Workshops and 4 Tutorials
- Workshops : All Things Cellular, C2B(I)D, HotMiddlebox, NS Ethics
- Tutorials: Open Hardware Networking, Network Verification, Cloud Storage, P4
- 29 Posters and 19 Demos + Industrial Demos
2. Socials
- Student dinner at the Lord's ground (today)
- Conference Banquet at the brewery (tomorrow)
- Breakfast, lunches, coffee breaks etc.
3. New this year
- Mentoring Moments
- Topic Preview Sessions
- All the background you need to have to understand on upcoming paper session - in 15 minutes
- Mentoring & Previews
- A chance for PhDs and Post-docs to meet with faculty outside their university
- Often during lunch
4. PC Chairs
- Goal: an excellent program
- Experience Track
- 49 Member PC
- Chosen for technical excellence, review quality, topic coverage, diversity
- Emphasis throughout o seeing submissions strengths and potential
5. Reviewing
- 242 papers received 3 first-round review
- 138 papers advanced to second round
- 37 outside expert reviewers
- Papers discussed online at each round's end
- PC discussed 74 papers at pc meeting
- PC and outside experts wrote over 110 reviews!
6. Best Paper Award
- Decided at conference this year
- Allows committee to consider community's reception of work
- Winner(s) announced at 8:45 am on third day of conference.
** The Sigcomm 2015 Awards **
Award Chair : Bruce maggs
- Test of time award
- Title : A Clean Slate 4D Approach to Network Control and Management - ACM SIGCOMM computer communication review V35 N5 (2005), pp. 41-54
- Winners : Albert Greenberg, Gisli Hjalmtysson, David A. Mltz, Andy Myers, Jennifer Rexford, Geoffrey Xie, Hong Yan, Jibin Zhan, and Hui Zhang
- Title : Sizing Router Buffers, Proc, SIGCOMM 2005
- Winners : Guido Appenzeller, Isaac Keslassy, and Nick McKeown
- SIGCOMM Dissertation Award
- Title : Transport Architectures for an Evolving Internet
- Winner : Keith Winstein (Advisor Hari Balakrishnan, MIT)
- SIGCOMM lifetime achievement Award
- Title : For pioneering the theory and practice of operating carrier and datacenter networks
- Winner : Albert Greenberg
SIGCOMM 2015 kicks off in London
Welcome to England, where Britannia rules the (radio?) waves, and SIGCOMM 2015 is kicking off this morning at Imperial College London. The conference is being live-streamed, and Layer 9 is hosting a backchannel discussion on Slack. (To join the discussion, please enter your email address above, or visit the Slack signup page.)
As the conference unfolds, the scribes will be live-blogging every session and posting the scribe notes here. We welcome commentary and other contributions from anybody; if you would like to post on layer9.org and do not already have a posting account, please email <signup at layer9 dot org> to request one.
As the conference unfolds, the scribes will be live-blogging every session and posting the scribe notes here. We welcome commentary and other contributions from anybody; if you would like to post on layer9.org and do not already have a posting account, please email <signup at layer9 dot org> to request one.
Wednesday, October 29, 2014
Session 3, Paper 1: Reclaiming the Brain: Useful OpenFlow Functions in the Data Plane
Authors: Michael Borokhovich (Ben Gurion University, Israel), Liron Schiff (Tel Aviv University, Israel), Stefan Schmid (TU Berlin and T-Labs, Germany)
Link to paper: Reclaiming the Brain: Useful OpenFlow Functions in the Data Plane
SDN simplifies network management by providing programmatic interface to a logically centralized controller, allowing splitting network into a “dumb” data plane and a “smart” control plane. This, however, comes with a cost: fine-grain control of the data plane would introduce computational overhead and latency in the control plane. This paper investigates which functionalities to add in the OpenFlow data plane (“south”), making it smarter to reduce interactions with the control plane and network more robust.
SDN simplifies network management by providing programmatic interface to a logically centralized controller, allowing splitting network into a “dumb” data plane and a “smart” control plane. This, however, comes with a cost: fine-grain control of the data plane would introduce computational overhead and latency in the control plane. This paper investigates which functionalities to add in the OpenFlow data plane (“south”), making it smarter to reduce interactions with the control plane and network more robust.
The approach in this paper relies on a simple template called “SmartSouth”, an in-band graph DFS traversal implemented using the match-action paradigm, using fast failover technique. In general, monitoring and communication functions are added to the “south” to make it more robust, proactively react to link failures and reduce interaction with control plane. Specifically, functions which are provided in the south:
- Topology snapshot: collects current view of network topology; fault-tolerant, no connectivity assumed; single connection to controller is required
- Blackhole detection: detects connectivity loss, regardless of the causes (e.g., physical failure, configuration errors, unsupervised carrier network errors). Two implementations are proposed:
- multiple DFS traversals, each with different time-to-live TTL, using binary search to find the point where packet is lost. Complexity: log n
- smart in-band counter: counter is read and updated during packet processing, and counter value can be written to packet header field; proactively install one smart counter per switch port; two DFS traversals needed: first traversal will go back and forth once on new link, second traversal will detect the blackhole (link with counter of value 1).
- Critical node detection: check if a node is critical for connectivity; non-critical node may be removed for maintenance and energy conservation; cheaper than snapshot; only one DFS traversal with root.
- Anycast: supports specification of multiple unknown destinations; it is extendable to specify service chains; useful to find an alternative path to the controller when link fails. Complexity: one DFS traversal
- no new hardware or protocol features are required
- keep states formally verifiable
- some techniques are possibly extendable to other functions (e.g., using smart counter to infer network load).
This work serves as the first step and more discussions on how to partition functionalities between data plane and control plane are encouraged.
HotNets 2014: Infrastructure Mobility: A What-if Analysis
Scribe notes by Athina Markopoulou.
Infrastructure Mobility: A What-if Analysis
Q : If you make the AP mobile, things may break down (physically)easier. How do you tradeoff between higher throughput and higher chance of failure.
A: These risks, as well as psychological discomfort, are increasing with the use of robots in our life.
Q: The question is not about psychology aside, it is about reliability.
A: We can optimize for different utility functions. So far, we optimized throughput, but we could include reliability in our objective function.
Q: What is your baseline for comparison? Could you get the same benefits by simply using MIMO?
A: We are currently using a single antenna. Mobility is complementary to MIMO.
Q: All 3 papers require help from participants. How reliable are these participants and how sensitive is the outcome to optimal choices?
A: The precise placement of AP is not critical, since there is a lot of diversity, and many positions of the AP are good enough.
Q: Rather than mobilizing the AP, we could move the antennas, etc.
A: Yes, but restricted mobility means less opportunity.
Q: You seem to need a lot of computation in real time. Where should this computation be done?
A: the computation bottleneck is the search space. It can be done at the local AP. In case of multiple APs, it can happen on the cloud.
Infrastructure Mobility: A What-if Analysis
Today’s wireless access networks cannot keep up with the demand. Today’s network infrastructure (APs, cell towers) is static while users may be mobile. The main idea proposed in this paper is to make the infrastructure itself mobile, in order to exploit diversity. There is a wide range of mobility options (tethering - on the order of feet, ceiling railing - on order of meters, cell tower drones - order of km) and timescales to adapt. The authors put a disclaimer that they do not know the killer app yet and they consider this as a bottom-up enabling research. The author argued that their idea can be made practical through robots (“robotic WiFi”) and that it can provide compelling gains (in terms of SNR variation, throughput gains and other metrics, as evidenced by experimental results for micro-mini-macro mobility) without actually moving the APs much. He compared the approach to overprovisioning and argued that mobility is complementary to density. He also envisioned that the monitoring and control of mobility should be coordinated and optimized by the cloud. He listed challenges including: how to move the AP, how to coordinate this decision with other optimizations (e.g. channel selection, coding etc).
Q&A (during panel discussion, some of them addressed to all papers)
Q : If you make the AP mobile, things may break down (physically)easier. How do you tradeoff between higher throughput and higher chance of failure.
A: These risks, as well as psychological discomfort, are increasing with the use of robots in our life.
Q: The question is not about psychology aside, it is about reliability.
A: We can optimize for different utility functions. So far, we optimized throughput, but we could include reliability in our objective function.
Q: What is your baseline for comparison? Could you get the same benefits by simply using MIMO?
A: We are currently using a single antenna. Mobility is complementary to MIMO.
Q: All 3 papers require help from participants. How reliable are these participants and how sensitive is the outcome to optimal choices?
A: The precise placement of AP is not critical, since there is a lot of diversity, and many positions of the AP are good enough.
Q: Rather than mobilizing the AP, we could move the antennas, etc.
A: Yes, but restricted mobility means less opportunity.
Q: You seem to need a lot of computation in real time. Where should this computation be done?
A: the computation bottleneck is the search space. It can be done at the local AP. In case of multiple APs, it can happen on the cloud.
HotNets 2014: An AS for Us
Scribe notes by Athina Markopoulou.
PEERING: An AS for Us (presented by Ethan-Katz Bassett)
Q: How do we know that we measure the Internet and not your infrastructure?
A: Right now documentation, we also plan to keep logs.
Q: Have you thought about using IPv6?
A: We plan to look into that.
Q: Are ISPs concerned about their policies being inferred/exposed?
A: To our experience, there is no pushback from ISPs. They know their own policy but they don’t have the global picture, so they actually want more visibility.
This work developed a testbed that allows researchers to configure an ISP (PEERING) and experiment with it. PEERING has its own AS number and IP address space and it peers with real ISPs. In particular, PEERING routers peer with 6 universities and providers. Richer connectivity is provided via peers at AMS-IX (Amsterdam Internet Exchange) and Phoenix-IX. When researchers configure this AS are allowed to only advertise prefixes that PEERING owes, not other people’s prefixes (so as to not become transit). The speaker then explained the use of the testbed through the motivating example of ARROW.
PEERING provides a sweet spot between realism (running things over the Internet) and control (by configuring this ISP). It can be used to enable running experiments for inter-domain routing research. The speaker concluded with a call to the community to use the testbed and propose new features.
Q& A (during panel discussion, some of them addressed to all papers)
Q: How do we know that we measure the Internet and not your infrastructure?
A: Right now documentation, we also plan to keep logs.
Q: Scalability issues?
A: It mainly depends on the number of prefixes we own.
Q: Have you thought about using IPv6?
A: We plan to look into that.
Q: Are ISPs concerned about their policies being inferred/exposed?
A: To our experience, there is no pushback from ISPs. They know their own policy but they don’t have the global picture, so they actually want more visibility.
HotNets 2014: Crowdsourcing Access Network Spectrum Allocation Using Smartphones
Scribe notes by Athina Markopoulou.
Crowdsourcing Access Network Spectrum Allocation Using Smartphones
Q&A (During Panel Discussion, some of them addressed to all papers):
Crowdsourcing Access Network Spectrum Allocation Using Smartphones
(presented by Jinghao Shi)
The main idea proposed in this paper was to use a smartphone “within proximity” of the primary device (laptop or tablet) to collect measurements (channel utilization and WiFi scan results) without disrupting the primary device. The key participants in the PocketSniffer system are: the phone, the laptop, the PocketSniffer AP, the PocketSniffer server. Challenges that need to be addressed include the following:
- Physical proximity: use phone next to laptop to collect measurements on the laptop’s behalf.
- Incentives for the phone: one way is to use the user’s own phone to collect measurements for the user’s laptop. If Bob’s phone is used to help Alice’s laptop, credits can be offered to Bob, to use later in exchange for QoS.
- Measurement efficiency: pick the phone to use based on criteria, such as proximity and battery level.
- Measurement validation: the phone may provide wrong measurements (lazy, selfish). Solution: trust the AP + cross validation.
The author also talked about the bigger picture (global information, cooperation, use of game theory, interaction of wifi-cellular) and their implementation on Nexus 5 and a public testbed.
Q: How do you decide which phone, within proximity of the laptop, to use?
A: We can have this information and pick the closest one.
Followup Q: Using the closest phone to the laptop may not be the best proxy for measurements, due to fading. Locationsvery close to each other (exact distance depending on frequency/wavelength) may have very different signal strengths. Have you actually done measurements to validate that?
A: Yes, we pick the closer phone. Our measurements, so far, put the phone on top of the laptop. For the purposes of picking which channel to use this is good enough.
Q: Where do you envision the computation to happen?
A: Where to implement the algorithm (to decide what phone to use for measurement and what WiFi channel to use) must take into account security and privacy, concerns. In enterprise environments, there is a centralized controller.
Theia: (Simple and Cheap) Networking for Ultra-Dense Data Centers
Paper Title: Theia: (Simple and Cheap) Networking for Ultra-Dense Data Centers
Authors: Meg Walraed-Sullivan, Jitu Padhye (Microsoft Research), Dave Maltz (Microsoft)
Presenter: Meg Walraed-Sullivan
Paper Link: http://conferences.sigcomm.org/hotnets/2014/papers/hotnets-XIII-final152.pdf
Ultra-Dense Data Centers UDDCs are expensive to build, therefore, more CPUs are packed into a rack, which poses a number of challenges: 1) power and cooling 2) failure recovery , 3) tailoring applications 4) networking problem. This talk focuses in networking problem raised from packing a huge number of CPUs into a rack. Theia suggests we rethink the ToR architecture used in many data centers that doesn't scale to connect thousands of CPUs. We should get rid of star topology for fixed direct connect topology. The upside is that it is way cheaper, reduces the power requirement significantly and requires smaller physical space. However, this cause a lose of full bisection bandwidth and constraints the flexibility of the topology. Theia suggests replacing switches with patch panels for connectivity among sub-racks and to the rest of the data center. It is clear that over-subscription is unavoidable.
We require a direct topology that minimizes the through traffic and supports a wide range of graph size. The best fit is circular graph but it might not be the best options.
Questions:
Authors: Meg Walraed-Sullivan, Jitu Padhye (Microsoft Research), Dave Maltz (Microsoft)
Presenter: Meg Walraed-Sullivan
Paper Link: http://conferences.sigcomm.org/hotnets/2014/papers/hotnets-XIII-final152.pdf
Ultra-Dense Data Centers UDDCs are expensive to build, therefore, more CPUs are packed into a rack, which poses a number of challenges: 1) power and cooling 2) failure recovery , 3) tailoring applications 4) networking problem. This talk focuses in networking problem raised from packing a huge number of CPUs into a rack. Theia suggests we rethink the ToR architecture used in many data centers that doesn't scale to connect thousands of CPUs. We should get rid of star topology for fixed direct connect topology. The upside is that it is way cheaper, reduces the power requirement significantly and requires smaller physical space. However, this cause a lose of full bisection bandwidth and constraints the flexibility of the topology. Theia suggests replacing switches with patch panels for connectivity among sub-racks and to the rest of the data center. It is clear that over-subscription is unavoidable.
We require a direct topology that minimizes the through traffic and supports a wide range of graph size. The best fit is circular graph but it might not be the best options.
Questions:
Q: what is the topology of between the racks ?
A: Imaginary. We don't have a DC yet
Q: Different racks should be able to make different trade offs? What do you think about having heterogeneous racks.
A: It's hard to do it, but we might consider it.
Q: At this scale, don't you think you need clever replacement
as oversubscribe is unavoidable?
A: We need clever placement, we do it today. We will keep it in mind.
PIAS: Practical Information-Agnostic Flow Scheduling for Data Center Networks
Paper Title: PIAS: Practical Information-Agnostic Flow Scheduling for Data Center Networks
Authors: Wei Bai, Li Chen, Kai Chen (Hong Kong University of Science and Technology), Dongsu Han (KAIST), Chen Tian (HUST), Weicheng Sun (Hong Kong University of Science and Technology and SJTU)
Presenter: Li Chen
Paper Link: http://conferences.sigcomm.org/hotnets/2014/papers/hotnets-XIII-final91.pdf
Q( Brighten Godfrey UIUC) Do you think putting storage at switches in DataCenter and use them in the way proposed in "Revisiting Resource Pooling: The Case for In-Network Resource Sharing" paper would yield a better flow completion time ?
A: Yes, it is possible
Authors: Wei Bai, Li Chen, Kai Chen (Hong Kong University of Science and Technology), Dongsu Han (KAIST), Chen Tian (HUST), Weicheng Sun (Hong Kong University of Science and Technology and SJTU)
Presenter: Li Chen
Paper Link: http://conferences.sigcomm.org/hotnets/2014/papers/hotnets-XIII-final91.pdf
Existing data center network
flow scheduling schemes minimize flow completion time FCT, however,
they assume a prior knowledge of flow size information to
approximate ideal preemptive shortest job first and require
customized switch hardware. PIAS minimizes FCT without requiring
prior knowledge of flow size by leveraging multilevel feedback queue
(MLFQ) that exists in many commodity switches. The goal is to have an
information agnostic approach that minimizes FCT and being readily
deplorable. Initially, a flow gets the highest priority and is demoted as it transfer more bytes. This is achieved by tagging packets and keeping per flow state and using switches queues.
There is a challenge on how to choose a demotion threshold. This is addressed by modeling it as a static FCT minimization problem. The second challenge exists if there is a mismatch between traffic distribution and is mitigated by using Explicit Congestion Control ECN.
At the end, the system is practical and effective, information antagonistic and achieve FCT minimization, and readily deploy-able.
Questions:
Q( Brighten Godfrey UIUC) How you handle the situation when you have different hardware that has different number of queues ? do you
use the minimum ?
A : Yes, currently we use the minimum.
A: Yes, it is possible
Revisiting Resource Pooling: The Case for In-Network Resource Sharing
Paper Title: Revisiting Resource Pooling: The Case for In-Network Resource Sharing
Authors: Ioannis Psaras, Lorenzo Saino, George Pavlou (University College London)
Presenter: Ioannis Psaras
Paper Link: http://conferences.sigcomm.org/hotnets/2014/papers/hotnets-XIII-final109.pdf
Resource pooling principle is leveraged to manage shared resources in networks. The main goal is to maintain stability and guarantee fairness. TCP effectively deal with uncertainty by suppressing demand and moving traffic as fast as the path's slowest link. The approach taken in this paper is to push as much traffic in the network, once we hit a bottleneck then we store temporary in routers caches and detour accordingly. Note, in-network storage *caches* are not used for as temporary storage for the most popular content, instead, it is used to store incoming content in temporarily. The assumptions are: 1) contents have name, 2) clients send network-layer contents. In this approach, clients regulate traffic that is pushed in the network, instead of senders. Fairness and stability is achieved in three phases: 1) push data phase, 2) cache & detour phase, and 3) back-pressure phase. Evaluation shows that there is high availability of detours in real typologies.
Questions:
Q: In the table that shows the available detour paths in real typologies, 2 hops detour availability means that there is no 1 hop but there 2 hop?
A: Yes
Q( Brighten Godfrey UIUC) Do you think putting storage at switches in DataCenter and use them in the way you suggested would yield a better flow completion time ?
A: Yes, it makes sense. The approach could fit to datacenter.
Authors: Ioannis Psaras, Lorenzo Saino, George Pavlou (University College London)
Presenter: Ioannis Psaras
Paper Link: http://conferences.sigcomm.org/hotnets/2014/papers/hotnets-XIII-final109.pdf
Resource pooling principle is leveraged to manage shared resources in networks. The main goal is to maintain stability and guarantee fairness. TCP effectively deal with uncertainty by suppressing demand and moving traffic as fast as the path's slowest link. The approach taken in this paper is to push as much traffic in the network, once we hit a bottleneck then we store temporary in routers caches and detour accordingly. Note, in-network storage *caches* are not used for as temporary storage for the most popular content, instead, it is used to store incoming content in temporarily. The assumptions are: 1) contents have name, 2) clients send network-layer contents. In this approach, clients regulate traffic that is pushed in the network, instead of senders. Fairness and stability is achieved in three phases: 1) push data phase, 2) cache & detour phase, and 3) back-pressure phase. Evaluation shows that there is high availability of detours in real typologies.
Questions:
Q: In the table that shows the available detour paths in real typologies, 2 hops detour availability means that there is no 1 hop but there 2 hop?
A: Yes
Q( Brighten Godfrey UIUC) Do you think putting storage at switches in DataCenter and use them in the way you suggested would yield a better flow completion time ?
A: Yes, it makes sense. The approach could fit to datacenter.
Subscribe to:
Posts (Atom)
