| |

Data Swarms

September 17, 2014, Storage Developer Conference, Santa Clara, CA—Yogexh Vedpathak from Cleversafe talked about the need for new approaches to moving data as the large datasets like video dominate traffic. As usual, the tradeoffs are between performance and costs.

Popular data like images and videos are dominating Web traffic and usually are only active for hours. The appearance of TV-type shows on line exacerbates the pipeline problems. The traditional method for supplying data was to have a central server feeding the data to clients. Some of the challenges for this approach are in the highly redundant data flows. Throughput slows as the number of users increases and the distant users affect throughput for all.

An alternative is to have a content distribution network with multiple servers to feed the clients. This puts the data closer to the downloaders and helps to balance the server loads but needs expensive infrastructure and creates many copies of the data. Nevertheless, creating a broadcast storage system calls for scalable performance to keep the storage load bounded. The system has to ensure delivery of unique packets and not cost a fortune.

One alternative is to use the bit torrent protocol. A .torrent file has file and tracker information. The tracker helps downloaders find each other and act as a swarm. The swarm allows downloaders to coordinate with each other to download pieces of the files. A system using bit torrent would start from a single server to seed the system. The default seeder is a gateway that communicates with the backend storage and caches the entire file.

This seeder then acts as a buffer if the file is popular and there are not enough peers or many free riders. The default seeder initiates transfers of separate chunks of the file to multiple clients. These clients then transfer their portions to other clients, so the server only has to ship out small portions of the file at a time.

Establishing a swarm is fairly simple. You can get the seed server from any IaaS provider and set up the proper interfaces to your storage system. This approach makes sense if some objects are frequently used, and if the bit torrent overhead is less than the file size. It is possible to set up a trackerless implementation. The tracker is SPoF in most P2P networks, but you can use the Mainline DHT, which is a distributed method to find peers. Instead of using the announce key, use nodes and the default seeder will act as the bootstrapping node.

In this type of sharing network, some users will be free riders. Conscious free riders tell the others that they are unwilling or unable to upload data due to limited CPU or bandwidth. These users about a second of latency for every additional 20 percent of users. Oblivious free riders do not tell that they are unwilling or unable, or have malfunctioning or uncooperative clients. These users add 1-2 seconds of latency for every 10 percent and seriously degrade the transfer process.

As a result, the uncooperative peers should be identified and excluded from the swarm. Exclusion prevents new peers from talking to these users. In addition, the various classes of users can be billed at different rates. Normal billing applies to users for direct access to the object. A discount will apply for a swarm with no free riders, but free riders can get to the swarm for free. Take down to destroy the swarm requires the default seeder to track active peers. If there are no active peers, the file can be cleared from the cache and the server can nullify the nodes in the torrent file.
 

Similar Posts