Add dms lecture notes
This commit is contained in:
1 parent
3b64218fe5
commit
ac5e4639d8
74 files changed
+16109
No files matched your search
@@ -0,0 +1 @@
|
||||
{"config":{"lang":["en"],"separator":"[\\s\\-]+","pipeline":["stopWordFilter"],"fields":{"title":{"boost":1000.0},"text":{"boost":1.0},"tags":{"boost":1000000.0}}},"docs":[{"location":"","title":"Umbra Notes","text":""},{"location":"books/designing_data_intensive_applications/part1/chapter1/","title":"Chapter 1: Reliable, Scalable and Maintainable Applications","text":"<p>Many applications today are data-intensive, as opposed to compute-intensive. Raw CPU power is rarely a limiting factor for these applications.</p> <p>A data-intensive application is built from the following building blocks</p> <ul> <li>Store data so that they, or another application can find it again later (databases)</li> <li>Remember the result of an expensive operation, to sped up reads (caches)</li> <li>Allow users to search data by keyword or filter it in various ways (search indexes)</li> <li>Send a message to another process, to be handled asynchronously (stream processing)</li> <li>Periodically crunch a large amount of accumulated data (batch processing)</li> </ul>"},{"location":"books/designing_data_intensive_applications/part1/chapter1/#thinking-about-data-systems","title":"Thinking about Data Systems","text":"<p>Database and a message queue are quite similar. They both store data for some time - though they have very different access patterns which means different performance characteristics and thus very different implementations.</p> <p>Boundaries between these implementations are becoming slightly blurred. There are data-stores that are also used as message queues (Redis) and there are messages queues with database-like durability guarantees (Apache Kafka).</p> One possible architecture for data system that combines several components <p>When you combine several tools in order to provide a service, the service's interface or application programming interface (API) usually hides those implementation details from clients.</p> <ul> <li>Reliability: The system should continue to work correctly (performing the correct function at the desired level of performance) even in the face of adversity (hardware or software faults, even human error).</li> <li>Scalability: As the system grows (in data volume, traffic volume or complexity), there should be reasonable ways of dealing with that growth.</li> <li>Maintainability: Over time, many different people will work on the system (engineering and operations, both maintaining current behaviour and adapting the system to new use cases), and they should all be able to work in it productively.</li> </ul>"},{"location":"books/designing_data_intensive_applications/part1/chapter1/#reliability","title":"Reliability","text":"<ul> <li>The application performs the function that the user expected.</li> <li>It can tolerate the user making mistakes or using the software in unexpected ways.</li> <li>Its performance is good enough for the required use case, under the expected load and data volume.</li> <li>The system prevents any unauthorized access and abuse.</li> </ul> <p>Things that ca go wrong are called faults. Systems that anticipate faults and can cope with them are called fault-tolerant or resilient. Fault tolerance does not mean making a system tolerant of all faults, but only tolerating certain types of faults.</p> <p>NOTE: A fault is not the same as a failure. </p> <ul> <li>A fault is defined as one component of the system deviating from its spec.</li> <li>A failure is when the system as a whole stops providing the required service to the user,</li> </ul> <p>It is impossible to to reduce the probability of a fault to zero; therefore it is best to design fault-tolerance mechanisms that prevent faults from causing failures.</p>"},{"location":"books/designing_data_intensive_applications/part1/chapter1/#hardware-faults","title":"Hardware Faults","text":"<p>Hard disks are reported as having a mean time to failing (MTTF) of about 10 to 50 years. So on a storage cluster with 10,000 disks, we should expect on average one disk to die per day.</p> <p>A good combatant for this is redundancy. Disks may be set up in RAID configurations, servers can have dual power supplies etc. When a component dies, the redundant component can take it's place whilst the broken one is being replaced. This approach cannot complete prevent hardware problems from causing failures, but it is well understood and can often keep a machine running uninterrupted for years.</p> <p>However, as data volumes and applications' computing demands have increased, more applications have begun using larger number of machines, which proportionally increase the rate of hardware faults. Moreover, in some cloud platforms such as AWS it is fairly common for virtual machine instances to become unavailable without warning as the platforms are designed to prioritise flexibility and elasticity over single-machine reliability.</p> <p>Hence there is a move toward systems that can tolerate the loss of entire machines, by using software fault-tolerance teLine truncated
|
||||
Reference in new issue
Block a user