|
[AERHM99] L. Alvisi, E.N. Elnozahy, S. Rao, S.A. Husain, and A.D. Mel. “An Analysis of Communication-Induced Checkpointing ,” in Digest of Papers, FTCS-29, The 29th Annual International Symposium on Fault-Tolerant Computing, 1999. [AFF03] A. Agbaria, A. Freund, and R. Friedman. “Evaluating Distributed Checkpointing Protocols,” in Proceedings of the 23rd International Conference on Distributed Computing Systems, 2003. [ALRL04] A. Avizienis, J.-C. Laprie, B. Randell, and C.Landwehr. Basic Concepts and Taxonomy of Dependable and Secure Computing, ” IEEE Trans. on Dependable and Secure Computing, 1(1): 11-33,2004. [Andrews99] Gregory R. Andrews. “Foundations of Multithreaded, Parallel, and Distributed Programming,”Addison-Wesley, 2000. [AS05] Adnan Agbaria and William H. Sanders.“Application-Driven Coordination-Free Distributed Checkpointing,”in Proceedings of the 25th IEEE International Conference on Distributed Computing Systems, 2005. [BHMR95] R. Baldoni, J. M. Helary, A. Mostefaoui, M. Raynal, “On Modeling Consistent Checkpoints and the Domino Effect in Distributed Systems,” Tech. Rep.RR-2564, IRISA, July 1995. [BL88] B. Bhargava and S. R. Lian. “ Independent checkpointing and concurrent rollback for recovery—An optimistic approach, “ In Proceedings of the Sixth International Conference on Data Engineering, pp 182-189 ,1988. [CL85] K.M. Chandy and L. Lamport. “Distributed Snapshots: Determining Global States of Distributed Systems, " ACM Trans. on Computer Systems,3(1):63-75, Feb. 1985. [CR72] K.M. Chandy and C.V. Ramamoorthy. “Rollback and Recovery Strategies for Computer Programs,"IEEE Trans. on Computers, 21(6):546-556, June 1972. [EAWJ02] E.N. Elnozahy, L. Alvisi, Y.-M. Wang and D.B. Johnson.“A Survey of Rollback-Recovery Protocols in Message-Passing System, " ACM Computing Surveys, 34(3): 375-408, Sept. 2002. [LWK03] C.Y. Lin, S.C. Wang and S.Y. Kuo. “An Efficient Time-Based Checkpointing Protocol for Mobile Computing Systems over Mobile IP,” Mobile Network Applications, 8:687-697, 2003. [Neumann] Peter G. Neumann. “Illustrative Risks to the Public in the Use of Computer Systems and Related Technology,”http://www.csl.sri.com/users/neumann/illustrative.html [Plank93] J.S. Plank. “Efficient Checkpointing on MIMD Architectures,” Ph.D. thesis, Princeton University,1993. [PT01] J. S. Planck and M. G. Thomason, “Processor allocation and checkpoint interval selection in cluster computing systems,” Journal of Parallel and distributed Computing, 2001. [Randell75] B. Randell. “System structure for software fault tolerance,” IEEE Trans. on Software. Engineering.1(2):220-232, 1975. [ROC] The Berkeley/Stanford Recovery-Oriented Computing Project. http://roc.cs.berkeley.edu/ [Tsai05] Jichiang Tsai. “An Efficient Index-Based Checkpointing Protocol with Constant-Size Control Information on Messages,”IEEE Trans. on Dependable and Secure Computing, 2(4): 287-296,Oct.-Dec. 2005. [Wang93] Y.-M Wang. “Space Reclamation for Uncoordinated Checkpointing in Message-Passing Systems,” Ph.D. Thesis, University of Illinois, Department of Computer Science, 1993. [Wang97] Y.-M Wang. “Consistent Global Checkpoints that Contain A Set of Local Checkpoints,” IEEE Trans. on Computers, 46(4):456-468, 1997. [WF93] K. F. Wong and M. A. Franklin, “Distributed computing systems and checkpointing,” in 2nd Int. Symp. On High Performance Distributed Computing (HPDC’93). IEEE CS Press, pp. 224–23, July 1993.
|