TitleDynamic Workflow Management and Monitoring Using DDS
Publication TypeConference Paper
Year of Publication2010
AuthorsPan, P., A. Dubey, and L. Piccoli
Conference Name7th IEEE International Workshop on Engineering of Autonomic & Autonomous Systems (EASe)

Large scientific computing data-centers require a distributed dependability subsystem that can provide fault isolation and recovery and is capable of learning and predicting failures to improve the reliability of scientific workflows. This paper extends our previous work on the scientific workflow management systems by presenting a hierarchical dynamic workflow management system that tracks the state of job execution using timed state machines. Workflow monitoring is achieved using a reliable distributed monitoring framework, which employs publish-subscribe middleware built upon OMG Data Distribution Service Standard. Failure recovery is achieved by stopping and restarting the failed portions of workflow directed acyclic graph.

