At the heart of experimental high energy physics (HEP) is the development of
facilities and instrumentation that provide sensitivity to new phenomena. Our
understanding of nature at its most fundamental level is advanced through the
analysis and interpretation of data from sophisticated detectors in HEP
experiments. The goal of data analysis systems is to realize the maximum
possible scientific potential of the data within the constraints of computing
and human resources in the least time. To achieve this goal, future analysis
systems should empower physicists to access the data with a high level of
interactivity, reproducibility and throughput capability. As part of the HEP
Software Foundation Community White Paper process, a working group on Data
Analysis and Interpretation was formed to assess the challenges and
opportunities in HEP data analysis and develop a roadmap for activities in this
area over the next decade. In this report, the key findings and recommendations
of the Data Analysis and Interpretation Working Group are presented.
Particle physics has an ambitious and broad experimental programme for the
coming decades. This programme requires large investments in detector hardware,
either to build new facilities and experiments, or to upgrade existing ones.
Similarly, it requires commensurate investment in the R&D of software to
acquire, manage, process, and analyse the shear amounts of data to be recorded.
In planning for the HL-LHC in particular, it is critical that all of the
collaborating stakeholders agree on the software goals and priorities, and that
the efforts complement each other. In this spirit, this white paper describes
the R&D activities required to prepare for this software upgrade.
Complex scientific workflows can process large amounts of data using
thousands of tasks. The turnaround times of these workflows are often affected
by various latencies such as the resource discovery, scheduling and data access
latencies for the individual workflow processes or actors. Minimizing these
latencies will improve the overall execution time of a workflow and thus lead
to a more efficient and robust processing environment. In this paper, we
propose a pilot job based infrastructure that has intelligent data reuse and
job execution strategies to minimize the scheduling, queuing, execution and
data access latencies. The results have shown that significant improvements in
the overall turnaround time of a workflow can be achieved with this approach.
The proposed approach has been evaluated, first using the CMS Tier0 data
processing workflow, and then simulating the workflows to evaluate its
effectiveness in a controlled environment.