DORAEMON T1-WC

→ America/Los_Angeles
Description

https://u-tokyo-ac-jp.zoom.us/j/98126050083

Link to recording

 

Quick recap

This meeting focused on planning the DORAEMON Thrust 1 (T1) water Cherenkov data challenge. Patrick explained the overall structure of DORAEMON, which aims to standardize data sets and create open data challenges using AI models to improve physics analysis. The group discussed Cesar's work using LUCiD, a new differentiable simulator, to generate realistic water Cherenkov events that could be used for public data challenges. Key decisions included whether to include detailed particle track segmentation (case C) versus simpler event classification (case A), and the need to expand the current particle list to include gamma rays, more proton decay modes, and additional hadronic interactions. Benda suggested adding calibration and detector shift challenges to address data-MC shift issues, though implementation of complex detector variations like coil failures would require significant additional work. The team agreed to document their requirements in the shared Google Doc and meet again the following week to finalize proposals before the upcoming workshop.

Next steps

Benda Xu

  • Add ideas and proposals for calibration and detector parameter variations (e.g., detector shifts, data-MC consistency) to the Google Doc, formalizing the discussion into a proposal for the DORAEMON data challenge.

Patrick

  • Ask the DORAEMON group if there is an existing proposal format/example to follow for the data challenge.

Takuya

  • Prepare a list of particle data sets (including additional particles and event types needed) in a Google Doc or spreadsheet, to be used as the final list for the data challenge.

Collaboration

  • All (Patrick, Takuya, Benda Xu, Shuoyu): Contribute to the Google Doc to define tasks, metrics, and dataset definitions for the T1 Water Cherenkov data challenge proposal.
  • All: Meet weekly at the same time next week to continue planning for the workshop.

Summary

DORAEMON Project Planning Discussion

Patrick and Takuya discussed the DORAEMON structure, which has been divided into thrusts with water Cherenkov (WC) that still needs to be defined. Patrick explained that much of the work has been happening on the US side with Cesar, Omar, Junjie, and Kazu, but the team needs to determine what would be most useful for Hyper-K.

DORAEMON Physics AI Project

Patrick explained the DORAEMON project, which is based on an open data challenge, aimed at organizing community efforts to solve physics problems using AI models. The project has three thrusts, with Thrust 1 focusing on data-driven reconstruction, which has seen the most development so far within Hyper-K. Patrick noted that while they previously attempted to standardize data sets by consolidating them on a common cluster, this hasn't been consistently adopted by the Hyper-K Reconstruction Group, making it challenging to expand to the public community.

Hyper-K Physics Metrics Discussion

Patrick discussed a working document prepared by Cesar, Omar, and Kazu, along with Benda and himself, outlining tasks and metrics for Hyper-K physics, including vertex and momentum reconstruction, classification, and interaction reconstruction. He explained the importance of defining metrics for evaluating performance on these tasks, particularly for multi-particle events, using parameters like IoU for pixel grouping and position, direction, and momentum reconstruction errors. Patrick also reviewed Cesar's presentation at NPML, which demonstrated the use of the LUCiD differentiable simulator to generate realistic water Cherenkov events, though the physics representation needs further validation. The discussion highlighted the need to define detailed labels and segmentation approaches for multi-particle events to improve physics analysis sensitivity.

Particle Track Reconstruction Approaches

Patrick and Takuya discussed different approaches to particle track reconstruction in Hyper-K data, with Patrick noting that Mathias's previous multi-ring segmentation work only considered whole tracks rather than individual particle scattering and track segmentation. Takuya suggested that including more detailed track information (case C in Cesar's slides) could help constrain uncertainties in high-energy analysis, particularly for pion scattering. The team agreed they need to make decisions about which data generation and truth label approach to use and communicate this to the rest of the group for the data challenge.

Hyper-K Pilot Production Discussion

The team discussed a pilot production using super-K-like geometry and considered switching to HK far detector and IWCD, while also maintaining interest in Super-K and WCTE existing experiments. They reviewed a list of simulated event types including single particles, multi-particle combinations, and specialized samples, questioning if the current list was complete and considering limitations in data size and organizational work for future challenges. The discussion highlighted missing gamma events and differences in multi-particle event creation method used by Mathias in HK. Takuya asked about the absence of muon+pizero events in the initial set, which is just because it was not a completely informed set yet.

LUCiD Data Generation Strategy

Benda Xu proposed using LUCiD to generate data for a MC consistency challenge by leveraging different phases of SK (Super-Kamiokande) data, including parameters from SK6 and SK7.5, with the possibility of deliberately modifying the LUCiD setup to create more challenging scenarios. Patrick raised concerns about implementing missing physics in LUCiD, particularly coil failures, and questioned the feasibility of using SK data to improve LUCiD models while maintaining public access to the generated datasets.

Detector Model Development Discussion

Patrick and Benda Xu discussed the development of a model to describe changes in PMT response based on position in the detector and magnetic field, which they plan to implement eventually in general. They noted that while particle and photon propagation are already implemented in their LUCiD model, detector variations and calibration shifts still need to be developed. The team agreed to document all desired features in their Google Doc for the upcoming workshop, focusing on "blue skies" ideas without specific time/personpower constraints, before prioritizing.

Workshop Task Definition Planning

The team discussed plans for defining tasks and datasets for an upcoming workshop and challenge proposal. Benda proposed using different detector physics models (WCSim vs LUCiD) to create MC challenges, while Takuya will handle particle dataset definitions. Patrick emphasized the need to formalize their discussions into a proposal for DORAEMON, including tasks, metrics, and dataset definitions. The group agreed to meet weekly, with Benda and Takuya confirming availability for next week's meeting.