DORAEMON T1-WC

America/Los_Angeles
Description

https://u-tokyo-ac-jp.zoom.us/j/98126050083

Link to recording

Quick recap

The meeting focused on defining the data sets and challenges for the T1 data challenge in water Cherenkov detectors. Kazuhiro emphasized the need to critically select data sets based on specific reconstruction challenges, rather than generating samples just because it's possible. He advised against mixing interesting studies with the core data challenge and urged the team to start from scratch rather than using existing lists. Tashiro shared a tentative list of particle samples but was asked to reconsider it based on the critical needs of Hyper-K. The discussion also covered the validation of the LUCiD simulation's new electronics modeling and the need for a baseline benchmark using fiTQun. Technical questions about metrics for multi-ring segmentation and the generation of samples for fiTQun tuning were addressed. The importance of involving experts like Ryo for the baseline and ensuring the data sets are ready for an upcoming workshop was highlighted.

Next steps

Omar A. Alterkait

  • Implement functionality in LUCiD to output per-photon information on scattering and reflection, and to ensure readiness to generate "electron bomb" samples for fiTQun tuning.

Patrick

  • Follow up with the student (Mathias) working on multi-ring segmentation to explore metrics beyond the Dice score, such as Panoptic Quality (PQ), for handling ghost rings.

Riya

  • Validate the new electronics modeling in LUCiD simulation against SK models, specifically checking the PR made by Cesar.

Takuya

  • Reconsider and propose a curated list of particle samples for the data challenge, starting from scratch based on critical reconstruction challenges, not just existing proposals.

Collaboration

  • Patrick and Takuya: Define the use case for each data set (training, benchmarking, physics application) and generate configurations accordingly.
  • Patrick and Takuya: Communicate with Ryo to define the samples needed for fiTQun tuning and ensure their generation before the workshop in Kyushu.
  • Patrick and Takuya: Find a time slot to meet with Cesar (US) to discuss data generation configurations.
  • Patrick and Takuya: Consult with Hyper-K experts to prioritize the top reconstruction challenges for the data challenge.

Summary

Genesis Sample Preparation Discussion

Takuya discussed preparing a list of samples and a tentative option table, which is based on a proposal by Cesar but includes additional samples. Patrick inquired about clarifying the distinction between data used for training versus testing or benchmarking, particularly for mu pi zero applications in physics, such as proton decay.

Data Challenge Sample Generation Strategy

The team discussed sample generation for a data challenge, with Kazuhiro emphasizing the need to have specific reasons for including each sample rather than generating data simply because it's possible. He argued that every sample should serve a unique purpose and be critically evaluated, distinguishing between smaller research studies and the main data challenge objectives. Kazuhiro stressed the importance of defining a common benchmark that the community can use, while avoiding the mixture of interesting but unrelated research studies with the main data challenge goals.

Critical Reconstruction Challenges Prioritization

Kazuhiro emphasized the need to focus on critical reconstruction challenges rather than building on existing work, urging the team to generate new priorities from scratch. Takuya agreed to reconsider the approach and start fresh for the data challenge. The discussion highlighted the importance of identifying key challenges in areas like clustering, multi-ring reconstruction, PID, and momentum reconstruction, with experts being tasked to prioritize these based on their expertise. Patrick mentioned Benda's interest in calibration and domain shift mitigation for reconstruction algorithms' robustness to detector configuration variations.

First Release Focus and Validation

Kazuhiro emphasized that the current focus should be on completing the first release without expanding the scope, and he urged the team to validate the realistic electronic simulations claimed by Cesar for the dataset. Patrick discussed the need to curate a list based on Hyper-K's needs, particularly regarding machine learning and AI exploration. Takuya mentioned considering data science activities and proposed making a proposal this week.

Electronics Modeling Dataset Validation

Kazuhiro emphasized the need to validate the new electronics modeling and configure datasets based on specific use cases, such as benchmarking or training, without using cost as a limitation. Patrick and Riya discussed ongoing work on validating LUCiD simulation and checking Cesar's implemented models. The team also considered generating common samples at the generator level for different detectors, though Kazuhiro questioned its relevance for T1. It was decided that the Japan team should prioritize data generation configurations and coordinate directly with Cesar for this task.

Neutrino Reconstruction Approach Discussion

Kazuhiro explained that while Monte Carlo samples of neutrino interactions can be generated using particle bomb (a random combination of particles), the team typically avoids using neutrino event generators due to their known inaccuracies and biases. He clarified that the reconstruction algorithm's primary focus should be on reconstructing visible particles rather than attempting to reconstruct neutrino properties directly, as this approach helps avoid learning biases from unreliable generators. The discussion highlighted a disagreement between Kazuhiro and Takuya regarding the approach to neutrino reconstruction, with Kazuhiro advocating for separating visible particle reconstruction from neutrino property inference.

T2 Data Challenge and Multi-ring Metrics Discussion

Kazuhiro explained that the T2 data challenge is already addressing the task of modeling neutral properties using event generators with reconstructed particle kinematics. Patrick raised a question about metrics for multi-ring segmentation, specifically regarding the use of Dice scores and the need for a better metric that can account for missed rings and fake "ghost rings" added by algorithms. Kazuhiro confirmed that proposing new metrics like DICE is acceptable but emphasized the need to discuss how metrics would work from both researcher and maintenance perspectives.

DORAEMON Project Sample Optimization

The team discussed generating particle samples and detector geometries for the DORAEMON project, with Kazuhiro emphasizing the need to focus on generic challenges rather than exact experiment replicas. They identified the need to optimize fiTQun as a baseline benchmark for water Cherenkov, with Ryo agreeing to lead this effort and requiring electron bomb simulations for fiTQun tuning. The team agreed to create a detailed list of required sample features, including photon scattering and reflection information, with Omar A. Alterkait confirming these modifications should be implementable in LUCiD. Kazuhiro stressed the importance of consulting with Ryo and preparing the lookup table before the upcoming meeting to ensure productive workshop sessions.