rottentomatoes.com/a/segment_anything_in_medical_images: An Exercise in Peer Review! 🍅⭐

Hello! It's that time of the month again!

FYI: Someone mistook the due date of this blog post and thought it was required before the next seminar! Sorry!!

Our topic this time around is peer reviewing scientific subject material. Reviews are a quality assurance mechanism to help editors decide whether a given manuscript should be accepted to a journal or conference based on peer and editorial standards. 

Anecdotally, whether standards are consistently, justifiably applied during the feedback process seems to be hit-or-miss. To keep things nice and clean, we'll try for a hit with today's blog entry by applying the review template from the International Joint Conference on Artificial Intelligence (IJCAI), assigning 1-10 star/s for each non-comment criterion...and yes, it is taking a lot of self-discipline to not do the fun thing - devolving into baseless critique and juicy back-and-forth à la OpenReview comment sections. Thanks for asking.

As per usual, Ma et. al's influential MedSAM proposal, which introduces a fine-tuned foundation model for universal medical image segmentation, will take one for the team and play victim of my opinion.

Ground Rules 📏

Our last seminar's slide pack had some solid resources for reading and writing as a peer reviewer, so I'll begin with a few reminders before we get into the meat & potatoes of QA'ing.

Reviewers are typically roped in for their specialised/technical knowledge in the field, so we don't need to worry as much about grammar and style unless truly inhibiting. That's the editor's job. Our primary objectives can instead be organised by service to the other parties involved, as follows:

  • Assist the journal editors with recommendations toward an acceptance decision based on the paper substance - AKA are the claims of the authors supported?
  • Assist the paper authors with constructive feedback so they can improve their work and ultimately qualify for publication.

Within the IJCAI review template, I'll rate each criterion with these goals in mind and discuss their ratings in the free-text comment to the authors, secondarily backed by PLOS's general guidelines for effective feedback, and incorporating my viewpoint and evaluation of potential.

Playing fair with the actual publication context of the paper in question, our review will also gloss over formatting/layout deviations and we'll employ a healthy dose of suspended disbelief to pretend the work has not since been iterated upon with incremental model versions and variations.

Review 👁️👁️

Criteria

Description

Rating

Relevance

·          Is the paper fully within the scope of the conference?

·          Will the questions and results of the paper be of interest to researchers in the field?

⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐

Significance

·          Is this a significant advance in the state of the art?

·          Is this a paper that people are likely to read and cite?

·          Does the paper address an important problem?

·          Does it open new research directions?

·          Is it a paper that is likely to have a lasting impact?

⭐⭐⭐⭐⭐⭐⭐⭐⭐

Originality

Reviewers should recognise and reward papers that propose genuinely new ideas. As a reviewer you should try to assess whether the ideas are truly new. Novel combinations,

adaptations or extensions of existing ideas are also valuable.

⭐⭐⭐⭐⭐⭐⭐⭐

Technical Quality

·          Are the results technically sound?

·          Are there obvious flaws in the conceptual approach?

·          Are claims well-supported by theoretical analysis or experimental results?

·          Are the experiments well thought out and convincing?

·          Will it be possible for other researchers to replicate these results? Is the evaluation appropriate?

·          Did the authors clearly assess both the strengths and weaknesses of their approach?

⭐⭐⭐⭐⭐⭐⭐⭐

Clarity and quality of writing

·          Is the paper clearly written?

·          Is there a good use of examples and figures?

·          Is it well organized?

·          Are there problems with style and grammar?

·          Are there issues with typos, formatting, references, etc.?

 

It may be better to advise the authors to revise a paper and submit to a later conference, than to accept and publish a poorly written version. However, if the paper is likely to be accepted, please make suggestions to improve the clarity of the paper, and provide details of typos

⭐⭐⭐⭐⭐⭐⭐⭐

Scholarship (Scientific Context)

·          Does the paper situate the work with respect to the state of the art?

·          Are relevant papers cited, discussed, and compared to the presented work?

⭐⭐⭐⭐⭐⭐⭐⭐⭐

 Comments to Authors:

The authors’ employment of fine-tuning techniques in this manuscript showcases an important development for AI use cases in healthcare. State-of-the-Art (SoTA) results highlight a possible paradigm shift for the task of medical image segmentation. The implications of this shift will be of interest to researchers in various (sub-) fields – e.g., general image segmentation, foundation model adaptation, and domain transfer methods. Paper content therefore fits well into conference scope and represents a significant advancement for the domain task in question.

The theoretical bases of the paper are not novel but their careful combination, application, and rigorous evaluation give way to novel insights and opportunities for AI and medical experts, alike. The breadth and conditions of experimentation lend great credibility to the study outcomes. Thoroughly detailed baseline model training substantiates their use as valid representations for current SoTA results. Technical methods and result comparisons are justified with respect to the core contribution and how it addresses the limitations of current/traditional approaches. These issues are well discussed, give shape to the paper’s contribution, and contextualise this contribution amongst established and emerging yet similar solutions.

The provision of source data and training code alongside the paper supports reproducibility. Providing more detail to substantiate decisions in preprocessing and tuning (including hyperparameter settings) processes used to train the final model would strengthen this.

It would be particularly interesting to understand how the authors decided on intensity range, resolution, and dimension standardisation for cross-modal dataset merging. Data augmentation is a common and critical practice in various computer vision applications in medical imagery due to sample shortages. Because the final model was trained without augmentation, it would also be interesting to understand whether image transforms were considered, tested, and what their observed impact was on model performance and generalisation. Altogether, it is expected that a proposal paper will primarily focus on the best version of the model being proposed. Nonetheless, supplementary information about what did not work/unexpected results during the learning process could offer insights for other transfer learning cases in computer vision (especially leveraging SAM – of which there are many other fine-tuning efforts, even for medical imagery alone).

The authors briefly mention the limitations of the dataset, posing potential impacts due to modality imbalances. Direct references/examples to empirical results capturing this limitation would be helpful - i.e., possible vs. actual. The paper would also benefit from further reflection on other limitations, such as computational efficiency (training details and scaling observations could be supplemented with inference time, cost, and hardware specifications) and interactivity (considering other prompt types) – as each of these contribute to overall utility in a clinical setting. Image dimensionality limits are justified (e.g., electing to adapt a 2D model vs. developing a 3D-native model), but this argument is superseded by the authors’ later work on the 3D/video-native SAM2. Such discussion could also be expanded on to highlight specific research directions for future work.

Otherwise, the paper is clearly written. Ideas are logically sequenced with appropriate levels of detail for each section. Result figures and tables are presented well and easy to parse for non-experts. There is minor room for improvement in terms of external structure and content organisation. The authors should consider moving some of the discussion under Results to Methods (e.g., experiment choice discussions) to avoid repetition and support better readability. Re-ordering Methods to precede the Results section would also assist the overall paper flow. Addition of a section to specifically address Limitations is recommended if revisions are made as per previous feedback.

Overall Score: [9] – An excellent paper, a very strong accept.

Confidence: [6] – I don’t have complete knowledge of the area but can assess the value of the work

Additional Commentary: Publication Impact & Reception

In the realm of medical image segmentation, the publication and reception of MedSAM is best described as pivotal. For context, the paper was published in January 2024 and since then, has been recognised as an unprecedented success in adapting Meta AI’s Segment Anything Model (SAM) meaningfully to the medical domain. The paper's results offered various insights to the research community and provided an empirical basis for what was already suspected - that foundation models show strong fundamental promise in medical applications but domain-specific fine-tuning is crucial. This catalysed a wave of adaptations (just look at the top submissions of the CVPR 2024 challenge!), including prompt-specific adapters, 3D extensions, efficiency optimisations, and earlier this year - MedSAM2 - which integrates additional modalities like videos or temporal sequences, e.g., slice-by-slice 3D volume processing, by re-using Meta's SAM2 framework.

As a natural consequence of its remarkable performance, MedSAM has influenced how researchers think about foundation models in medicine. The paper has been cited in successive works as a landmark model with specific respect to the future of zero-shot learning and domain adaptation in clinical AI. Overall, publication established generalist architectures as strong baselines for medical image segmentation in a scientific context where specialist models previously reigned supreme (and who doesn't love a good coup?! 👑) 

Comments

  1. Really appreciated the structured and insightful review here especially how it balances technical critique with constructive feedback.

    ReplyDelete

Post a Comment

Popular posts from this blog

Managing Oneself: A Lesson in Herding Cats (except it's actually only one cat...and the cat has acute time management issues that the vet is starting to think are less of a professional weakness and more of an inherent character flaw)

Personal Intro, Motivation, and Learning Goals