RSSAmplifier

Blog

shunk031.me

shunk031.me

shunk031.meRSS feed ↗183 posts

Latest posts

Activation Steeringにおける文章崩壊の抑制に向けた初期検討

Where Do Vision-Language Models Preserve and Use Depth Information?

Our Presentations at YANS 2026

We will present the following papers at YANS 2026: 陳 冠帆, 川田 拓朗, 守田 竜梧, 北田 俊輔, 彌冨 仁. “Where Do Vision-Language Models Preserve and Use Depth Information?” . 西山 天, 川田 拓朗, 北田 俊輔, 永井 大地, 彌冨 仁. “Activation Steeringにおける文章崩壊の抑制に向けた初期検討” .

DesignCorrection-R1: Learning to Reason for Graphic Design Correction

Generation and Evaluation of Editable Graphical Abstracts for Academic Papers

Our Presentations at MIRU 2026

We will present the following papers at MIRU 2026: Shunsuke Kitada, Shoma Iwai, Ryota Yoshihashi, Atsuki Osanai. “DesignCorrection-R1: Learning to Reason for Graphic Design Correction” . Takuro Kawada, Shunsuke Kitada, Hitoshi Iyatomi. “Generation and Evaluation of Editable Graphical Abstracts for Academic Papers” .

Accepted our papers to ACL2026 SRW

The following papers have been accepted to the ACL 2026 Student Research Workshop (SRW) . Ryuhei Miyazato, Shunsuke Kitada, and Kei Harada. “EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors” . Michito Takeshita, Takuro Kawada, Takumi Ohashi, Shunsuke Kitada, and Hitoshi Iyatomi. “A11y-Compressor: A Framework for Enhancing the Efficiency of…

A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction

EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors

Compressed-a11y: 視覚的文脈の再構成と冗長性削減による GUI エージェント観測の効率化

MLP の重みを反映した Sparse Autoencoder の初期化手法の提案

SciGA-Vec: 学術論文におけるベクタ画像形式の Graphical Abstract データセット

ステアリングベクトルは日本語LLMを堅牢に制御できるか?

Our Presentations at ANLP 2026

We will present the following papers at ANLP 2026: 川田 拓朗, 北田 俊輔, 彌冨 仁. “SciGA-Vec: 学術論文におけるベクタ画像形式の Graphical Abstract データセット” . 菊谷 幹, 北田 俊輔, 原 聡. “MLP の重みを反映した Sparse Autoencoder の初期化手法の提案” . 北田 俊輔, 原 聡. “ステアリングベクトルは日本語LLMを堅牢に制御できるか?” . 竹下 理斗, 川田 拓朗, 大橋 巧, 北田 俊輔, 彌冨 仁. “Compressed-a11y: 視覚的文脈の再構成と冗長性削減による GUI エージェント観測の効率化” .

Accepted our papers to Findings of CVPR 2026

The following papers have been accepted to the Findings of CVPR 2026: Takuro Kawada, Shunsuke Kitada, Sota Nemoto, Hitoshi Iyatomi. “SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers” . Daichi Nagai, Ryugo Morita, Shunsuke Kitada, Hitoshi Iyatomi. “TAUE: Training-free Noise Transplant and Cultivation Diffusion Model” .

A Coding-Agent-Friendly Environment Is Friendly to Humans Too: ghq × gwq × fzf

As your number of repositories grows through git clone , for both work and personal projects, it becomes harder to remember where you cloned things and which directory you’re currently working in. You end up wasting time on cd and shell completion. On top of that, once you start creating branches for development or checking out other branches to review Pull Requests, the overhead of switching…

ICCV2025 Conference Participation Report

Our Paper and Workshop Accepted at ICCV 2025, the Top International Conference on Computer Vision (Participation Report) Hello, I’m Shunsuke Kitada ( @shunk031 ), and I work at LY Corporation on research and development for image generation and design generation. From October 19 to 23, 2025, I attended and presented at the International Conference on Computer Vision, ICCV 2025 , held in Hawaii,…

[Invited Talk] テストを書かない研究者に送る “最初にテストを書く” 実験コード入門 – オレオレ最強 main.py から抜け出すために –

堅牢.py #1 堅牢.py はPythonを安全に扱うための技術に興味を持つエンジニアのイベントです。今回が初開催です!! 資料 - Slides

Our Presentations at IBIS 2025

We will present the following paper at IBIS 2025: 永井 大地, 守田 竜梧, 北田 俊輔, 彌冨 仁. “潜在拡散モデルでのレイアウト指定生成に向けたアテンションに基づく初期ノイズ最適化” .

潜在拡散モデルでのレイアウト指定生成に向けたアテンションに基づく初期ノイズ最適化

TAUE: Training-free Noise Transplant and Cultivation Diffusion Model

On Python Type Hints that Save Everything!

Table of Contents Introduction The difficulty of a dynamically typed language What are type hints? Example 1: Basic type hints for function arguments and return values Example 2: When the return value may be one of several types Example 3: Type hints for collection types Example 4: Optional type Example 5: Callable , Any , and other type hints Benefits of using type hints 1) Improved readability &…

Testable Dotfiles Management With Chezmoi

This article explains an approach to dotfiles management that emphasizes testability, using the author’s dotfiles repository shunk031/dotfiles as a case study. Table of Contents Introduction Dotfiles and Dotfiles Repositories The Problem of Untested Scripts My Repository’s Approach: Testable Configuration Architecture Design: Testable Configuration Repository Structure Design…

Foundation Data for Industrial Tech Transfer

We are happy to announce the ICCV 2025 workshop “FOUND: Foundation Data for Industrial Tech Transfer” has been accepted! This workshop will be held in conjunction with ICCV 2025, one of the top conferences in computer vision. The workshop will highlight the critical role of data in building foundation models, with a particular focus on robust dataset construction and industry-driven applications.…

ローカル LLM を用いた AI エージェントの現状と課題

WRIME-TC:時間的文脈による書き手と読み手の感情分析の強化

日本語の広告画像向けタイポグラフィ属性の生成型解析:小規模 VLM の LoRA 微調整による検証

GenGA:学術論文における編集可能な Graphical Abstract の自動生成に関する初期検討

Our Presentations at YANS 2025

We will present the following papers at YANS 2025: 陳 冠帆, 田中 優太朗, 中川 翼, 北田 俊輔, 川田 拓朗, 彌冨 仁. “WRIME-TC:時間的文脈による書き手と読み手の感情分析の強化” . 川田 拓朗, 北田 俊輔, 彌冨 仁. “GenGA:学術論文における編集可能な Graphical Abstract の自動生成に関する初期検討” . 中町 礼文, 浦宗 龍生, 吉橋 亮太, 和田 有輝也, 北田 俊輔, 牧田 光晴. “日本語の広告画像向けタイポグラフィ属性の生成型解析:小規模 VLM の LoRA 微調整による検証” . 竹下 理斗, 川田 拓朗, 大橋 巧, 北田 俊輔, 彌冨 仁. “ローカル LLM を用いた AI…

Our Presentations at MIRU 2025

We will present the following papers at MIRU 2025: 川田 拓朗, 北田 俊輔, 根本 颯汰, 彌冨 仁. “グラフィカルアブストラクト推薦と評価の統合ベンチマーク” . Jiahao Zhang, Ryota Yoshihashi, Shunsuke Kitada, Atsuki Osanai, Yuta Nakashima. “VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction” .

VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction

グラフィカルアブストラクト推薦と評価の統合ベンチマーク

Appointed as Researcher at The University of Electro-Communications

I’m thrilled to announce that, in addition to my main role at LYCorp., I’ve just started as an part-time researcher at The University of Electro-Communications! Under the guidance of Prof. Satoshi Hara , I’ll be working on the K Program “Red Teaming Framework for Misalignment in Large Language Models.” Looking forward to making an impact in both industry and academia!

Accepted our Workshop Proposal on Foundation Data to ICCV 2025

Our workshop proposal “FOUND: Foundation Data for Foundation Models” has been accepted at ICCV 2025, one of the top conferences in computer vision. This workshop will highlight the critical role of data in building foundation models, with a particular focus on robust dataset construction and industry-driven applications. More Information about the Workshop It is a great honor to serve…

Published "Image Generation with Python" from Impress Books

I’m thrilled to announce the release of my new book, “Image Generation with Python”! With the rapid advancements in technology, image generation has become more accessible than ever. This comprehensive guide delves into the fascinating world of image generation, focusing on diffusion models and Stable Diffusion 🌟 Whether you’re an AI engineer, a student interested in…

書籍「Pythonで学ぶ画像生成」

Our Presentations at ANLP 2025

We will present the following paper at ANLP 2025: 川田 拓朗, 根本 颯汰, 北田 俊輔, 彌冨 仁. “SciGA: 学術論文における Graphical Abstract 設計支援のための統合データセット” .

SciGA: 学術論文における Graphical Abstract 設計支援のための統合データセット

Our Presentations at IPSJ 2025

We will present the following paper at IPSJ 2025: 永井 大地, 守田 竜梧, 北田 俊輔, 彌冨 仁. “潜在拡散モデルにおける生成対象の個別配置と生成による画質改善の試み” .

潜在拡散モデルにおける生成対象の個別配置と生成による画質改善の試み

SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers

Accepted our journal paper to IEEE Access

The following paper has been accepted to the IEEE Access . This paper is an extended version of the content presented at the “Practical ML for Low Resource Settings Workshop @ ICLR 2024” . Sota Nemoto, Shunsuke Kitada, Hitoshi Iyatomi. “Majority or Minority: Data Imbalance Learning Method for Named Entity Recognition” .

Majority or Minority: Data Imbalance Learning Method for Named Entity Recognition

This paper is an extended version of the content presented at the “Practical ML for Low Resource Settings Workshop @ ICLR 2024” . We are honored to receive a rating of 9, which is in the top 15% of accepted papers, and a strong accept. We also gave an oral presentation in the ICLR workshop. arXiv IEEE SCImago

VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction

ECCV 2024 速報

I contributed a summary on creative graphic design to the ECCV2024 report on cvpaper.challenge group. You can find my summary on pages 75 to 81 below: https://hirokatsukataoka.net/temp/presen/241004ECCV2024Report_finalized.pdf#page=75 https://hirokatsukataoka.net/temp/presen/241004ECCV2024Report_finalized.pdf#page=81

Our Presentations at YANS 2024

We will present the following papers at YANS 2024: 川田 拓朗, 根本 颯汰, 北田 俊輔, 彌冨 仁. “学術論文における Graphical Abstract 自動生成の初期検討” . 水口 徳人, 北田 俊輔, 守田 竜梧, 彌冨 仁. “text-to-image 拡散モデルにおける誘導 attention map を用いた画像生成手法の提案” . 永井 大地, 根本 颯汰, 北田 俊輔, 彌冨 仁. “潜在拡散モデルにおけるプロンプトを用いた配色制御の試み” . 小川 剛毅, 根本 颯汰, 北田 俊輔, 彌冨 仁. “大規模言語モデルを用いたオノマトペ付与による日本語音声データセットの拡張” .

text-to-image 拡散モデルにおける誘導 attention map を用いた画像生成手法の提案

大規模言語モデルを用いたオノマトペ付与による日本語音声データセットの拡張

学術論文における Graphical Abstract 自動生成の初期検討

潜在拡散モデルにおけるプロンプトを用いた配色制御の試み