DONGGUK UNIVERSITY

Research

iSN Lab builds the system software and networking substrate that AI actually runs on — from GPU clusters and datacenter fabrics down to edge devices and IoT. We measure real systems to find where they waste time, bandwidth, and cost, then redesign the scheduling, runtime, and network layers they sit on so the same hardware does more work.

저희 연구실은 AI가 실제로 구동되는 바탕 — GPU 클러스터와 데이터센터 네트워크에서 엣지 디바이스와 IoT에 이르기까지 — 을 이루는 시스템 소프트웨어와 네트워킹 인프라를 연구합니다. 실제 시스템을 직접 측정해 시간과 대역폭, 비용이 새는 지점을 찾고, AI 워크로드를 실행하는 스케줄링·런타임·네트워크 계층을 다시 설계해 같은 하드웨어로 더 많은 일을 해내도록 하는 연구들을 진행하고 있습니다.

System for AI

Our work spans three research areas:

연구실에서는 크게 세 가지 연구를 수행하고 있습니다:

01

AI System

Large language models stress every layer of the stack at once — GPUs, host memory, and the network between them. We study how LLM training, fine-tuning, and serving actually behave on real infrastructure, then redesign the GPU scheduling, job placement, and network paths that support them so the same cluster serves more requests for less.

대규모 언어 모델은 GPU와 호스트 메모리, 그리고 그 사이를 잇는 네트워크까지 스택의 모든 계층에 한꺼번에 부담을 줍니다. 우리는 LLM의 학습·파인튜닝·서빙이 실제 인프라에서 어떻게 동작하는지 직접 측정하고, 그 실행을 뒷받침하는 GPU 스케줄링과 작업 배치, 네트워크 경로를 다시 설계합니다. 같은 클러스터로 더 적은 비용에 더 많은 요청을 처리하여 기존 시스템보다 더 빠른 학습과 추론을 목표로 합니다.

Research topics

연구 주제

  • Resource management (GPU utilization and network) for AI model training and inference
  • Mixture-of-Experts (MoE) model serving and expert routing efficiency
  • Agentic AI systems: scheduling, orchestration, and tool-use pipelines
  • LLM structure analysis and model compression (sparsification, pruning) for system resource efficiency
  • AI 모델 학습·추론을 위한 GPU 및 네트워크 자원 관리 최적화
  • Mixture-of-Experts (MoE) 모델 서빙 및 전문가 라우팅 효율화
  • 에이전틱 AI 시스템: 스케줄링, 오케스트레이션, 도구 사용 파이프라인
  • 시스템 자원 효율을 높이기 위한 LLM 구조 분석 및 모델 압축

Publication highlights

대표 연구 논문

  • Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMsEMNLP 2026 (main)DOI
  • Accurate Simulation of Distributed Training Jobs with Network Contention ModelingIEEE MASCOTS 2026DOI
  • Prediction-based GPU Sharing for Distributed TrainingFGCS 2026DOI
  • Prediction of the Resource Consumption of Distributed Deep Learning SystemsACM SIGMETRICS 2022DOI
All publications 전체 논문 목록
Back to research areas 연구 분야 목록으로

02

Cloud and Datacenter Networking System

AI and cloud workloads need networks that are not merely fast but predictable. We build datacenter and software-defined networking systems that raise throughput, hold tenants apart under contention, and stay programmable — from the software switch on a single host up to the fabric connecting thousands.

AI와 클라우드 워크로드에 필요한 네트워크는 빠르기만 해서는 부족하고 예측 가능해야 합니다. 우리는 단일 호스트의 소프트웨어 스위치부터 수천 대를 잇는 데이터센터 패브릭까지, 처리량을 높이고 경합 상황에서도 테넌트 간 격리를 지키면서 프로그래밍 가능성을 잃지 않는 네트워킹 시스템을 구축합니다.

Research topics

연구 주제

  • Container and Kubernetes networking optimization in public clouds
  • AI for networking: learning-based traffic prediction, configuration, and control
  • Intelligent traffic splitting and load balancing in datacenter networks
  • Programmable network virtualization and software-defined networking
  • Bandwidth isolation and control plane engineering in virtual networks
  • 퍼블릭 클라우드 환경의 컨테이너 및 쿠버네티스 네트워킹 최적화
  • AI for networking: 학습 기반 트래픽 예측과 네트워크 설정·제어 자동화
  • 데이터센터 네트워크의 지능형 트래픽 분할 및 로드 밸런싱
  • 프로그래머블 네트워크 가상화 및 소프트웨어 정의 네트워킹(SDN)
  • 가상 네트워크의 대역폭 격리 및 제어 평면 엔지니어링

Publication highlights

대표 연구 논문

  • Revisiting Traffic Splitting for Software Switch in DatacenterACM SIGMETRICS 2025DOI
  • Control Channel Isolation in SDN Virtualization: A Machine Learning ApproachIEEE/ACM CCGrid 2023DOI
  • Machine Learning-Based Prediction Models for Control Traffic in SDN SystemsIEEE TSC 2023DOI
  • TeaVisor: Network Hypervisor for Bandwidth Isolation in SDN-NVIEEE TCC 2023DOI
All publications 전체 논문 목록
Back to research areas 연구 분야 목록으로

03

Edge Computing and IoT

At the edge, resources are scarce, devices are unalike, and latency budgets are counted in milliseconds. We design system software for IoT, edge platforms, and the cloud-edge pipelines between them, so that AI models train and infer efficiently, and packets arrive without delay, on constrained hardware.

엣지 환경에서는 자원이 부족하고, 디바이스는 제각각이며, 지연 예산은 밀리초 단위로 계산됩니다. 우리는 IoT와 엣지 플랫폼, 그리고 그 사이를 잇는 클라우드-엣지 파이프라인을 위한 시스템 소프트웨어를 설계합니다. 제한된 하드웨어 위에서도 AI 모델이 효율적으로 학습하고 추론하며, 패킷이 지연 없이 처리되도록 하는 것이 목표입니다.

Research topics

연구 주제

  • Packet processing for IoT and edge-native container environments
  • System support for personalized federated learning applications
  • Model lightweighting and efficient AI for resource-constrained devices
  • IoT 및 엣지 네이티브 컨테이너 환경을 위한 패킷 처리
  • 엣지 디바이스 내 개인화 학습 시스템 지원
  • 자원 제약 디바이스를 위한 모델 경량화 및 효율적 AI

Publication highlights

대표 연구 논문

  • Xronos: Heterogeneity-Aware Tensor Parallelism for Collaborative LLM Fine-Tuning on Edge CPUsIEEE MASCOTS 2026DOI
  • Parameter-Efficient 12-Lead ECG Reconstruction from a Single LeadMICCAI 2025DOI
  • Intelligent Packet Processing for Performant Containers in IoTIEEE IoT Journal 2024DOI
All publications 전체 논문 목록
Back to research areas 연구 분야 목록으로