Blog(1)
Blog(1)
Research Areas(0)
Publications(22)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
Efficient Compositional Multi-tasking for On-device Large Language Models
On-device System of Compositional Multi-tasking in Large Language Models
News(1)
Others(0)