Blog Post

Microsoft Foundry Blog
1 MIN READ

Re: Finetune Small Language Model (SLM) Phi-3 using Azure Machine Learning

mimillet disable flash attention if your GPU does not support it.  Here 

model_kwargs = dict(
    use_cache=False,
    trust_remote_code=True,
    attn_implementation="flash_attention_2",  # loading the model with flash-attenstion support
    torch_dtype=torch.bfloat16,
    device_map=None
)
Published Jul 01, 2024
Version 1.0
No CommentsBe the first to comment