OER·harvester

← Back to the library
GitHub NOTEBOOK resource

Phi35 Moe Demo

Repository 21 Lessons, Get Started Building with Generative AI

Licence
OPEN MIT
Authors
Microsoft (microsoft)
Updated
2025-11-17 · GitHub
Language
en detected
Length
86 words
Open original ↗
{ }
Jupyter notebook Converted to a read-only view · code is not executed

Source: Phi35 Moe Demo · GitHub · microsoft/generative-ai-for-beginners Authors: Microsoft (microsoft) Licence: MIT — https://spdx.org/licenses/MIT.html

In [1]
# pip install transformers
In [2]
# pip install torch torchvision torchaudio -U
In [3]
# pip install flash-attn --no-build-isolation
In [4]
# ! pip install flash_attn -U
In [5]
from torch import bfloat16
import transformers
In [6]
model_id = "../Phi3MOE"
In [7]
model = transformers.AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=bfloat16,
    device_map='auto'
)
In [8]
model.eval()
In [9]
tokenizer = transformers.AutoTokenizer.from_pretrained(model_id)
In [10]
# generate_text = transformers.pipeline(
#     model=model, tokenizer=tokenizer,
#     return_full_text=False,  # if using langchain set True
#     task="text-generation",
#     # we pass model parameters here too
#     temperature=0.1,  # 'randomness' of outputs, 0.0 is the min and 1.0 the max
#     top_p=0.15,  # select from top tokens whose probability add up to 15%
#     top_k=0,  # select from top 0 tokens (because zero, relies on top_p)
#     max_new_tokens=2048,  # max number of tokens to generate in the output
#     repetition_penalty=1.1  # if output begins repeating increase
# )

pipe = transformers.pipeline(

    "text-generation",

    model=model,

    tokenizer=tokenizer,

)
In [11]
generation_args = {

    "max_new_tokens": 512,

    "return_full_text": False,

    "temperature": 0.3,

    "do_sample": False,

}
In [12]
sys_msg = """You are a helpful AI assistant, you are an agent capable of using a variety of tools to answer a question. Here are a few of the tools available to you:

- Blog: This tool helps you describe a certain knowledge point and content, and finally write it into Twitter or Facebook style content
- Translate: This is a tool that helps you translate into any language, using plain language as required

To use these tools you must always respond in JSON format containing `"tool_name"` and `"input"` key-value pairs. For example, to answer the question, "Build Muliti Agents with MOE models" you must use the calculator tool like so:

```
```json

{
    "tool_name": "Blog",
    "input": "Build Muliti Agents with MOE models"
}

```

Or to translate the question "can you introduce yourself in Chinese" you must respond:

```json

{
    "tool_name": "Search",
    "input": "can you introduce yourself in Chinese"
}

```

Remember just output the final result, output in JSON format containing `"agentid"`,`"tool_name"` , `"input"` and `"output"`  key-value pairs .:

```json

[


{   "agentid": "step1",
    "tool_name": "Blog",
    "input": "Build Muliti Agents with MOE models",
    "output": "........."
},

{   "agentid": "step2",
    "tool_name": "Search",
    "input": "can you introduce yourself in Chinese",
    "output": "........."
},
{
    "agentid": "final"
    "tool_name": "Result",
    "output": "........."
}
]

```

The users answer is as follows.
"""
In [13]
```python
def instruction_format(sys_message: str, query: str):
    # note, don't "</s>" to the end
    return f'<|system|> {sys_message} <|end|>\n<|user|> {query} <|end|>\n<|assistant|>'
In [14]
query ='Write something about Generative AI with MOE , translate it to Chinese'
In [15]
input_prompt = instruction_format(sys_msg, query)
In [16]
input_prompt
In [17]
import torch

torch.cuda.empty_cache() 

import os

os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True "
In [18]
# res = generate_text(input_prompt)

output = pipe(input_prompt, **generation_args)
In [19]
output[0]['generated_text']