Caution
You will be deploying 3 services in total; project, lab, and redis. You should expect 3 services, 3 deployments, 2 horizontal pod autoscalers, and a single virtual service that will route based on the path. The only incremental addition is the project. Extend your existing kustomize overlay to deploy the project. You can copy your kustomize scripts from lab and extend to add the project. This will stop you from impacting your lab4 autograder.
Caution
You will need to install git-lfs in order to pull the model locally. https://git-lfs.github.com/. You will be pulling your model from huggingface. You can see a tutorial for pulling https://huggingface.co/docs/hub/en/repositories-getting-started#cloning-repositories.
Caution
torch does not support Intel-based Mac's anymore. If you have migrated from an intel-based mac at some point to a new ARM-based Mac you might run into issues. A simple test to verify this is to run the following poetry run python -c "import platform; print(platform.machine())" This should show arm64 for a mac to work as expected. Similarly arm64 is not supported on Windows, a recent student had issues with a new Microsoft Surface Pro; such computers are not supported by torch. If you have this issue the best solution is to launch a virtual machine (either on your computer or in a cloud environment) and do the work on a virtual machine.
Note
The model for your lab was roughly 1 MB, while the model for the project is 1 GB. Think about the implications this has for the resources required for your application to run given you have to load your model. Adjust limits accordignly without wasting resources.
Note
One requirement of the project is that the system is fast when scaling. This implies that new pods are coming online. If the model has to be pulled from huggingface, and that model is quite large (~500 MB), then it will take a long time for a new pod to come online. During scaling events such as when k6 is being run this will lead to a significant amount of latency as existing pods won't be able to keep up with the load. You should bake the model into the image similarly to how we have done for the lab. The best practice to handle this would be to mount the model from shared storage instead, but that's a lot of extra work for students to understand.
Note
The training script is provided as a reference. You will not need to run the training yourself as it requires significant GPU resources.
Note
Your image will be fairly large due to torch and the model being roughly 500 MB each. You may run into some networking issues pushing to ecr.
Expose both your project and lab over your virtual service to show that we can support multiple services
Since we wrote the lab where all endpoints are /lab by mounting and the project is setup the same way except for /project we can simply route based on the url path to a particular service.
Provided is a minimal Virtual Service definition that you can extend for multiple matches.
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: external-access
namespace: winegarj
spec:
gateways:
- istio-ingress/winegarj-gateway
hosts:
- winegarj.mids255.com
http:
- match:
- uri:
prefix: "/lab"
route:
- destination:
host: lab-prediction-service
port:
number: 8000
The goal of project is to take everything you have learned in this class and deploy a fully functional prediction API accessible to end users.
You will:
- Utilize
Poetryto define your application dependancies - Package up an existing NLP model (DistilBERT) for running efficient CPU-based sentiment analysis from
HuggingFace - Create a
FastAPIapplication to serve prediction results from user requests - Test your application with
pytest - Utilize
Dockerto package your application as a logic unit of compute - Cache results with
Redisto protect your endpoint from abuse - Deploy your application to
AWSwithKubernetes - Use
k6to load test your applicationk6 run -e NAMESPACE=${NAMESPACE} --summary-trend-stats "min,avg,med,max,p(90),p(95),p(99),p(99.99)" load.js
- Use
Grafanato visualize and understand the dynamics of your system
Below is an outline of the overall architecture:
flowchart TB
subgraph Entire [ ]
subgraph Userland [ ]
User(User)
Developer(Developer)
end
subgraph AWS [AWS]
subgraph Account [Account]
subgraph ECR [Elastic Container Registery]
subgraph repo [Container Repo]
i1(image tag:4f925d7)
i2(image tag:29a727c):::fa
end
end
subgraph k8s [Elastic Kubernetes Service]
subgraph istio [Namespace: istio-ingress]
gw(your-name-gateway)
end
subgraph subgraph_padding1 [ ]
subgraph cn [Namespace: your-name-here]
direction TB
subgraph subgraph_padding2 [ ]
NPS2(ClusterIP: lab-prediction-service):::nodes
NPS3(ClusterIP: project-prediction-service):::nodes
subgraph PD [Lab Deployment]
direction TB
IC1(Init Container: verify-redis-dns)
IC2(Init Container: verify-redis-ready)
FA(Lab FastAPI Container):::fa
IC1 --> IC2 --> FA
end
subgraph ProD [Project Deployment]
direction TB
IC3(Init Container: verify-redis-dns)
IC4(Init Container: verify-redis-ready)
FA2(Project FastAPI Container):::fa
IC3 --> IC4 --> FA2
end
NPS1(ClusterIP: redis-service):::nodes
RD(Redis Deployment)
VS(VirtualService)
VS <--> NPS2
VS <--> NPS3
NPS1 <--->|Port 6379| PD
NPS1 <-->|Port 6379| ProD
NPS1 <-->|Port 6379| RD
NPS2 <-->|Port 8000| PD
NPS3 <-->|Port 8000| ProD
end
end
end
i2 -..- FA
i1 -...- FA2
end
end
end
end
gw <---> User
VS <--> gw
Developer -.->|aws sts get-caller-identity --profile ucberkeley-sso | AWS
Developer -.->|aws sts get-caller-identity --profile ucberkeley-student | Account
Developer -.->|"aws ecr get-login-password --region us-west-2 --profile ucberkeley-student | docker login --username AWS --password-stdin 650251712107.dkr.ecr.us-west-2.amazonaws.com"| ECR
Developer -.->|aws eks update-kubeconfig --name eks-datasci255-students --profile ucberkeley-student | k8s
Developer -->|docker push| repo
classDef nodes fill:#68A063
classDef subgraph_padding fill:none,stroke:none
classDef inits fill:#cc9ef0
classDef fa fill:#00b485
style cn fill:#B6D0E2;
style RD fill:#e6584e;
style PD fill:#FFD43B;
style ProD fill:#FFD43B;
style k8s fill:#b77af4;
style AWS fill:#00aaff;
style Account fill:#ffbf14;
style ECR fill:#cccccc;
style repo fill:#e7e7e7;
style Userland fill:#ffffff,stroke:none;
style Entire fill:#ffffff,stroke:none;
class subgraph_padding1,subgraph_padding2 subgraph_padding
class IC1,IC2,IC3,IC4 inits
-
Write pydantic models to match the specified input model
{ "text": ["example 1", "example 2"] }
-
Write pydantic models to match the specified output model
{ "predictions": [ [ { "label":"POSITIVE", "score":0.7127904295921326 }, { "label":"NEGATIVE", "score":0.2872096002101898 } ], [ { "label":"POSITIVE", "score":0.7186233401298523 }, { "label":"NEGATIVE", "score":0.2813767194747925 } ] ] }
-
Pull the following model locally to allow for loading into your application. Put this at the root of your project directory for an easier time.
- Add the model files to your
.gitignoresince the file is large, and we don't want to managegit-lfsand incur costs for wasted space.HuggingFaceis hosting the model for us.
- Add the model files to your
-
Create and execute
pytesttests to ensure your application is working as intended -
Build and deploy your application locally (Hint: Use
kustomize) -
Push your image to
ECR.- Use a prefix based on your namespace, and call the image
project
- Use a prefix based on your namespace, and call the image
-
Deploy your application to
EKSsimilar to lab 4- Make sure to adjust your virtual service to expose both
projectandlab
- Make sure to adjust your virtual service to expose both
-
Run
k6against your application with the providedload.js -
Capture screenshots of your
grafanadashboard for your service/workload during the execution of yourk6script -
Feel extremely proud about all the learning you went through over the semester and how this will help you develop professionally and enable you to deploy an API effectively during your capstone. There is much to learn, but getting the fundamentals are key.
Please review the train.py to see how the model was trained and pushed to HuggingFace as an artifact store for models and their associated configuration.
This model took 5 minutes to transfer learn on 2x A4000 GPUs with a 256 batch size, taking 15 GB of memory on each GPU.
Training on CPUs would likely have taken several days. The given implementation allows for maximum text sequences of 512 tokens for each input.
Do not try to run the training script on your local machine.
Model loading examples are provided in example.py. In this file, we directly load the model from HuggingFace; however, this is extremely inefficient given the size of the underlying model (256 MB) for a production environment.
We will pull down the model locally as part of our build process.
Model prediction pipelines are included in the transformers API provided by HuggingFace, which dramatically reduces the complexity of the Inferencing application.
An example is provided in mlapi/example.py and is instrumented already in your main.py application.
We provide you with a pytest file, test_mlapi.py, which has the structure of how you should design your pydantic models.
You will have to do some reverse engineering so that your model matches our expectations.
Do not run poetry update it will take a long time due to the handling of torch dependencies.
Do a poetry install instead.
You might need to install git lfs https://git-lfs.github.com/
All code will be graded off your repo's main branch and EKS deployment.
No additional forms or submission processes are needed.
All items are conditional on a 95% cache rate, and after a 10 minute sustained load:
| Criteria | 0% | 50% | 90% | 100% |
|---|---|---|---|---|
| Functional API | No Endpoints Work | Some Endpoints Functional | Most Endpoints Functional | All Criteria Met |
| Caching | No Attempt at Caching | Caching system instantiated but not used | Caching system created but missing some functionality | All Criteria Met |
| Kubernetes Practices | No Attempt at Deployments | Deployments exist but lack key functionality | Kubernetes deployment mostly functional | All Criteria Met |
| Testing | No Testing is done | Minimal amount of testing done. No testing of new endpoints. | Only "happy path" tested and with minimal cases | All Criteria Met |
| Passing Provided Tests | Pydantic model does not adhere to our given pytest file | Pydantic model somewhat passes pytest file | Pydantic model mostly passes pytest file | All Criteria Met |
| Model Loading | Model loads from hugging face on API instantiation | N/A | N/A | Model is loaded into the container at build |
| Predict Endpoint Performance | Endpoint performs at 1 request/second | Endpoint performs at 5 requests/second | Endpoint performs at 9 request/second | Endpoint performs at 10 requests/second |
| Predict Endpoint Latency @ 10 Virtual Users | p(99) < 10 seconds | p(99) < 5 seconds | p(99) < 3 seconds | p(99) < 2 seconds |
This project will take approximately ~10 hours.