Migrating from arvados-node-manager to crunch-dispatch-cloud¶
- Table of contents
- Migrating from arvados-node-manager to crunch-dispatch-cloud
Choose a node¶
The dispatch service can run on any host that can connect to the Arvados API service, the cloud provider's API, and the SSH service on cloud VMs. In the following example it runs on the same node as the API server and controller.
Update cluster configuration file¶
/etc/arvados/config.yml, add configuration items for the dispatch service.
Clusters: uuid_prefix: CloudVMs: BootProbeCommand: "mount | grep /mnt/scratch" SSHPort: "2222" SyncInterval: 1m TimeoutIdle: 2m TimeoutBooting: 10m TimeoutProbe: 5m TimeoutShutdown: 30s ImageID: "image-12345678" Driver: Azure DriverParameters: SubscriptionID: XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX subscription_id: XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX # not needed after #14745 ClientID: XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX key: XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX # not needed after #14745 (same value as ClientID) ClientSecret: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX secret: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX # not needed after #14745 (same value as ClientSecret) TenantID: XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX tenant_id: XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX # not needed after #14745 CloudEnv: AzurePublicCloud cloud_environment: AzurePublicCloud # not needed after #14745 ResourceGroup: zzzzz resource_group: zzzzz Location: centralus region: centralus # not needed after #14745 (same value as Location) Network: zzzzz Subnet: zzzzz-subnet-private StorageAccount: example storage_account: example # not needed after #14745 BlobContainer: vhds blob_container: vhds # not needed after #14745 DeleteDanglingResourcesAfter: 20 delete_dangling_resources_after: 20 # not needed after #14745 Dispatch: PrivateKey: "..." StaleLockTimeout: 1m PollInterval: 10s ProbeInterval: 10s MaxProbesPerSecond: 10 InstanceTypes: x1lg: ProviderType: x1.large VCPUs: 16 RAM: 128G Scratch: 128G Price: 1.23 ManagementToken: "example-secret-management-token" NodeProfiles: apiserver: # references ARVADOS_NODE_PROFILE in environment file (see below). arvados-dispatch-cloud: Listen: ":9005"
Create the host configuration file
Stop and disable the crunch-dispatch-slurm service, and uninstall the package to make sure it doesn't start after the next reboot/upgrade.
# systemctl stop crunch-dispatch-slurm # systemctl disable crunch-dispatch-slurm # apt-get remove crunch-dispatch-slurm
Containers that have already been locked and submitted to SLURM will make their way through the SLURM queue, but newly queued containers will be left for crunch-dispatch-cloud to run.
# apt-get install crunch-dispatch-cloud
Verify the service is running¶
$ token="example-secret-management-token" $ curl -H "Authorization: Bearer $token" http://localhost:9005/metrics