Nov 5–6 · Toronto
And the people building <with> them — one room, on the Toronto waterfront.

Every session is a working practitioner walking through what they run in production — architecture, failure modes, and the fix.
Whether you build the frameworks or build on them, you’re in the room with the people solving the problem one layer over.
Cohere, Vector, RBC, Shopify, Wealthsimple and the startups between them — two days on the waterfront.
What you ship on it — products, workflows, and agents in regulated industries. Living with what happens next.
In a world where models are released daily, how do we build an effective eval approach, or know when the right time is to swap models? This session is designed to help engineers move more quickly, build better products, and save on costs

VP AI Engineering, Vector Institute
Running your own AI stack became an urgent question after the temporary deprecation of Anthropic’s Mythos, but not every workload should move, and this track is as much about deciding well as about migrating.

Founder TMLS, MLOps World
There is no settled architecture for deciding what an agent should store, retrieve, revise, and forget. In this track, attendees will leave with a stronger understanding for choosing, implementing, and evaluating memory components in an agent system.

Co-Founder & CTO, Hudson Labs
AI usage continues to grow, yet leaders are struggling to quantify the ROI of their investment. In this track you’ll hear from engineers who have debugged and optimized their costs.

SVP AI & Operations, Wisedocs
AI systems can increasingly self-evaluate and modify their own prompts, tools, memory, code, data, and workflows. This track will focus on ensuring these feedback loops produce genuine, repeatable improvement rather than benchmark overfitting or unstable changes.

Co-Founder & CTO, Hudson Labs
Being more productive than you were before coding agents makes it easy to be fooled into thinking you’re already at best practice. This track covers coding agents on complex codebases, secure AI-assisted development, and the code review bottleneck with concrete case studies from speakers who’ve gone beyond the hype.

Founder TMLS, MLOps World
The field of AI is moving at breakneck speed, and while teams race to capture value, the biggest opportunities often lie with big, bold ideas and creative, outside-the-box approaches

Founder TMLS, MLOps World
Vector Institute
ABOUT THE SPEAKER:
Kathryn Hume leads AI Engineering at the Vector Institute, where she builds tools and platforms that further industry adoption of frontier AI. Prior to joining Vector, she held multiple technology leadership roles in financial services and AI, including Chief Technology Officer for Tangerine Bank, Vice President of Technology for RBC’s mobile, online banking, and direct investing businesses, and Interim Head of Borealis AI, RBC’s AI research institute. Before joining RBC, she held leadership roles at AI startups integrate.ai and Fast Forward Labs (acquired by Cloudera).
Kathryn is a recognized author and public speaker, with work appearing at TED, the Harvard Business Review, and the Globe & Mail. She has served as a guest lecturer and adjunct professor — teaching courses on applied AI and responsible AI — at the Harvard Business School, Stanford, the University of Toronto, and the University of Calgary. She is a mother of two young boys, a board member at CanadaHelps active in the charity sector, and an advocate for women in technology and educational reform in the age of generative AI.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Elastic
ABOUT THE SPEAKER:
Abhimanyu Anand is a Senior Data Scientist at Elastic, where Abhi works on the Agent Builder platform, across areas like context management, agent memory, and self improving agents. Throughout his career, Abhi’s work has spanned productionizing large-scale AI systems, such as agentic applications, chatbots, and recommender systems, for user bases exceeding hundreds of millions of people. Abhi has been involved in NLP for close to 10 years, through his Master’s, research, and industry experience, with his current focus on AI agents, continual learning, and large scale inference.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Most teams building self improving agents follow a similar trajectory: the agent runs, reflects, updates an artifact, and the metric rises. In production, the metric can still rise while a later incident traces back to a learned change that passed review but was evaluated against the wrong signals.
This talk examines different self learning systems that may change the agent context, the agent harness, or model weights. Each offers different gains and failure modes. Memory can absorb sycophancy and inflate its reward signal. Skill libraries can overfit to the last task or be tainted by one bad run. Harness changes can discard useful behaviour as new tools arrive. Weight updates persist longest and are the hardest to reverse.
I’ll cover the controls that make self iteration viable in production: bounded, reversible updates; provenance for every learned artifact; promotion gates; and evaluations that separate improvement from noise.
WHAT YOU’LL LEARN:
Teams are starting to let agents learn from production use, but offline benchmarks and approval workflows do not show whether a learned change will hold up in live systems. This talk gives practitioners a way to decide what agents may change, what evidence is needed before promotion, and how to detect and reverse regressions when that evidence is incomplete.
TD Bank
ABOUT THE SPEAKER:
ML engineering leader bridging cutting-edge research and enterprise-scale production. Currently leading ML Systems & Engineerg group at TD Bank (AI Platform).
I am passionate about building high-performance teams & delivering AI solutions that make a difference. Academic foundation includes a PhD in Machine Learning from UWaterloo and a Postdoctoral Fellowship at Stanford.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Wealthsimple
ABOUT THE SPEAKER:
Diederik van Liere is the Chief Technology Officer at Wealthsimple,Canada’s leading financial innovator serving more than 4 million Canadians and holding approximately $155 billion in assets under administration as of June 30, 2026.. He is a technology leader who has built impactful data science and engineering teams for global brands at scale. He joined Wealthsimple in 2018 as Head of Data Science and Engineering and assumed the role of Chief Technology officer in March 2023. Prior to joining Wealthsimple, Diederik built complex data pipelines and infrastructure as the Director of Data Science at Shopify. He also helped build the data platform at the Wikimedia Foundation (the custodian of Wikipedia and sister projects).
Diederik holds a PhD from the Rotterdam School of Management in the Netherlands, in the area of complex networks and was a Postdoctoral Researcher at Toronto’s Rotman School of Management.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Disco
ABOUT THE SPEAKER:
Wendy is a principal data scientist at Disco, a Toronto based educational technology startup. She has over a decade of experience building data products and teams at scale. Previously, she has held data leadership positions at Shopify, PagerDuty, and Wattpad.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Disco’s MCP server lets learning operators and administrators manage and generate insight about their programs through AI agents instead of dashboards, which means those agents need to be trustworthy, not just capable. This talk walks through how we built our eval process for MCP tools, from defining what ‘correct’ looks like for operator-facing actions to structuring a dataset that catches failures before operators do.
WHAT YOU’LL LEARN:
Most organizations building on MCP are still treating it like a plumbing problem (get the tools connected, ship it) without a rigorous way to know whether the agent is actually doing the right thing when it touches real operator data and workflows. This talk gives practitioners a concrete, repeatable process for building evals that catch agent failures before they reach production, using a real multi-phase MCP rollout as the case study.
Manulife
ABOUT THE SPEAKER:
I’m a Senior Applied AI Scientist at Manulife, where I work as an internal Forward Deployed Engineer — sitting close to the business and closer to the metal — building agentic AI products tied to real business outcomes.
Previously, I owned RLOps from thought to production at Prem Labs for their agent finetuning platform, now live on AWS Marketplace. I’m an early contributor to Hugging Face’s trl library, the most widely used LLM finetuning library in the world.
I completed my Master in Applied Science from Queen’s University under the supervision of Dr. Ali Etemad. My research on Deepfake Detection was partially funded by the Vector Institute Scholarship in AI. Published in NeurIPS and TMLR. Reviewer at NeurIPS and ICASSP 2026. Listen to this podcast to know more.
In the 5pm–9am part of my life, I advise and help traditional businesses become AI-ready.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
This session reframes “Loop Engineering” as applied cybernetics — the same observe→diff→act reconciliation pattern that powers Kubernetes controllers, now applied to AI agent loops. Attendees will learn how to map K8s design principles (declarative goals, idempotent actions, crash-only design, admission webhooks as HITL gates) directly onto agent infrastructure, and walk away with a concrete checklist for building production agent loops that self-heal, externalize state, and fail gracefully. We’ll cover real failure modes — thrashing, stale reads, infinite reconciliation, conflicting controllers — and the fixes that transfer from a decade of K8s production learnings.
WHAT YOU’LL LEARN:
Publicus
ABOUT THE SPEAKER:
Akash Shetty is the co-founder and CTO of Publicus, a Canadian GovTech company building AI-native infrastructure for government procurement.
At Publicus, Akash leads the development of production systems that transform fragmented contracts, tenders, invoices, amendments, task authorizations, webpages, and PDFs into structured, traceable intelligence for businesses and public-sector organizations. His work spans multimodal document processing, retrieval and ranking, agent orchestration, evaluation, data lineage, human review, and the deployment of trustworthy AI workflows in environments involving public money and institutional accountability.
Publicus has completed federal work supporting Buy Canada decision-making and is working on procurement modernization through Innovative Solutions Canada and Public Services and Procurement Canada. Akash approaches AI as a builder shaped by physics and first-principles thinking, with a focus on moving beyond impressive prototypes toward systems that can be verified, operated, and trusted in the real world.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Publicus processes procurement data spread across RFPs, contracts, amendments, award notices, invoices, task authorizations, webpages, and other records. Our early AI workflows optimized for producing a useful answer: retrieve relevant context, reason over it, generate structured output, and evaluate the result. That worked until we tried to increase automation in workflows where an answer can be semantically correct while its evidence is incomplete, stale, contradictory, or simply the wrong source.
We redesigned the system around an evidence ledger rather than the generated answer. Extracted facts and agent claims became typed artifacts carrying source document, location, provenance, transformation history, confidence, and conflicts. Retrieval became evidence acquisition rather than context stuffing. Verification became a separate stage that reconciles claims across records. Unsupported or conflicting outputs cross an explicit human-review boundary instead of being silently converted into fluent prose.
The redesign also changed how we evaluate the system. Instead of relying primarily on answer quality or a single LLM-as-judge score, we measure the pipeline at several boundaries: field-extraction F1, claim-to-source agreement, evidence-grounding accuracy, conflict detection precision/recall, unsupported-output rate, lineage coverage, bilingual performance, latency, cost per document, and human-review burden.
I’ll walk through the production architecture, the failure modes that forced the redesign, where deterministic code replaced model judgment, how we preserve evidence through multi-step agent execution, and how we decide whether a new model or workflow is actually safe to promote.
Attendees will leave with a practical pattern for building agents where correctness means more than producing the right-looking answer: designing typed evidence artifacts, separating retrieval from verification, constructing evals around failure boundaries, and defining human escalation rules for high-stakes AI systems.
raindrop.ai
ABOUT THE SPEAKER:
Manav is an accomplished ML engineer and technical leader with deep experience in evaluations, reinforcement learning, information retrieval, and applied AI systems. At Raindrop AI, he works across machine learning, product, and software engineering, with a focus on monitoring and building AI agents for the real world. He previously led engineering at Dippy AI and Datapoint AI, where he scaled consumer AI systems to 10M+ users.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
To turn the art of building AI agents into a science, you need a way to measure progress — and measuring an agent turns out to be far harder, and far more important, than measuring a model. This talk is about how to actually evaluate agents: what to measure, how the field gets it wrong, and why good measurement is what eventually lets agents improve themselves.
What we’ll explore:
WHAT YOU’LL LEARN:
I have a unique vantage point on how some of the world’s leading AI companies are creating their AI agent lifecycles and developing their agent improvement flywheels.
Google DeepMind
ABOUT THE SPEAKER:
Shashank works as a Research Engineer on the Foundational Research Incubation team at Google DeepMind on post-training systems for foundation models and agents. Previously, he co-founded Dice Health, an AI automation platform for veterinary healthcare providers, and grew it from bootstrapping to acquistion. His research at Meta AI received the NeurIPS 2022 best paper award, at the most prestigious AI research conference. Outside of work, he also runs the Toronto Machine Learning systems group which covers topics like training, inference and deployment of AI systems from consumer devices to large datacenters. He has served as a technical advisor to the Vector Institute and various seed and Series A stage startups.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
WHAT YOU’LL LEARN:
Practitioners can start to think about how they can use self-improving AI agents for tasks relevant to them
Wisedocs
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Model, Data and Objective Drift have traditionally ML Engineering problems that have had a ceiling where we could automate. This talk is a case study of how Wisedocs automated large portions of model development across a 1500+ document types that face standard MLOps challenges. The talk will have three parts, a focus on domain knowledge, the ROI of agentic ML development and sample harnesses and skills we use to enable these capabilities.
WHAT YOU’LL LEARN:
Agentic ML is an evolving form of automating ML workflows while maintaining strong data validation, efficiency and cost consideration
Wisedocs
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Model, Data and Objective Drift have traditionally ML Engineering problems that have had a ceiling where we could automate. This talk is a case study of how Wisedocs automated large portions of model development across a 1500+ document types that face standard MLOps challenges. The talk will have three parts, a focus on domain knowledge, the ROI of agentic ML development and sample harnesses and skills we use to enable these capabilities.
WHAT YOU’LL LEARN:
Agentic ML is an evolving form of automating ML workflows while maintaining strong data validation, efficiency and cost consideration
Wisedocs
ABOUT THE SPEAKER:
Afseen Syeda is a Lead Prompt Engineer at Wisedocs, where she works alongside the machine learning team to design and refine prompts that make AI more effective at processing and reasoning over complex medical documents. Her focus is getting the most out of AI models through thoughtful prompting, from extraction accuracy to workflow efficiency.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Managing 10,000+ prompts across dozens of use cases is a challenging task when new LLMs get released weekly. This talk covers the infrastructure, automation and LLM based prompt that allows Wisedocs to automatically update prompts with new product releases, classify error categories and empower prompt engineers to focus on emerging, challenging problems, rather than maintanance.
WHAT YOU’LL LEARN:
Prompts are a key
Wisedocs
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Managing 10,000+ prompts across dozens of use cases is a challenging task when new LLMs get released weekly. This talk covers the infrastructure, automation and LLM based prompt that allows Wisedocs to automatically update prompts with new product releases, classify error categories and empower prompt engineers to focus on emerging, challenging problems, rather than maintanance.
WHAT YOU’LL LEARN:
Prompts are a key
Hyperdimensional Computing Labs
ABOUT THE SPEAKER:
Prashanth works at the intersection of data infrastructure, graphs, and agentic systems. He’s constantly thinking about how graph and vector representations can be combined so that agents retrieve better context than either is able to provide on its own.
Prashanth regularly builds benchmarks and end-to-end case studies that put the tradeoffs between tools and techniques across different paradigms on the record, enabling architectural decisions to be based on tangible evidence rather than claims.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
What if AI systems could learn continuously without repeatedly retraining the model underneath them? At HDC Labs, we are exploring hyperdimensional computing as a representation layer for online, associative learning. Neural models provide powerful learned representations and general capabilities; hyperdimensional representations provide a lightweight structure in which new observations, relationships, corrections, and outcomes can be incorporated as they arrive. Because learning can happen through simple composable operations over these representations, useful behavior can emerge from relatively few examples rather than large retraining datasets.
That same compositional structure also creates an opportunity for greater explainability. Instead of treating every update as an opaque change to model weights, hyperdimensional representations can preserve the primitives and associations that contributed to a result, making retrieval and reasoning easier to inspect. In this talk, we will show how HDC can complement modern ML systems with three capabilities that matter increasingly in the agentic age: explainable representations, continuous online learning, and sample-efficient adaptation.
WHAT YOU’LL LEARN:
Practitioners increasingly need AI systems that can incorporate new information after deployment without continuously retraining or fine-tuning large models. Hyperdimensional computing offers a lightweight representation layer where new examples and associations can be learned online, often from very little data, while keeping the underlying operations composable and inspectable. This can make deployed systems cheaper to update, faster to adapt, and easier to understand, .especially in agentic, edge, and rapidly changing environments.
Third Bit
ABOUT THE SPEAKER:
Dr. Greg Wilson is a programmer, author, and educator based in Toronto. He co-founded Software Carpentry, which has taught software skills to tens of thousands of researchers worldwide, and has authored or edited over a dozen books. Greg is a Fellow of the Python Software Foundation and a recipient of ACM SIGSOFT’s Influential Educator of the Year award.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Most organizations don’t know how to assess the productivity of their software developers, which means that many claims about the impact of AI on productivity are vacuous. This talk analyzes some common mistakes, and describes a few things companies can do to figure out what is and isn’t actually cost-effective.
WHAT YOU’LL LEARN:
It will help them avoid the “simple but wrong” trap when it comes to assessing AI tooling and practices.
Tangerine
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
WHAT YOU’LL LEARN:
They will see what a leading digital bank is doing to bring in Evals and merge with TDD practices
Vector institute
ABOUT THE SPEAKER:
Dr. Shaina Raza is an Applied Machine Learning Scientist, Responsible AI, at the Vector Institute, specializing in Responsible AI, multimodal and agentic AI evaluation, fairness, AI safety, and sustainable AI. She leads and contributes to research on trustworthy AI evaluation frameworks, deepfake detection, bias mitigation, and responsible deployment of advanced AI systems. Her work spans both academic research and industry-facing initiatives, including Horizon Europe projects and open-source Responsible AI tools. She has been recognized as a Responsible AI Leader of the Year 2025 and an Emerging Leader 2026, and was named among Elsevier’s Top 2% Scientists in 2025.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
WHAT YOU’LL LEARN:
Practitioners will learn how to evaluate AI systems beyond accuracy—covering fairness, safety, robustness, explainability, and sustainability—so they can identify risks earlier, choose more efficient evaluation strategies, and make better deployment decisions.
Rakuten Kobo
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
When your BI team has 1000+ workbooks and a growing request backlog, the answer isn’t more analysts — it’s rethinking the architecture. In this talk we share how Rakuten Kobo’s Data team replaced a centralized BI bottleneck with a code-first model where dashboards are version-controlled files authored end-to-end by AI. We cover the [long-term] strategy, tooling decision, the AI-guided authoring workflow built on Claude Code and Cloud, the real operating boundaries discovered through production-scale data. Attendees will leave with a concrete architecture pattern, an honest account of where the approach has limits, and a staged AI coding model they can adapt to any BI environment.
WHAT YOU’LL LEARN:
We are on the forefront of this shift from BI tools to open source (with the option to still buy commercial license) and this success story can empower and encourage others to follow the suit and start transforming their organizations. This is very practical presentation where companies can mimmic in a few weeks.
CIBC
ABOUT THE SPEAKER:
Head of AI Research at CIBC, with twelve years at the financial sector spanning every stage of the path research must travel to become value. Erin set her research priorities with business accountability, production relevancy, and clear governance, so frontier work is scoped from the outset against what institutions can operationalize, govern and monetize.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Most AI benchmarks are optimized for public datasets and generic tasks, but enterprise teams operate in a very different reality: domain-specific workflows, regulated data, multimodal inputs, and business-critical quality thresholds. This talk presents the CIBC Benchmark, an enterprise evaluation framework designed to test AI models and workflows against real enterprise tasks, enterprise-relevant datasets, and standardized evaluation metrics.
WHAT YOU’LL LEARN:
Staff ML Engineer
Cohere
Principal Engineer
Shopify
CIBC
ABOUT THE SPEAKER:
Head of AI Research at CIBC, with twelve years at the financial sector spanning every stage of the path research must travel to become value. Erin set her research priorities with business accountability, production relevancy, and clear governance, so frontier work is scoped from the outset against what institutions can operationalize, govern and monetize.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Most AI benchmarks are optimized for public datasets and generic tasks, but enterprise teams operate in a very different reality: domain-specific workflows, regulated data, multimodal inputs, and business-critical quality thresholds. This talk presents the CIBC Benchmark, an enterprise evaluation framework designed to test AI models and workflows against real enterprise tasks, enterprise-relevant datasets, and standardized evaluation metrics.
WHAT YOU’LL LEARN:
Rakuten Kobo
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
When your BI team has 1000+ workbooks and a growing request backlog, the answer isn’t more analysts — it’s rethinking the architecture. In this talk we share how Rakuten Kobo’s Data team replaced a centralized BI bottleneck with a code-first model where dashboards are version-controlled files authored end-to-end by AI. We cover the [long-term] strategy, tooling decision, the AI-guided authoring workflow built on Claude Code and Cloud, the real operating boundaries discovered through production-scale data. Attendees will leave with a concrete architecture pattern, an honest account of where the approach has limits, and a staged AI coding model they can adapt to any BI environment.
WHAT YOU’LL LEARN:
We are on the forefront of this shift from BI tools to open source (with the option to still buy commercial license) and this success story can empower and encourage others to follow the suit and start transforming their organizations. This is very practical presentation where companies can mimmic in a few weeks.
Vector institute
ABOUT THE SPEAKER:
Dr. Shaina Raza is an Applied Machine Learning Scientist, Responsible AI, at the Vector Institute, specializing in Responsible AI, multimodal and agentic AI evaluation, fairness, AI safety, and sustainable AI. She leads and contributes to research on trustworthy AI evaluation frameworks, deepfake detection, bias mitigation, and responsible deployment of advanced AI systems. Her work spans both academic research and industry-facing initiatives, including Horizon Europe projects and open-source Responsible AI tools. She has been recognized as a Responsible AI Leader of the Year 2025 and an Emerging Leader 2026, and was named among Elsevier’s Top 2% Scientists in 2025.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
WHAT YOU’LL LEARN:
Practitioners will learn how to evaluate AI systems beyond accuracy—covering fairness, safety, robustness, explainability, and sustainability—so they can identify risks earlier, choose more efficient evaluation strategies, and make better deployment decisions.
Tangerine
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
WHAT YOU’LL LEARN:
They will see what a leading digital bank is doing to bring in Evals and merge with TDD practices
Third Bit
ABOUT THE SPEAKER:
Dr. Greg Wilson is a programmer, author, and educator based in Toronto. He co-founded Software Carpentry, which has taught software skills to tens of thousands of researchers worldwide, and has authored or edited over a dozen books. Greg is a Fellow of the Python Software Foundation and a recipient of ACM SIGSOFT’s Influential Educator of the Year award.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Most organizations don’t know how to assess the productivity of their software developers, which means that many claims about the impact of AI on productivity are vacuous. This talk analyzes some common mistakes, and describes a few things companies can do to figure out what is and isn’t actually cost-effective.
WHAT YOU’LL LEARN:
It will help them avoid the “simple but wrong” trap when it comes to assessing AI tooling and practices.
Wisedocs
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Managing 10,000+ prompts across dozens of use cases is a challenging task when new LLMs get released weekly. This talk covers the infrastructure, automation and LLM based prompt that allows Wisedocs to automatically update prompts with new product releases, classify error categories and empower prompt engineers to focus on emerging, challenging problems, rather than maintanance.
WHAT YOU’LL LEARN:
Prompts are a key
Wisedocs
ABOUT THE SPEAKER:
Afseen Syeda is a Lead Prompt Engineer at Wisedocs, where she works alongside the machine learning team to design and refine prompts that make AI more effective at processing and reasoning over complex medical documents. Her focus is getting the most out of AI models through thoughtful prompting, from extraction accuracy to workflow efficiency.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Managing 10,000+ prompts across dozens of use cases is a challenging task when new LLMs get released weekly. This talk covers the infrastructure, automation and LLM based prompt that allows Wisedocs to automatically update prompts with new product releases, classify error categories and empower prompt engineers to focus on emerging, challenging problems, rather than maintanance.
WHAT YOU’LL LEARN:
Prompts are a key
Wisedocs
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Model, Data and Objective Drift have traditionally ML Engineering problems that have had a ceiling where we could automate. This talk is a case study of how Wisedocs automated large portions of model development across a 1500+ document types that face standard MLOps challenges. The talk will have three parts, a focus on domain knowledge, the ROI of agentic ML development and sample harnesses and skills we use to enable these capabilities.
WHAT YOU’LL LEARN:
Agentic ML is an evolving form of automating ML workflows while maintaining strong data validation, efficiency and cost consideration
Wisedocs
ABOUT THE SPEAKER:
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Model, Data and Objective Drift have traditionally ML Engineering problems that have had a ceiling where we could automate. This talk is a case study of how Wisedocs automated large portions of model development across a 1500+ document types that face standard MLOps challenges. The talk will have three parts, a focus on domain knowledge, the ROI of agentic ML development and sample harnesses and skills we use to enable these capabilities.
WHAT YOU’LL LEARN:
Agentic ML is an evolving form of automating ML workflows while maintaining strong data validation, efficiency and cost consideration
Disco
ABOUT THE SPEAKER:
Wendy is a principal data scientist at Disco, a Toronto based educational technology startup. She has over a decade of experience building data products and teams at scale. Previously, she has held data leadership positions at Shopify, PagerDuty, and Wattpad.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Disco’s MCP server lets learning operators and administrators manage and generate insight about their programs through AI agents instead of dashboards, which means those agents need to be trustworthy, not just capable. This talk walks through how we built our eval process for MCP tools, from defining what ‘correct’ looks like for operator-facing actions to structuring a dataset that catches failures before operators do.
WHAT YOU’LL LEARN:
Most organizations building on MCP are still treating it like a plumbing problem (get the tools connected, ship it) without a rigorous way to know whether the agent is actually doing the right thing when it touches real operator data and workflows. This talk gives practitioners a concrete, repeatable process for building evals that catch agent failures before they reach production, using a real multi-phase MCP rollout as the case study.
Wealthsimple
ABOUT THE SPEAKER:
Diederik van Liere is the Chief Technology Officer at Wealthsimple,Canada’s leading financial innovator serving more than 4 million Canadians and holding approximately $155 billion in assets under administration as of June 30, 2026.. He is a technology leader who has built impactful data science and engineering teams for global brands at scale. He joined Wealthsimple in 2018 as Head of Data Science and Engineering and assumed the role of Chief Technology officer in March 2023. Prior to joining Wealthsimple, Diederik built complex data pipelines and infrastructure as the Director of Data Science at Shopify. He also helped build the data platform at the Wikimedia Foundation (the custodian of Wikipedia and sister projects).
Diederik holds a PhD from the Rotterdam School of Management in the Netherlands, in the area of complex networks and was a Postdoctoral Researcher at Toronto’s Rotman School of Management.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
TD Bank
ABOUT THE SPEAKER:
ML engineering leader bridging cutting-edge research and enterprise-scale production. Currently leading ML Systems & Engineerg group at TD Bank (AI Platform).
I am passionate about building high-performance teams & delivering AI solutions that make a difference. Academic foundation includes a PhD in Machine Learning from UWaterloo and a Postdoctoral Fellowship at Stanford.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Hyperdimensional Computing Labs
ABOUT THE SPEAKER:
Prashanth works at the intersection of data infrastructure, graphs, and agentic systems. He’s constantly thinking about how graph and vector representations can be combined so that agents retrieve better context than either is able to provide on its own.
Prashanth regularly builds benchmarks and end-to-end case studies that put the tradeoffs between tools and techniques across different paradigms on the record, enabling architectural decisions to be based on tangible evidence rather than claims.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
What if AI systems could learn continuously without repeatedly retraining the model underneath them? At HDC Labs, we are exploring hyperdimensional computing as a representation layer for online, associative learning. Neural models provide powerful learned representations and general capabilities; hyperdimensional representations provide a lightweight structure in which new observations, relationships, corrections, and outcomes can be incorporated as they arrive. Because learning can happen through simple composable operations over these representations, useful behavior can emerge from relatively few examples rather than large retraining datasets.
That same compositional structure also creates an opportunity for greater explainability. Instead of treating every update as an opaque change to model weights, hyperdimensional representations can preserve the primitives and associations that contributed to a result, making retrieval and reasoning easier to inspect. In this talk, we will show how HDC can complement modern ML systems with three capabilities that matter increasingly in the agentic age: explainable representations, continuous online learning, and sample-efficient adaptation.
WHAT YOU’LL LEARN:
Practitioners increasingly need AI systems that can incorporate new information after deployment without continuously retraining or fine-tuning large models. Hyperdimensional computing offers a lightweight representation layer where new examples and associations can be learned online, often from very little data, while keeping the underlying operations composable and inspectable. This can make deployed systems cheaper to update, faster to adapt, and easier to understand, .especially in agentic, edge, and rapidly changing environments.
Google DeepMind
ABOUT THE SPEAKER:
Shashank works as a Research Engineer on the Foundational Research Incubation team at Google DeepMind on post-training systems for foundation models and agents. Previously, he co-founded Dice Health, an AI automation platform for veterinary healthcare providers, and grew it from bootstrapping to acquistion. His research at Meta AI received the NeurIPS 2022 best paper award, at the most prestigious AI research conference. Outside of work, he also runs the Toronto Machine Learning systems group which covers topics like training, inference and deployment of AI systems from consumer devices to large datacenters. He has served as a technical advisor to the Vector Institute and various seed and Series A stage startups.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
WHAT YOU’LL LEARN:
Practitioners can start to think about how they can use self-improving AI agents for tasks relevant to them
raindrop.ai
ABOUT THE SPEAKER:
Manav is an accomplished ML engineer and technical leader with deep experience in evaluations, reinforcement learning, information retrieval, and applied AI systems. At Raindrop AI, he works across machine learning, product, and software engineering, with a focus on monitoring and building AI agents for the real world. He previously led engineering at Dippy AI and Datapoint AI, where he scaled consumer AI systems to 10M+ users.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
To turn the art of building AI agents into a science, you need a way to measure progress — and measuring an agent turns out to be far harder, and far more important, than measuring a model. This talk is about how to actually evaluate agents: what to measure, how the field gets it wrong, and why good measurement is what eventually lets agents improve themselves.
What we’ll explore:
WHAT YOU’LL LEARN:
I have a unique vantage point on how some of the world’s leading AI companies are creating their AI agent lifecycles and developing their agent improvement flywheels.
Publicus
ABOUT THE SPEAKER:
Akash Shetty is the co-founder and CTO of Publicus, a Canadian GovTech company building AI-native infrastructure for government procurement.
At Publicus, Akash leads the development of production systems that transform fragmented contracts, tenders, invoices, amendments, task authorizations, webpages, and PDFs into structured, traceable intelligence for businesses and public-sector organizations. His work spans multimodal document processing, retrieval and ranking, agent orchestration, evaluation, data lineage, human review, and the deployment of trustworthy AI workflows in environments involving public money and institutional accountability.
Publicus has completed federal work supporting Buy Canada decision-making and is working on procurement modernization through Innovative Solutions Canada and Public Services and Procurement Canada. Akash approaches AI as a builder shaped by physics and first-principles thinking, with a focus on moving beyond impressive prototypes toward systems that can be verified, operated, and trusted in the real world.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Publicus processes procurement data spread across RFPs, contracts, amendments, award notices, invoices, task authorizations, webpages, and other records. Our early AI workflows optimized for producing a useful answer: retrieve relevant context, reason over it, generate structured output, and evaluate the result. That worked until we tried to increase automation in workflows where an answer can be semantically correct while its evidence is incomplete, stale, contradictory, or simply the wrong source.
We redesigned the system around an evidence ledger rather than the generated answer. Extracted facts and agent claims became typed artifacts carrying source document, location, provenance, transformation history, confidence, and conflicts. Retrieval became evidence acquisition rather than context stuffing. Verification became a separate stage that reconciles claims across records. Unsupported or conflicting outputs cross an explicit human-review boundary instead of being silently converted into fluent prose.
The redesign also changed how we evaluate the system. Instead of relying primarily on answer quality or a single LLM-as-judge score, we measure the pipeline at several boundaries: field-extraction F1, claim-to-source agreement, evidence-grounding accuracy, conflict detection precision/recall, unsupported-output rate, lineage coverage, bilingual performance, latency, cost per document, and human-review burden.
I’ll walk through the production architecture, the failure modes that forced the redesign, where deterministic code replaced model judgment, how we preserve evidence through multi-step agent execution, and how we decide whether a new model or workflow is actually safe to promote.
Attendees will leave with a practical pattern for building agents where correctness means more than producing the right-looking answer: designing typed evidence artifacts, separating retrieval from verification, constructing evals around failure boundaries, and defining human escalation rules for high-stakes AI systems.
Manulife
ABOUT THE SPEAKER:
I’m a Senior Applied AI Scientist at Manulife, where I work as an internal Forward Deployed Engineer — sitting close to the business and closer to the metal — building agentic AI products tied to real business outcomes.
Previously, I owned RLOps from thought to production at Prem Labs for their agent finetuning platform, now live on AWS Marketplace. I’m an early contributor to Hugging Face’s trl library, the most widely used LLM finetuning library in the world.
I completed my Master in Applied Science from Queen’s University under the supervision of Dr. Ali Etemad. My research on Deepfake Detection was partially funded by the Vector Institute Scholarship in AI. Published in NeurIPS and TMLR. Reviewer at NeurIPS and ICASSP 2026. Listen to this podcast to know more.
In the 5pm–9am part of my life, I advise and help traditional businesses become AI-ready.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
This session reframes “Loop Engineering” as applied cybernetics — the same observe→diff→act reconciliation pattern that powers Kubernetes controllers, now applied to AI agent loops. Attendees will learn how to map K8s design principles (declarative goals, idempotent actions, crash-only design, admission webhooks as HITL gates) directly onto agent infrastructure, and walk away with a concrete checklist for building production agent loops that self-heal, externalize state, and fail gracefully. We’ll cover real failure modes — thrashing, stale reads, infinite reconciliation, conflicting controllers — and the fixes that transfer from a decade of K8s production learnings.
WHAT YOU’LL LEARN:
Elastic
ABOUT THE SPEAKER:
Abhimanyu Anand is a Senior Data Scientist at Elastic, where Abhi works on the Agent Builder platform, across areas like context management, agent memory, and self improving agents. Throughout his career, Abhi’s work has spanned productionizing large-scale AI systems, such as agentic applications, chatbots, and recommender systems, for user bases exceeding hundreds of millions of people. Abhi has been involved in NLP for close to 10 years, through his Master’s, research, and industry experience, with his current focus on AI agents, continual learning, and large scale inference.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
Most teams building self improving agents follow a similar trajectory: the agent runs, reflects, updates an artifact, and the metric rises. In production, the metric can still rise while a later incident traces back to a learned change that passed review but was evaluated against the wrong signals.
This talk examines different self learning systems that may change the agent context, the agent harness, or model weights. Each offers different gains and failure modes. Memory can absorb sycophancy and inflate its reward signal. Skill libraries can overfit to the last task or be tainted by one bad run. Harness changes can discard useful behaviour as new tools arrive. Weight updates persist longest and are the hardest to reverse.
I’ll cover the controls that make self iteration viable in production: bounded, reversible updates; provenance for every learned artifact; promotion gates; and evaluations that separate improvement from noise.
WHAT YOU’LL LEARN:
Teams are starting to let agents learn from production use, but offline benchmarks and approval workflows do not show whether a learned change will hold up in live systems. This talk gives practitioners a way to decide what agents may change, what evidence is needed before promotion, and how to detect and reverse regressions when that evidence is incomplete.
Vector Institute
ABOUT THE SPEAKER:
Kathryn Hume leads AI Engineering at the Vector Institute, where she builds tools and platforms that further industry adoption of frontier AI. Prior to joining Vector, she held multiple technology leadership roles in financial services and AI, including Chief Technology Officer for Tangerine Bank, Vice President of Technology for RBC’s mobile, online banking, and direct investing businesses, and Interim Head of Borealis AI, RBC’s AI research institute. Before joining RBC, she held leadership roles at AI startups integrate.ai and Fast Forward Labs (acquired by Cloudera).
Kathryn is a recognized author and public speaker, with work appearing at TED, the Harvard Business Review, and the Globe & Mail. She has served as a guest lecturer and adjunct professor — teaching courses on applied AI and responsible AI — at the Harvard Business School, Stanford, the University of Toronto, and the University of Calgary. She is a mother of two young boys, a board member at CanadaHelps active in the charity sector, and an advocate for women in technology and educational reform in the age of generative AI.
TALK TITLE:
TRACK:
SUB TOPIC:
ABSTRACT:
08:30
09:45 - 10:25
10:25 - 10:55
10:55 - 10:40
Cost Management & ROI · Agent Deployment & Observability
11:45 - 12:30
Cost Management & ROI · Agent Deployment & Observability ·
Coding Agents & Dev Workflows
12:30 - 1:30 PM
1:30 - 2:15
2:20 - 2:50
2:55 - 3:25
3:25 - 3:55
3:55 - 4:25
4:30 - 5:00
8:45 AM - 9:45 AM
9:45 - 10:25
10:25 -10:55
10:55 - 11:40
Coding Agents & Dev Workflows · Evals & Benchmark
11:45 - 12:30
Recursive Self-Improvement (RSI) · AI Sovereignty
Evals & Benchmarks
12:30 - 1:30
1:30 - 2:00
Recursive Self-Improvement (RSI) · Evals & Benchmarks
2:05 - 2:35
Recursive Self-Improvement (RSI) · Evals & Benchmarks
ML & AI practitioners across two days
in senior, lead, or leadership roles
speakers from teams shipping agents now
peer-curated tracks, zero intro talks
All prices in CAD. Both days, all tracks, meals.
Plus taxes & fees
Total $471.90
Plus taxes & fees
Total $424.67 / seat
Plus taxes & fees
Total $412.86
Talk selection is peer-reviewed by the practitioner committee. We want the debugging, the scaling, the war stories — not the pitch deck. Submissions close September 1st.
MLOps North is hosted at RBC WaterPark Place.
Expect deep sessions on running agents inside a regulated enterprise.
Gold Sponsor
Gold Sponsor
Gold Sponsor
88 Queens Quay W, Toronto, ON M5J 0B6. Minutes from Union Station and the core.
Getting there
8-minute walk from Union Station (GO, UP Express, subway). Billy Bishop airport is a 15-min cab.
“Building Agents” is the stack — frameworks, memory, evals, orchestration. “Building with Agents” is what you ship on it — products, workflows, agents in regulated industries.
Yes — the call for proposals is open until September 1st and is peer-reviewed by the practitioner committee. Sponsorship tiers run from Startup Zone to Gold; a few spots remain.
Business Leaders: C-Level Executives, Project Managers, and Product Owners will get to explore best practices, methodologies, principles, and practices for achieving ROI.
Engineers, Researchers, Data Practitioners: Will get a better understanding of the challenges, solutions, and ideas being offered via breakouts & workshops on Natural Language Processing, Neural Nets, Reinforcement Learning, Generative Adversarial Networks (GANs), Evolution Strategies, AutoML, and more.
Job Seekers: Will have the opportunity to network virtually and meet over 30+ Top Al Companies.
Ignite what is an Ignite Talk?
Ignite is an innovative and fast-paced style used to deliver a concise presentation.
During an Ignite Talk, presenters discuss their research using 20 image-centric slides which automatically advance every 15 seconds.
The result is a fun and engaging five-minute presentation.
You can see all our speakers and full agenda here