Hey, I’m Kenneth. # Software engineer. AI in the loop. I’m a software engineer at Siemens EDA. I build software that maps large chip designs onto programmable hardware (FPGAs), so teams can test them before manufacturing the chips. I also use AI for development and build agent skills and MCP integrations. [My work](https://kennethassogba.github.io/#work) Sceaux, France ## At Siemens EDA [Veloce proFPGA CS](https://blogs.sw.siemens.com/hardware-assisted-verification/2024/10/16/fpga-based-prototyping-from-do-it-yourself-to-an-essential-soc-verification-and-system-validation-tool/) I work on the compiler for FPGA prototyping, which maps a chip design across multiple FPGAs. ### Placement & partitioning I develop placement and partitioning algorithms that map a chip’s netlist across multiple FPGAs. ### Netlist qualification I’ve recently started working on netlist qualification, including clock handling. ### Performance optimization I redesigned circuit replication to run over 4× faster on production designs, and optimized graph pruning with speedups of up to 60× on large circuits. ### AI-assisted development I write agent skills that explain our codebase and engineering practices, and integrate MCP servers into the development workflow. ## After hours Team project ### La Bulle A voice coaching app with an editable recap and optional Notion follow-up. Voice OpenAI Cloudflare [Read the build note](https://kennethassogba.github.io/notes/building-la-bulle) [Open La Bulle](https://bulle.hodge-podge.workers.dev/?lang=en) Personal project ### This website I’m making this website AI-native, following Cloudflare’s recommendations: Markdown pages, content discovery, and WebMCP tools for agents to search and read it. For you **.html** For your agent **.md** shadcn/ui Markdown WebMCP [Read the build note](https://kennethassogba.github.io/notes/a-website-for-people-and-agents) [For agents](https://kennethassogba.github.io/agents.html) [ **cmake2graph** A Python tool for visualizing CMake target dependencies. ](https://github.com/kennethassogba/cmake2graph) ## Notes [All notes](https://kennethassogba.github.io/notes.html) - [ ### Making this website AI-native ](https://kennethassogba.github.io/notes/a-website-for-people-and-agents) Markdown pages, WebMCP tools, and content discovery based on Cloudflare’s recommendations. Oct 2026 2 min read - [ ### Agent skills for a large codebase ](https://kennethassogba.github.io/notes/repository-knowledge-as-agent-skills) Documenting the codebase and engineering practices for coding agents. Oct 2026 2 min read - [ ### Building La Bulle ](https://kennethassogba.github.io/notes/building-la-bulle) Scaffolding a voice coaching app with Codex, OpenAI Realtime, Cloudflare Workers, D1, and Resend. Sep 2026 6 min read Contact [Say hello](mailto:kennethassogba@gmail.com) Canonical: https://kennethassogba.github.io/ --- About # Hi, I’m Kenneth. ![Kenneth Assogba](https://kennethassogba.github.io/assets/img/me.jpg) I’m a software engineer at Siemens EDA. I build software that maps large chip designs onto programmable hardware (FPGAs), so teams can test them before manufacturing the chips. I also use AI for development and build agent skills and MCP integrations. ## FPGA prototyping I work on the compiler for FPGA prototyping, which maps a chip design across multiple FPGAs. I develop placement and partitioning algorithms that map a chip’s netlist across multiple FPGAs. I’ve recently started working on netlist qualification, including clock handling. I redesigned circuit replication to run over 4× faster on production designs, and optimized graph pruning with speedups of up to 60× on large circuits. [Siemens’ overview of Veloce proFPGA CS](https://blogs.sw.siemens.com/hardware-assisted-verification/2024/10/16/fpga-based-prototyping-from-do-it-yourself-to-an-essential-soc-verification-and-system-validation-tool/) explains the prototyping platform. ## AI-assisted development I write agent skills that explain our codebase and engineering practices, and integrate MCP servers into the development workflow. ## Experience Before Siemens, I built simulation software at CEA during my PhD in Applied Mathematics at École polytechnique. I grew up in Benin and now live in Sceaux, France. - 2023 — now: **Siemens EDA** FPGA prototyping · Placement & partitioning · C++ - 2020 — 2023: **CEA / École polytechnique** Simulation software · C++ · MPI & OpenMP · PhD - 2020: **Total** Wave propagation simulation and Python tooling ## Publications From my research at CEA. - [The Pn form of the Neutron Transport Problem Achieves Linear Scalability Through Domain Decomposition](https://kennethassogba.github.io/publications/mc-2023) · Accepted · 2023 - [Spherical Harmonics and Discontinuous Galerkin Finite Element Methods for the Three Dimensional Neutron Transport Equation: Application to Core and Lattice Calculation](https://kennethassogba.github.io/publications/nse-2023) · Nuclear Science and Engineering · 2023 - [Precise 3D Reactor Core Calculation Using Spherical Harmonics and Discontinuous Galerkin Finite Element Methods](https://kennethassogba.github.io/publications/physor-2022) · Proceedings of International Conference on Physics of Reactors (PHYSOR) · 2022 [Google Scholar](https://scholar.google.com/citations?user=zumTckUAAAAJ) · [ORCID](https://orcid.org/0000-0002-0635-7508) [Read as Markdown](https://kennethassogba.github.io/about/index.md) Canonical: https://kennethassogba.github.io/about.html --- # Notes AI-assisted development, developer tools, C++, and scientific computing. - [ ### Making this website AI-native ](https://kennethassogba.github.io/notes/a-website-for-people-and-agents) Markdown pages, WebMCP tools, and content discovery based on Cloudflare’s recommendations. Oct 2026 2 min read - [ ### Agent skills for a large codebase ](https://kennethassogba.github.io/notes/repository-knowledge-as-agent-skills) Documenting the codebase and engineering practices for coding agents. Oct 2026 2 min read - [ ### Building La Bulle ](https://kennethassogba.github.io/notes/building-la-bulle) Scaffolding a voice coaching app with Codex, OpenAI Realtime, Cloudflare Workers, D1, and Resend. Sep 2026 6 min read - [ ### Developing cmake2graph ](https://kennethassogba.github.io/notes/cmake2graph) CMake Dependency Visualization. Mar 2025 2 min read - [ ### Write a header-only object oriented interface around MPI ](https://kennethassogba.github.io/notes/write-interface-mpi) Write interface around MPI. May 2023 3 min read - [ ### Async communications are effective ](https://kennethassogba.github.io/notes/async-communications) Async communications. Apr 2023 2 min read Draft - [ ### Send an object via MPI using serialization ](https://kennethassogba.github.io/notes/mpi-using-serialization) Send an object via MPI using serialization. Mar 2023 1 min read [Subscribe via RSS](https://kennethassogba.github.io/feed.xml) [Read as Markdown](https://kennethassogba.github.io/notes/index.md) Canonical: https://kennethassogba.github.io/notes.html --- # For agents Markdown pages, content discovery, and WebMCP tools. Every page is available as Markdown. You can also read the content index or use the browser tools below. - [Start here](https://kennethassogba.github.io/llms.txt) `/llms.txt`: A concise map of this site. - [Everything in Markdown](https://kennethassogba.github.io/llms-full.txt) `/llms-full.txt`: Profile, projects, notes, and publications in one file. - [Content index](https://kennethassogba.github.io/api/content.json) `/api/content.json`: Titles, topics, dates, URLs, and plain text for local search. - [Profile](https://kennethassogba.github.io/api/profile.json) `/api/profile.json`: The same professional information you see on the website. - [API catalog](https://kennethassogba.github.io/.well-known/api-catalog) `/.well-known/api-catalog`: Read-only APIs and their OpenAPI description. - [RSS feed](https://kennethassogba.github.io/feed.xml) `/feed.xml`: Follow new writing without scraping the page. ## Content Signals I allow search, AI input, and model training in [robots.txt](https://kennethassogba.github.io/robots.txt). Content Signals declare these preferences to crawlers that support them. ## WebMCP In browsers with WebMCP support, `search_content` searches the notes and publications, and `read_page` reads a page as Markdown. Both tools are read-only. ## Hosting On GitHub Pages, use the explicit Markdown URLs. The optional Cloudflare adapter supports `Accept: text/markdown` and HTTP discovery headers. WebMCP support depends on the browser. [Read the implementation note](https://kennethassogba.github.io/notes/a-website-for-people-and-agents) [Read as Markdown](https://kennethassogba.github.io/agents/index.md) Canonical: https://kennethassogba.github.io/agents.html --- # Making this website AI-native Markdown pages, WebMCP tools, and content discovery based on Cloudflare’s recommendations. Author: Kenneth Assogba Date: 2026-10-05 Topic: AI & agents URL: https://kennethassogba.github.io/notes/a-website-for-people-and-agents I’m making this website AI-native, following [Cloudflare’s agent-readiness recommendations](https://blog.cloudflare.com/agent-readiness/) and [Double Slash’s implementation](https://double-slash.dev/articles/is-it-agent-ready/). For this site, that means an agent can find the content, read it as Markdown, and search the notes and publications. Here’s what I added. ## Markdown pages The notes are written in Markdown. The build generates full HTML pages with React and shadcn/ui, plus a Markdown version of each page. The HTML already contains the content before JavaScript runs. Every article has a **Read as Markdown** link and a copy button. The homepage is available at [/index.md](https://kennethassogba.github.io/index.md); this note is at [/notes/a-website-for-people-and-agents/index.md](https://kennethassogba.github.io/notes/a-website-for-people-and-agents/index.md). ## Content discovery The build generates these files: - [/llms.txt](https://kennethassogba.github.io/llms.txt): a list of pages and Markdown URLs; - [/llms-full.txt](https://kennethassogba.github.io/llms-full.txt): the complete text; - [/api/content.json](https://kennethassogba.github.io/api/content.json): titles, dates, topics, URLs and plain text; - [/api/profile.json](https://kennethassogba.github.io/api/profile.json): my profile and projects; - [/sitemap.xml](https://kennethassogba.github.io/sitemap.xml): the public pages; - [/feed.xml](https://kennethassogba.github.io/feed.xml): the RSS feed for new notes. The JSON endpoints are listed in [the API catalog](https://kennethassogba.github.io/.well-known/api-catalog), with an [OpenAPI description](https://kennethassogba.github.io/api/openapi.json). They are static files and don’t require an account or API key. ## WebMCP tools The search button and ⌘ / Ctrl K open a local search over the writing. The query stays in the browser. In browsers with WebMCP support, the page registers two tools: - `search_content`: searches the notes and publications using the same function as the search interface; - `read_page`: reads the Markdown of a known page on this site. Both tools are read-only. WebMCP is still experimental, so the Markdown files and JSON endpoints are also available directly. ## Content Signals I allow search, AI input, and model training. These preferences are declared in [robots.txt](https://kennethassogba.github.io/robots.txt): ```text Content-Signal: search=yes, ai-input=yes, ai-train=yes ``` Crawlers that support Content Signals can read these preferences. The signals themselves don’t enforce them. ## Hosting GitHub Pages serves the HTML and Markdown files. It can’t run the server code needed to return Markdown from an HTML URL using `Accept: text/markdown`, so agents should use the explicit Markdown URLs there. I included an optional Cloudflare Worker adapter for content negotiation. It also adds HTTP `Link` discovery headers and `Vary: Accept` so caches distinguish HTML from Markdown. The adapter is in the repository and isn’t deployed. DNS-based discovery would require a domain I control. The `github.io` subdomain doesn’t provide that access. The [For agents page](https://kennethassogba.github.io/agents.html) lists the resources. The build tests check the content exports, the browser tools, and content negotiation. ## References - [Cloudflare: Agent readiness](https://blog.cloudflare.com/agent-readiness/) - [Double Slash: Is it agent ready?](https://double-slash.dev/articles/is-it-agent-ready/) - [Is it agent ready? — scanner](https://isitagentready.com/) - [Content Signals](https://contentsignals.org/) - [RFC 9727 — API catalog discovery](https://www.rfc-editor.org/rfc/rfc9727.html) - [WebMCP proposal](https://github.com/webmachinelearning/webmcp) --- # Agent skills for a large codebase Documenting the codebase and engineering practices for coding agents. Author: Kenneth Assogba Date: 2026-10-05 Topic: AI & agents URL: https://kennethassogba.github.io/notes/repository-knowledge-as-agent-skills I work on the compiler for FPGA prototyping at Siemens EDA, mainly on placement and partitioning. I’ve also started working on netlist qualification, including clock handling. Alongside that work, I write agent skills and integrate MCP servers into our development workflow. The skills document the codebase and our engineering practices. A coding agent needs that information to work on an existing project: where to make a change, which constraints matter, and how to test it. ## What goes into a skill For a development task, a skill should answer: - Where does this kind of change belong? - Which constraints does the surrounding system rely on? - What is a good example already in the repository? - Which checks should run after the change? - When is it time to ask a person instead of guessing? Useful instructions name the relevant subsystem, point to an existing implementation, and give the command for the relevant checks. ## MCP integrations A skill contains instructions. An MCP server exposes tools. I use skills to explain how to work in the repository, and MCP integrations to connect tools used during development. The instructions should say when to use a tool and how to check its result. ## Review and testing I use AI to assist development. The resulting changes still need code review, tests, and performance checks. I’m also [making this website AI-native](https://kennethassogba.github.io/notes/a-website-for-people-and-agents), with Markdown pages and tools for agents to search and read the content. --- # Building La Bulle Scaffolding a voice coaching app with Codex, OpenAI Realtime, Cloudflare Workers, D1, and Resend. Author: Kenneth Assogba Date: 2026-09-27 Topic: AI & agents URL: https://kennethassogba.github.io/notes/building-la-bulle [La Bulle](https://bulle.hodge-podge.workers.dev/?lang=en) is a voice coaching app I built with Séb and Fano for the X-IA hackathon. Séb brought the Kedo micro-coaching protocol: 14 questions, asked one at a time, with room to think. You describe a situation, talk it through, then get an editable recap. You can keep a few notes for the next session, send the recap by email, or continue in Notion. I worked with Codex on the development. Here’s how the project took shape over September 26 and 27. ## Scaffold I started by writing down the user flow, the architecture, and the expected behavior during a call. Those documents became the basis for the implementation: how a session starts, what happens when someone asks for time, what gets saved, and which actions need a click from the person using the app. The scaffold was small: ```text public/ HTML, CSS, and browser JavaScript worker/ TypeScript API and model calls migrations/ D1 schema changes tests/ API and browser-controller tests docs/ Architecture, behavior, and test scenarios wrangler.jsonc Cloudflare configuration ``` The frontend uses plain HTML, CSS, and JavaScript. A Cloudflare Worker serves the static files and handles `/api/*`. D1 stores sessions, messages, approved notes, and the state of background jobs. The Worker calls the provider APIs with `fetch`. Wrangler runs the Worker and D1 locally, applies the database migrations, and deploys the app. The build command runs a dry-run deployment; publishing is a separate manual step. The first version already had text coaching, a voice call, and notes. From there, I worked through the call behavior, added French and English, and built the Notion flow. The editable recap and email came next. I kept the docs alongside the code as these decisions changed. ## Models The app uses three OpenAI models: | Model | What it does | | --- | --- | | `gpt-4.1-mini` | Text coaching, proposed notes, recaps, and the three Notion agents, through the Responses API | | `gpt-realtime-2.1` | The voice conversation over WebRTC, with the `marin` voice | | `gpt-4o-transcribe` | Transcription of the person’s speech during the call | For text, the Worker sends the recent conversation and the notes the person has approved. Responses use a strict JSON schema. The text coach returns a reply and a conversation decision: clarify, rephrase, explore, or finish. The app uses that decision to handle the end of a session. Voice has a separate path. The browser sends its WebRTC offer to the Worker, which creates the Realtime call with OpenAI and returns the answer. After that, audio travels directly between the browser and OpenAI. The API key stays in the Worker. Transcription supplies the written record. It doesn’t control when the voice model answers, and a failed transcription doesn’t stop the call. The app saves the received transcript when the call ends; it doesn’t record audio files. ## Getting the call to behave properly The first implementation mixed browser timers with Realtime turn-taking. It broke the conversation. I [removed the browser turn controller](https://github.com/kennethassogba/hodge-podge/commit/b6666f83099242b719d0a276ce70700c17412aa8) and let Realtime handle ordinary turns: ```json { "type": "semantic_vad", "eagerness": "medium", "create_response": true, "interrupt_response": true } ``` The browser requests the greeting. After that, Realtime detects the end of a turn and creates the next response. The person can interrupt the coach without waiting for a button. An explicit request for a pause needs different behavior. If someone says “give me a moment”, the model can call `pause_coaching`. The app lets the acknowledgement finish, then starts a timer. It asks once whether the person is ready to continue. If they start speaking first, it cancels the timer. If they ask to resume on their own, it doesn’t schedule a reminder. I also had to handle incomplete responses. The original 300-token output limit could cut off a spoken explanation. I raised it to 2,048 tokens and added one retry for a response truncated by the token limit or a temporary server error. Before retrying, the app removes the incomplete output from the conversation context and saved transcript. A normal interruption by the person doesn’t trigger that retry. Hanging up closes the microphone immediately. Saving the transcript happens afterward, with retries for temporary failures. The app waits for a confirmed save before opening the recap. On mobile, it requests a screen wake lock during the call and releases it afterward. That prevents automatic screen sleep when the browser permits it; it doesn’t make the call work with the phone locked. These were separate changes, each with a specific scenario to reproduce. “Improve the voice experience” would have been too vague to implement or verify. ## Recaps and email The recap is generated from the selected conversation. It distinguishes decisions from ideas that came up, and can say that no action was decided. The person can edit it before copying it, sending it, or using it in Notion. The recap and the coach’s memory are separate. Generating a recap doesn’t silently add it to the next session. A proposed note only becomes memory after the person chooses to keep it. For email, the Worker calls Resend using a server-side key and a verified sending domain. It sends the edited recap, with a UTF-8 transcript attachment if requested. The email module is small: it builds the plain-text and escaped HTML bodies, adds the optional attachment, and calls the Resend API. A fingerprint of the content supplies an idempotency key, so repeated requests for the same email reuse the key. D1 records the send status. The app doesn’t keep the recipient address or email body in that send record. A response from Resend confirms acceptance, rather than delivery to the inbox. ## Continuing in Notion Notion is optional. Coaching, the recap, and email work without connecting it. After OAuth, the person selects up to three pages and submits the edited recap as their intention. The full coaching transcript isn’t passed to this flow. Three sequential model calls then do the work: 1. Read the selected material and investigate the intention, citing exact excerpts. 2. Check the findings against those sources. Unsupported findings stop the job and request more context. 3. Prepare a proposal the person can edit, such as a meeting outline or a decision rule. The Worker checks that cited excerpts exist in the retrieved text. The job’s stage and results live in D1. A scheduled Worker resumes pending jobs, and a database lease prevents two executions from advancing the same job at once. Publishing is a separate request after the person has reviewed the proposal. It creates a new Notion page. The agents don’t edit the source pages. ## Testing as I went The API tests use Miniflare with a real local D1 database and simulated provider responses. They cover ownership checks, approved memory, recap edits, email attachments, duplicate sends, and the Notion publication flow. Browser tests run the actual voice controller with simulated WebRTC events, including late events, interruption, pause cancellation, and hangup during a retry. Those tests can check application behavior without spending API credits. They can’t tell me whether a spoken exchange sounds right. For that, I added separate scripts that exercise the real Realtime connection with synthetic audio, plus a browser harness that injects audio into the app’s WebRTC path. The development loop was concrete: reproduce a problem, change the relevant behavior with Codex, run the local checks, then try the call again. The scaffold got the app running. Most of the following work was in the parts between model calls: turn-taking, cancellation, saving, retries, and making sure the person’s edits were used. The [source code](https://github.com/kennethassogba/hodge-podge) includes the [architecture](https://github.com/kennethassogba/hodge-podge/blob/main/docs/architecture.md), [voice behavior](https://github.com/kennethassogba/hodge-podge/blob/main/docs/silence.md), and [recap and email implementation](https://github.com/kennethassogba/hodge-podge/blob/main/docs/apres-bulle.md). --- # Developing cmake2graph CMake Dependency Visualization. Author: Kenneth Assogba Date: 2025-03-27 Topic: CMake URL: https://kennethassogba.github.io/notes/cmake2graph ## A Journey into CMake Dependency Visualization As software projects grow larger, understanding dependencies between components becomes increasingly challenging. This is especially true for C++ projects using CMake, where target dependencies can quickly become complex. This led me to develop `cmake2graph`, a tool that visualizes CMake target dependencies as directed graphs. ## The Problem Working on large C++ codebases, I often encountered these challenges: - Difficulty understanding dependency relationships - Circular dependencies causing build issues - Complex CMake files with unclear target relationships - Hard to spot unnecessary dependencies ## The Solution: cmake2graph `cmake2graph` is a Python tool that: 1. Parses CMake files recursively 2. Extracts target dependencies 3. Builds a directed graph 4. Visualizes the relationships 5. Provides filtering options Here's a simple example: ```bash cmake2graph /path/to/project --output deps.png ``` ## Technical Implementation The tool uses several key technologies: - **NetworkX**: For graph creation and manipulation - **Matplotlib**: For visualization - **CMake Parser**: Custom implementation to extract dependencies ### Key Features - **Recursive Parsing**: Handles nested CMake files - **Dependency Filtering**: Focus on specific targets - **Depth Control**: Limit dependency chain depth - **External Library Filtering**: Focus on project-specific targets - **Multiple Output Formats**: Support for PNG, SVG, PDF ## Lessons Learned 1. **CMake Complexity**: CMake's flexibility makes parsing challenging 2. **Graph Layout**: Finding the right balance between aesthetics and clarity 3. **Performance**: Handling large projects requires optimization 4. **User Experience**: Balancing features vs. simplicity ## Future Development Several exciting possibilities lie ahead: ### 1. Automatic Link Error Resolution Currently planning to add features that: - Analyze linking errors - Suggest missing dependencies - Automatically fix common linking issues ```cmake # Before: Link error target_link_libraries(app core) # After: Automatically fixed target_link_libraries(app PRIVATE core missing_dependency ) ``` ### 2. CMake File Cleanup Future versions could: - Detect unused targets - Remove redundant dependencies - Standardize CMake syntax - Enforce modern CMake practices ### 3. Dependency Analysis Planning to add: - Cycle detection and breaking - Dependency impact analysis - Build time optimization suggestions - Target visibility recommendations ### 4. Integration Features Looking to integrate with: - IDE plugins - CI/CD pipelines - Build systems - Static analyzers ## Contributing The project is open source and welcomes contributions. Key areas where help is needed: - CMake parsing improvements - Graph visualization enhancements - Documentation - Test coverage - New feature implementation ## Conclusion `cmake2graph` started as a simple visualization tool but has grown into a platform for CMake dependency management. While it currently serves its basic purpose well, the potential for growth is significant. The future roadmap focuses on making C++ dependency management more maintainable, visual, and automated. Whether you're managing a small project or a large codebase, understanding and optimizing dependencies is crucial for maintainable software. ## Links - [GitHub Repository](https://github.com/kennethassogba/cmake2graph) - [PyPI Package](https://pypi.org/project/cmake2graph) --- # Write a header-only object oriented interface around MPI Write interface around MPI. Author: Kenneth Assogba Date: 2023-05-07 Topic: C++, MPI URL: https://kennethassogba.github.io/notes/write-interface-mpi **`tl;dr: I have written a simple header-only MPI interface in C++.`** The developments are availible here [human.mpi](https://github.com/kennethassogba/human.mpi) MPI or Message Passing Interface, is a standard for writing parallel programs in a distributed environment. It has become a reference in scientific computing and many legacy computing codes are progressively integrating MPI in view of the move to distributed computing architectures. ## Problem The addition of MPI communications within an existing computational code can lead to difficulties in readability and maintainability. This is partly because the physics (or math) + communications code mix is difficult to read. A good way to integrate MPI into existing code can be to encapsulate the MPI functions in a class with a simple interface. This is what Boost::MPI offers, for example. ## Proposal I have written a header-only interface - so easy to integrate in an existing code - which provides a simpler way to write MPI messages. This interface handle some of the more complex details of the library thus making it easier for developers to write parallel programs. In addition it is easier to port existing MPI programs to different platforms or environments, as the wrapper provide a consistent interface that is independent of the underlying implementation of MPI. ## A simple broadcast example ```cpp #include #include #include "human/mpi.hpp" int main() { human::mpi::communicator world(); auto rank = world.rank(); auto size = world.size(); auto root = world.root(); std::cout << "Process " << rank << "/" << size << std::endl; std::string msg; if (rank == root) msg = "Hello"; world.bcast(msg); std::cout << "Process" << rank << " " << msg << std::endl; return 0; } ``` Here, `msg` is sent to all non-root processes (0 by default). In reality the sending is done in two steps. First the size is broadcasted and the non-root resize the `msg` to the size received. Finally the `msg` content is sent. When the `communicator` instance goes out of scope (e.g., at the end of the `main` function), the destructor will be called, which will finalize the MPI library. The equivalent in pure MPI would be ```cpp #include #include #include int main(int argc, char* argv[]) { MPI_Init(&argc, &argv); int rank = 0; MPI_Comm_rank(MPI_COMM_WORLD, &rank); int size = 0; MPI_Comm_size(MPI_COMM_WORLD, &size); int root = 0; std::cout << "Process " << rank << "/" << size << std::endl; std::string msg; if (rank == root) msg = "Hello"; int msg_size = msg.size(); MPI_Bcast(&msg_size, 1, MPI_INT, root, MPI_COMM_WORLD); if (rank != root) msg.resize(msg_size); MPI_Bcast(const_cast(msg.c_str()), msg_size, MPI_BYTE, root, MPI_COMM_WORLD); std::cout << "Process" << rank << " " << msg << std::endl; MPI_Finalize(); return 0; } ``` The line `world.bcast(msg)` turns into at least 4 lines of code. ## A simple point-to-point communication Here is an example of how to use the wrapper to `send` a message between two processes. ```cpp std::string msg_sent, msg_recv; int other; if (world.rank() == world.root()) { msg_sent = "Hello"; other = 1; } else { msg_sent = "world!"; other = 0; } auto tag = 1; world.send(msg_sent, other, tag); world.recv(msg_recv, other, tag); std::cout << "P" << rank << " " << msg_sent << " " << msg_recv << std::endl; ``` There should be no deadlock problem as the messages are quite small. In the case of larger messages it is more appropriate to use non-blocking communications. ## Wrapping up I have write GitHub Actions to test the code on push and pull request. There is more to do, including writing tests and future developments are listed in the [roadmap](https://github.com/kennethassogba/human.mpi#roadmap). --- # Async communications are effective Async communications. Author: Kenneth Assogba Date: 2023-04-05 Topic: MPI URL: https://kennethassogba.github.io/notes/async-communications Status: Draft **`tl;dr: On a good cluster, with enough local work, the waiting time following async communications is negligible.`** ## Context In the paper I submitted to the mc conference, I apply a simple domain decomposition method for neutron transport solver [link]. The subdomains are coupled together at the innermost level of resolution, i.e. during the matrix-vector products. So there is a lot of communication and one would expect the overhead to be high. The communications are however made in non-blocking mode [Algorithm], and the waiting time measured is generally negligible [Table]. Let us present here the waiting times measured on a case with 900 million unknowns. ## Algorithm The Matrix-Vector product from the point of view of a subdomain ```cpp for(const auto& subdomain : domain_neighbors) { mpi::request_in[i] = mpi::world.irecv(x_upwind); // async // Copy outgoing part of x in the x_out buffer mpi::request_out[i] = mpi::world.isend(x_out); // async i++; } y += A_diag * x; // local work mpi::world.waitall(request_in); mpi::world.waitall(request_out); y += A_offd * x_upwind; ``` ## Cluster Topaze @TGCC - 864 nodes, each housing two - AMD EPYC 7763 2.45 GHz sockets - equipped with 64 cores each => 128 cores/node - Interconnect: InfiniBand HDR-100 network ## Results with no asynchonous progress (Draft) Table wait (request_in, request_out) time on process 0, last/2, last. number of communication or wait The waiting time on the process of rank time 0, last/2, last with last=size-1. ## Next - Enable [async progress](https://www.intel.com/content/www/us/en/docs/mpi-library/developer-guide-linux/2021-6/asynchronous-progress-control.html) - Try non-bloquing collectives. - Try a pipelined linear solver, e.g [pipelined BiCGstab](https://www.sciencedirect.com/science/article/abs/pii/S0167819117300406). --- # Send an object via MPI using serialization Send an object via MPI using serialization. Author: Kenneth Assogba Date: 2023-03-27 Topic: C++, MPI URL: https://kennethassogba.github.io/notes/mpi-using-serialization **`tl;dr: Let us discuss about MPI and serialization.`** To send an object instance with MPI, there are three main options: - Serialize the object into a byte string and send that. This involves converting the object into a binary format that can be reconstructed elsewhere. - Send each attribute of the object separately and reassemble it at the receiving end. - Register the object as an MPI data type. This defines how the object should be laid out in memory so MPI knows how to send its constituent parts. ## Serialization Let us discuss here about serialization. One uses the Boost.Serialization library to convert the object into a byte stream. Then `MPI_Send` can be used to send the stream, and `MPI_Recv` used to receive it. Finally, arrived at its destination, one deserialize the stream into an object again. (In progress) Present simple exemple of serialization → send → deserialization ## Ressources - [How to send a set object in MPI_Send](https://stackoverflow.com/questions/31014044/how-to-send-a-set-object-in-mpi-send) - [Can't get C++ Boost Pointer Serialization to work](https://stackoverflow.com/questions/28901596/cant-get-c-boost-pointer-serialization-to-work) - [boost serialization of dynamic arrays](https://stackoverflow.com/questions/21408521/boost-serialization-of-dynamic-arrays) - [Serialization/Deserialization of a Vector of Integers in C++](https://stackoverflow.com/questions/51230764/serialization-deserialization-of-a-vector-of-integers-in-c) --- # The Pn form of the Neutron Transport Problem Achieves Linear Scalability Through Domain Decomposition Large-scale neutron transport simulation Author: Kenneth Assogba Authors: Kenneth Assogba, Lahbib Bourhrara Date: 2023 Topic: Conference proceeding URL: https://kennethassogba.github.io/publications/mc-2023 - [Paper](https://kennethassogba.github.io/assets/docs/mc_2023.pdf) - [Conference](https://mc2023.com/) ![Distributed matrix-vector product](https://kennethassogba.github.io/assets/img/matrix_distributed-spmv.svg "Distributed matrix-vector product") Due to strong coupling between the angular moments, the spherical harmonics (Pn) formulation of the neutron transport equation has been neglected in favor of the discrete ordinates (Sn) form. In this work, we target large-scale neutron transport simulation using a combined discontinuous Galerkin (DG) - spherical harmonics approximation. We leverage the benefits of DG discretization to wrap the previous developed solver, called NYMO, in a canonical ghost-mesh based domain decomposition framework. The developed solver handles unstructured, curved and non-conforming meshes with vacuum and reflection as boundary conditions. Robustness, strong and weak scalability experiments have been conducted on the CEA's pre-exascale system Topaze. We reach and maintain a strong scaling efficiency of 100% up to 4096 cores and 80% up to 8192 cores. In particular, a calculation with 913 million degrees of freedom is performed in 101 seconds. Thus outperforming previously published results for Pn transport as well as many Sn and SPn solvers. ## Subdomains coupling The general Matrix-Vector product from the point of view of a subdomain ```cpp for(const auto& subdomain : domain_neighbors) { mpi::request_in[i] = mpi::world.irecv(x_upwind); // async // Copy outgoing part of x in the x_out buffer mpi::request_out[i] = mpi::world.isend(x_out); // async i++; } y += A_diag * x; mpi::world.waitall(request_in); mpi::world.waitall(request_out); y += A_offd * x_upwind; ``` ## Weak scaling experiment | Partitioning | | | | Simple | | | Geometric | | |------:|-------:|------:|--:|--------------------:|----------------:|--:|-----------------------:|----------------:| | $n_d$ | \#core | \#dof | | time (s) | efficiency (\%) | | time (s) | efficiency (\%) | | 16 | 1024 | 228M | | 145 | - | | 142 | - | | 32 | 2048 | 456M | | 154 | 94 | | 157 | 90 | | 64 | 4096 | 913M | | 146 | 99 | | 147 | 96 | ## Strong scaling experiment Strong scaling experiment on up to 128 domains using a total of 8192 CPU-cores. 0 means shared memory only calculation. | Partitioning | | | Simple | | | | Geometric | | | |---------------:|-------:|---|--------------------:|--------:|----------------:|--:|-----------------------:|--------:|----------------:| | $n_d$ | \#core | | time(s) | speedup | efficiency(\%) | | time(s) | speedup | efficiency(\%) | | 0 | 64 | | 10597 | - | - | | 10597 | - | - | | 4 | 256 | | 2557 | 4.1 | 104 | | 2658 | 4 | 100 | | 8 | 512 | | 1242 | 8.5 | 107 | | 1245 | 8.5 | 106 | | 16 | 1024 | | 667 | 15.9 | 99 | | 628 | 16.9 | 106 | | 32 | 2048 | | 302 | 35.1 | 110 | | 305 | 34.7 | 109 | | 64 | 4096 | | 146 | 72.6 | 113 | | 147 | 72.1 | 113 | | 128 | 8192 | | 101 | 104.9 | 82 | | 108 | 98.1 | 77 | --- # Spherical Harmonics and Discontinuous Galerkin Finite Element Methods for the Three Dimensional Neutron Transport Equation: Application to Core and Lattice Calculation Spherical harmonics and discontinuous Galerkin for the neutron transport equation. Author: Kenneth Assogba Authors: Kenneth Assogba, Lahbib Bourhrara, Igor Zmijarevic, Grégoire Allaire, Antonio Galia Date: 2023 Topic: Journal Paper URL: https://kennethassogba.github.io/publications/nse-2023 - [Paper](https://kennethassogba.github.io/assets/docs/nse_2023.pdf) - [Link](https://www.tandfonline.com/doi/abs/10.1080/00295639.2022.2154546) ![3D pin-cell](https://kennethassogba.github.io/assets/img/3d-mesh.svg "3D pin-cell") We combine spherical harmonics and discontinuous Galerkin to discretize the neutron transport equation. The method can handle all geometries describing the fuel elements without any simplification nor homogenisation. Moreover the use of matrix assembly-free method avoids building large sparse matrices, which enables to produce high-order solutions in small computational time and less storage usage. The resulting transport solver has a wide range of applications: it can be used for a core calculation as well as for a precise 281-groups lattice calculation accounting anisotropic scattering. --- # Precise 3D Reactor Core Calculation Using Spherical Harmonics and Discontinuous Galerkin Finite Element Methods Precise 3D Reactor Core Calculation Using Spherical Harmonics and Discontinuous Galerkin Finite Element Methods Author: Kenneth Assogba Authors: Kenneth Assogba, Lahbib Bourhrara, Igor Zmijarevic, Grégoire Allaire Date: 2022 Topic: Conference proceeding URL: https://kennethassogba.github.io/publications/physor-2022 - [Paper](https://kennethassogba.github.io/assets/docs/physor_2022.pdf) - [Link](https://www.ans.org/pubs/proceedings/article-51104/) - [Conference](https://www.ans.org/meetings/physor2022/) ![Takeda 3 core](https://kennethassogba.github.io/assets/img/takeda3.svg "Takeda 3 core") We study the use of Pn method in angle and Discontinuous Galerkin in space to solve 3D neutron transport problem. Pn method consists in developing the angular flux on truncated spherical harmonics basis. In this paper, we couple this method with the discontinuous finite elements in space to obtain a complete discretisation of the multigroup neutron transport equation. To investigate its precision, the method was applied to Takeda and C5G7 benchmark problems. These calculations point out that the proposed Pn -DG method is capable of producing accurate solutions in small computational time, and that it is able to handle complex 3D geometries.