Sign in to view Moses’ full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Moses’ full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Cupertino, California, United States
Sign in to view Moses’ full profile
Moses can introduce you to 10+ people at Apple
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
32K followers
500+ connections
Sign in to view Moses’ full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Moses
Moses can introduce you to 10+ people at Apple
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Moses
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Moses’ full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Activity
32K followers
-
Moses Pawar shared thisA slightly different post from what I normally share here. Today I am a very proud father. Ethan’s first song, “Orbiting You,” is now live on Apple Music and Spotify. He is in 9th grade, collaborated with vocalist Arshiya Ghosh, and played drums on the recording. The song was released through Spotlife Studio Records. Apple Music: https://lnkd.in/g5xCByxD Spotify: https://lnkd.in/g-KRUTqd Ethan has been learning music and playing drums since he was four years old. We have watched him spend years practicing, learning different styles and growing as a musician. It is special to see that work turn into a recording that people can now listen to. There is also a satisfying connection for me personally. I work closely with the Apple Music team, and they make extensive use of the AI and Data Platforms that my teams build at Apple. Seeing my 9th grader’s first release appear on Apple Music connects my work and family life in a way I did not expect. I am also glad his first release came through a collaboration. Music requires listening to others, adapting and contributing your part to something larger. Those are useful lessons well beyond music. Congratulations to Ethan, Arshiya and everyone at Spotlife Studio who helped bring “Orbiting You” to life.
-
Moses Pawar posted thisMeta-Harness This is a continuation of my thread on agentic system optimization. In this one I want to show how Meta-Harness works in practice. Quick recap of the setup. A harness, as we know, is the code and tools around a fixed model. It decides what to store, what to retrieve and what to show the model at each step. The optimizer here is called the proposer. It is a coding agent (Claude Code with Opus 4.6) that reads past attempts and writes new versions of the harness. Every past attempt is saved as a folder with its code, its score and its execution traces. Now to the part of the paper I found most interesting. This is a log from a 10-iteration run on TerminalBench-2, which is a benchmark of hard command-line tasks. The starting harness scored 64.4%. a) The proposer read a lot before it changed anything. In each iteration it opened between 69 and 99 files, 82 on average. About 41% of those were the code of earlier attempts and 40% were execution traces. Only 6% were score files. b) In iterations 1 and 2, it made two bug fixes. The first one removed leftover terminal markers that were confusing the agent and sending it into loops. The second one fixed a completion check that made the agent re-verify its work for 15 to 40 extra steps after it was already done. Both attempts also came with a new prompt. Both got worse, scoring 58.9% and 57.8%. c) In iteration 3, it compared the two failed attempts and looked at what they had in common. The bug fixes were different, but the new prompt was the same in both. It figured out that the new prompt was making the agent delete files it still needed, and that the bug fixes had been tested together with this bad change. So it went back to the original prompt and tested only the bug fixes. That attempt scored 63.3%, just 1.1 points below where it started, which showed its reasoning was right. d) After that, it kept trying smaller fixes. It finally landed on a safer change that added new behavior without removing anything, and that became the best harness in the run. This is how I would want an engineer on my team to debug. Two changes fail, you look at what they share, you pull that part out and you test again. This example shows the key principles of Meta-Harness. a) Keep the full history. Every past attempt, score and trace is saved. Without the two failed attempts, the proposer could not have spotted what they had in common. b) Give the proposer raw traces, not summaries. c) Let the proposer decide what to read. The folder is too big to fit in one prompt, so the proposer searches it and opens only what it thinks matters. d) Keep the search loop simple. There is no fixed rule for which earlier attempt to build on.
-
Moses Pawar shared thisContinuing my thread on agentic systems optimization. Harness optimization is all the rage so lets look at one technique. This one was specifically focussed on prompt and middleware optimization. BTW, when I think of harness I include prompt, tool interfaces, state handling and recovery logic. The optimizer, PRISM from a recent Airbnb paper , uses four techniques that I found useful. It is a evolutionary technique on the lines of GEPA but adapted to harnesses. It keeps a Pareto frontier of harnesses during the search. A harness stays on the frontier if no other harness is at least as good on both measures (task and reliability e.g. not getting stuck in loop) and better on one. a) It uses a three-way data split. One split is used to propose changes, a second is used to select the final harness, and a third is held out and used only to measure the final result. This separates the search process from the reported improvement. Say there are 300 past customer support conversations where the agent failed. The first 100 go into the first split. The optimizer reads these failures and proposes changes to the harness, such as a new rule in the prompt. The next 100 go into the second split. Each proposed harness runs on these conversations, and the one with the highest pass rate is kept. The last 100 go into the third split and are not used during the search. When the search is done, the chosen harness runs on these 100 conversations once, and that is the reported result. b) It measures worst-case lift in addition to average improvement by running the optimizer multiple times, selecting the harness that performs best on the validation split, and then measuring the 5th percentile improvement on held-out data c)It limits middleware changes to a small set of patterns. Middleware sits between the model and its tools. It can fix malformed arguments before the tool runs. It can reject an invalid call and ask the model to try again. It can also block a call until a required step has happened. For example, if the model sends expression but the tool expects math_expression, the middleware can rename the field before executing the call. d) It routes failures to the layer that should fix them. PRISM clusters failed trajectories by root cause and decides whether the change belongs in the prompt, middleware, or both. Where this could be improved so this can be more widely used 1. Only few benchmarks were covered. 2. The scope was limited to few middleware changes and feels like a proof-of-concept
-
Moses Pawar shared thisContinuing my thread on agentic systems optimization I found AgentFactory concept interesting. It is a recent paper to optimize the model and the agent workflow together, rather than treating the underlying model as fixed and only searching over prompts or workflow structure. a) AgentFactory optimizes both model configuration and agent workflow. The model side can include the foundation model, fine-tuning data, tuning method and hyperparameters, while the workflow is represented as executable code. b) An LLM acts as the optimizer. At each iteration it uses the task, performance targets and results from previous experiments to propose a new combination of model tuning and workflow design. c) The system can optimize multiple objectives such as accuracy, cost and latency. Across eight benchmarks covering reasoning, coding, mathematics, medicine and finance, the they claim average 9.1% improvement over existing automated agent-design methods. d) The two closest implementations I found in practice are Microsoft Agent Lightning and DSPy with an RL backend such as Arbor. Agent Lightning can train agents using RL or SFT and also supports automatic prompt optimization while keeping the existing agent harness, tools and control flow in the loop. DSPy provides prompt optimization through methods such as GEPA and MIPRO, while dspy. GRPO can optimize model weights through RL. Neither currently does the full AgentFactory optimization over model weights, prompts and workflow structure as one optimization problem. AWS and GCP now provide several pieces of this separately. Both have prompt optimization and reinforcement fine-tuning, and AWS AgentCore can use traces to optimize agent prompts and tool descriptions. I could not find a single managed AWS or GCP capability that jointly searches the agent workflow, prompts and model weights in one optimization loop.
-
Moses Pawar posted thisYesterday I wrote about GEPA which is a prompt optimizer that uses an LLM to read execution traces and evaluator feedback and rewrite one prompt at a time in a agentic system. By default, GEPA rotates through prompts , while choosing the candidate to mutate using an instance-level Pareto frontier over validation examples. In comparison AgentGrad another prompt optimizer focuses on two parts of multi-agent prompt optimization: deciding which agent to change and combining feedback from multiple failures. It works in four steps. a) For each failed example, AgentGrad performs sequential changes in reverse execution order. It adds a hint derived from the ground-truth answer to one agent at a time and reruns the downstream agents. The first intervention that makes the final system output correct identifies the agent whose prompt should be optimized. b) It compares that agent's original output with the output produced under the intervention. This difference is given to a LLM which converts the into a sample-level textual gradient describing how the prompt should change. c) AgentGrad groups similar gradients into semantic minibatches. An aggregator LLM converts each cluster into one what they call as generalized textual gradient (prompt change) representing a common failure pattern. d) A prompt optimizer generates a candidate prompt change edit from each generalized gradient. The edit is accepted only if it improves the corresponding semantic minibatch and also improves or maintains performance on the validation set. Observations a) GEPA chooses which prompts to optimize using round-robin selection. AgentGrad instead chooses individual agents and tests whether changing that agent's behavior improves the end-to-end failure. b) AgentGrad in my opinion changes local behavior and then tries to generalize it by grouping local changes. c) GEPA maintains multiple candidates rather than one evolving prompt set using Pareto frontier mechanism. It would be interesting to combine this with AgentGrad's approach.
-
Moses Pawar posted thisGEPA is a prompt optimizer for agentic systems that have more than one step. Here is how it works, with a simple example. The system checks claims about films against a movie database using two searches. One prompt summarizes the entries from the first search. A second prompt writes the query for the second search. Claim: "The director of 3 Idiots also directed Munna Bhai M.B.B.S." To check this, the system needs three entries: 3 Idiots, Munna Bhai M.B.B.S. and Rajkumar Hirani. The first search finds the two film entries. The summary covers the plot and cast of both films but does not mention the director. The second query is "Munna Bhai M.B.B.S. cast," which returns Sanjay Dutt's entry which is not useful. The system finds 2 of the 3 entries it needs. A normal evaluator returns a score of 0.67. GEPA's evaluator returns 0.67 plus a note: "Missing: the Rajkumar Hirani entry." GEPA then runs the current prompts on 3 training examples and collects each step's input, output and the evaluator notes. Gives these to an LLM, which rewrites one prompt. Lets say it adds a rule to the summary prompt: name every person the claim depends on, such as a film's director, even if the entries mention them only once. Reruns the same 3 examples. If the score goes up, it scores the new prompts on the full validation set and keeps them as a candidate. Picks the next candidate (prompts) to improve from the Pareto frontier across validation examples, not only the candidate with the best average. (The Pareto frontier is the set of options where improving performance on one example would mean doing worse on at least one other example. In GEPA, a prompt is on the frontier when no other prompt performs as well or better on every validation example, meaning each frontier prompt has some example where it is especially strong.) Observations: The context evaluator sends via text is more critical. "Missing: the Rajkumar Hirani entry" tells the LLM what to fix. A score of 0.67 does not. This got me thinking on how this would work with JEV models
-
Moses Pawar posted thisSuppose a student knows three valid ways to solve a problem. They can use algebra, draw a diagram, or reason through the numbers directly. Now imagine two ways of teaching the next set of problems. In the first approach, the teacher repeatedly shows one worked solution and asks the student to reproduce that method. The student gets better at the task, but may increasingly default to the demonstrated method even when the other approaches they already knew would work. In the second approach, the student attempts the problems using their existing methods. The teacher mainly provides feedback on whether the answer and reasoning are correct. If algebra works, it is reinforced. If the diagram works, that is also reinforced. Incorrect approaches are corrected. The student is still learning, but the learning process does not require replacing several useful strategies with one preferred strategy. The learning is more generalized. This is the type of distinction the RL’s Razor paper makes between SFT and on-policy RL. a. Suppose a model already has three successful ways to solve a coding problem. • Strategy A has 40% probability. • Strategy B has 35%. • Strategy C has 20%. • Incorrect behavior has 5%. With SFT, the training data may contain only Strategy B. The objective keeps increasing the probability of B, even though A and C were already valid solutions. The model improves on the new task, but its behavior can change more than necessary. b. With on-policy RL, the model generates trajectories from its own current behavior. If A, B, and C all succeed, all three can receive positive reward. The model can reduce the unsuccessful behavior while preserving much more of its existing distribution (probabilities) over successful strategies. This does not mean RL is inherently better than supervised learning. If we can construct an SFT datasets that preserved the model’s existing distribution over good behaviors while removing the bad ones, supervised learning could achieve the same objective. The broader principle is to change the model only as much as necessary to learn the new capability.
-
Moses Pawar posted thisAutoJev-27B ranks second on Hugging face on Jev benchmarks. Here is a quick summary of architecture - AutoJev-27B is a pretrained Qwen3.8-27B backbone with its generative vocabulary (LM head) readout replaced by a small bounded decision head. The decision head is a linear projection from the 5,120-dimensional final hidden state to 255 option logits. A logit is a raw score the model gives each option, and a softmax turns those scores into probabilities that add up to 1. a. Each of the 255 outputs maps to an option label such as A, B or C that the Qwen tokenizer encodes as a single token. If a question has three options, only those three logits are used and the rest are masked. b. The decision head is not randomly initialized. AutoJev copies the rows for the option label tokens from Qwen's original output layer. The model starts with weights that already relate to tokens such as A, B and C, and then the full model is fine-tuned. c. The context, the question and all options go into one input sequence. Qwen processes them together. AutoJev takes the hidden state at the last position and passes it to the decision head. d. A softmax over the active option logits, scaled by a calibrated temperature, gives the probability of each option. There is no decoding loop and no text is generated. e. The same setup handles yes/no questions.
-
Moses Pawar posted thisIceberg 1.12 introduces server-side scan planning work in Iceberg REST Catalog. I have mixed feeling not specifically about this change but moving more work towards a central control plans from engines. 1. What happens today a. Spark, Trino, Flink, and other engines use Iceberg client library for Iceberg-specific scan planning. b. The engine determines the SQL predicate and required columns, then passes them to Iceberg. c. The Iceberg client library runs inside the engine process and reads the snapshot, manifest list, manifests, file statistics, and delete-file metadata to produce FileScanTasks. d. The engine then converts those Iceberg tasks into Spark partitions, Trino splits, or the equivalent execution unit. 2. Server side planning a. The engine can send the snapshot, projection, filter, and other scan parameters to the REST Catalog. b. The catalog service performs the Iceberg scan planning and returns the resulting FileScanTasks. c. The API also supports asynchronous planning, polling, fetching large sets of scan tasks incrementally, cancellation, and returning temporary storage credentials. d. SQL optimization, join planning, split scheduling, Parquet execution, and other engine-specific work remain in Spark, Trino, or Flink. 3. Why this can be useful a. A centralized service can may be reuse decoded manifests, snapshot metadata, statistics, and indexes across many queries and engines. b. The REST service combines scan planning with authorization and credential vending, so clients can receive short-lived access only to the files required for a particular scan. c. Newer clients in lets say python etc also benefit because they no longer need the full Iceberg manifest-planning implementation locally. 4. Tradeoffs in my mind a. If we continue this trend clients ( e.g execute query) can now become concentrated in a shared service. Spark, Flink etc are heavily distributed systems and designed to scale. These distributed engines are also being made performant via projects like Comet/DataFusion. b. A failure or overload in the planning service could affect multiple engines at the same time. c. Small or already-cached tables may be faster to plan locally. d. Client and server versions must agree on expression semantics. Backward compatibility is important. To me, the important change is not that Iceberg is taking planning away from Spark or Trino. Much of the Iceberg-specific logic was already in the Iceberg client library. I am more concerned about what else we continue to move from engines to centralized control plane and impact of that on capacity management and optimizations that have existed and matured in distributed engines.
Experience & Education
-
Apple
******** ** *********** * ***** ***** ** ******** * ***** **** ********
-
****
****** ******** ** *********** * *************** ********* * ******* **
-
******* *********** *******
****** *********** ******* * **************
-
********** ** ******** **********
** ******** ******* undefined
-
-
**** ********* ** ******** **********
** ******** *******
-
View Moses’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Publications
-
Distributed musical performances: Architecture and stream management
ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP)
The DIP project investigates a versatile framework for the capture, recording, and replay of video, audio, and MIDI (Musical Instrument Digital Interface) streams in an interactive environment for collaborative music performance.
Other authors -
High resolution live streaming with the HYDRA architecture
ACM Computers in Entertainment (CIE)
HYDRA project (Highperformance Data Recording Architecture) focuses on the acquisition, transmission, storage, and rendering of high-resolution media such as H-quality video and multiple channels of audio.
Other authorsSee publication
Patents
-
FORMAT-AGNOSTIC STREAMING ARCHITECTURE USING AN HTTP NETWORK FOR STREAMING
Issued US 20120265853
Languages
-
Hindi
Full professional proficiency
-
Marathi
Native or bilingual proficiency
-
English
Native or bilingual proficiency
Recommendations received
28 people have recommended Moses
Join now to viewView Moses’ full profile
-
See who you know in common
-
Get introduced
-
Contact Moses directly
Other similar profiles
Explore more posts
-
Coby Benveniste
MarkeTeam.ai • 2K followers
State machines, state machines everywhere! If anyone is interested in #AI, building #AIAgents, or #Erlang and #Elixir, take a look at my latest talk from CodeBEAM Europe 2025 to see how we at MarkeTeam.ai build our agentic systems! And if you are interested in these kinds of things... meet me in Malaga at ElixirConf® EU 2026 in a couple of months and hear me talk about how we build distributed performance testing for our #LiveView platform! https://lnkd.in/dFXpqNCg
11
-
Bruno Aziza
IBM • 54K followers
The "Palantirization" of the everything. The a16z piece captures the current state of of 3 trends well. 𝟭️) 𝗧𝗵𝗲 𝗥𝗶𝘀𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗙𝗼𝗿𝘄𝗮𝗿𝗱-𝗗𝗲𝗽𝗹𝗼𝘆𝗲𝗱 𝗘𝗻𝗴𝗶𝗻𝗲𝗲𝗿 (𝗙𝗗𝗘): Software vendors are now embedding elite engineers directly into customer environments to get to production-ready AI faster. 𝟮) 𝗢𝘂𝘁𝗰𝗼𝗺𝗲-𝗕𝗮𝘀𝗲𝗱 "𝗕𝗶𝗴 𝗦𝘄𝗶𝗻𝗴𝘀": We’re moving away from seat-based licensing and toward mission-critical outcomes. 𝟯) 𝗧𝗵𝗲 𝗣𝗹𝗮𝘁𝗳𝗼𝗿𝗺 𝗦𝗽𝗶𝗻𝗲: High-touch delivery only scales if there is a reusable "spine" of data primitives underneath. Your platform needs to learn from every deployment. Piece at https://lnkd.in/g3enw9CU
13
5 Comments -
Venkatesh Sankaran
Saguna Consulting • 8K followers
Waymo raising capital at this scale reflects how the autonomous vehicle conversation has evolved. The problem is no longer model capability. It is operational depth. Autonomy today is constrained by questions of reliability, deployment economics, regulatory coordination, and public safety at scale. Once vehicles move beyond pilots and into daily city operations, the complexity shifts from algorithms to systems. Fleet management, incident handling, infrastructure variability, and long-term cost structures begin to dominate outcomes. Large funding rounds signal that the market understands this transition. Capital is being allocated not to prove feasibility, but to sustain execution across geographies and conditions. The next phase will separate technical achievement from commercial durability. Expansion will expose unit economics, insurance frameworks, and the true cost of maintaining consistency across environments. For the ecosystem, this is a familiar inflection point. Early breakthroughs are behind us. What remains is the harder work of turning autonomy into dependable infrastructure. That is where long-term value will be decided. #waymo #funding #tech
14
-
Romain Jourdan
Amazon Web Services (AWS) • 4K followers
"Listen to the Heroes" said Werner Vogels at his re:Invent 2024 Keynote. 🎙️ Here is a good opportunity, as our new episode landed with Chris Miller! I had the pleasure of sitting down with Chris —AWS Hero since 2021, AI Software Engineer at Workato, and one of the most authentic voices that I know in the Bay Area developer community. What started as an "accidental" DeepRacer win turned into a journey through classical AI, computer vision with DeepLens, and now cutting-edge multi-agent systems. Chris doesn't just build with AI—he thinks deeply about what it means to build responsibly with AI. 🔍 Here's what we explored: → The Road to re:Invent hackathon story (featuring AI imposters of Jeff Barr, Swami Sivasubramanian, and Werner Vogels) → Why "vibe coding" isn't enough—and what responsible AI development actually looks like → Multi-agent orchestration patterns, token management, and recursion limits → The reality of building in Silicon Valley during the AI boom → Kiro, autonomous agents and the future of AI-assisted development Chris shared something that resonated with me: "AI has extended my life expectancy as a tech worker. It's allowing me to do so much more than I ever would have been capable of on my own." But he's also clear-eyed about the challenges: security reviews, code quality, testing, and the importance of transparency when shipping AI-generated code. If you're navigating the intersection of AI and software development—or just curious about what's happening in the trenches—this conversation is worth your time. 🎧 Listen now on your favorite Podcast app or 📺 Watch on YouTube: https://lnkd.in/ekauKixs What's your experience with AI-assisted development? Are you seeing similar patterns in your work?
17
3 Comments
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content